Safety monitoring video intelligent analysis method based on multi-algorithm collaboration and unified architecture
This intelligent analysis method for security surveillance videos, employing a multi-algorithm collaboration and a unified architecture, utilizes a CNN-LSTM spatiotemporal fusion engine and attention mechanism to achieve deep joint representation of surveillance videos and weighted integration of the importance of multi-source features. This solves the problems of unclear event identification and risk assessment in existing technologies, thereby improving the accuracy and real-time performance of surveillance video analysis.
Patent Information
- Application Number
- CN202511157373.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-28
AI Technical Summary
Existing video surveillance systems are inadequate in terms of deep information fusion and unified decision-making among the results of multiple tasks, especially in terms of inaccurate target behavior analysis in dynamic scenes, which leads to unclear event identification and risk assessment.
A security surveillance video intelligent analysis method based on multi-algorithm collaboration and unified architecture is adopted. Spatiotemporal feature extraction and fusion are performed through the CNN-LSTM spatiotemporal fusion engine, and the results of multi-algorithm analysis are weighted and integrated by the attention mechanism. Event type identification and risk assessment are performed, and finally, structured security surveillance video analysis results are generated.
It improves the accuracy and anti-interference ability of event detection, reduces the false alarm rate, enhances the real-time performance and scalability of the system, and can reliably capture key behavioral information in complex scenarios.
Smart Images

Figure CN121033730A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent video monitoring, and particularly to a security monitoring video intelligent analysis method based on a multi-algorithm coordination and unified architecture. BACKGROUND
[0002] With the development of video monitoring technology and artificial intelligence algorithms, intelligent video analysis has gradually become one of the core application means in the fields of public safety, traffic management, industrial monitoring, etc. Traditional video monitoring systems mostly rely on manual patrol and rule-driven event detection mechanisms, which have problems such as low efficiency and slow response in actual application. In recent years, visual algorithms based on deep learning such as convolutional neural network (CNN) and long short-term memory network (LSTM) have been widely used in target detection, behavior recognition and anomaly detection tasks, effectively improving the automation and intelligence level of video analysis.
[0003] In the prior art, a single algorithm or a series structure is usually used to analyze the target and behavior in the monitoring video, which lacks the ability of coordinated optimization of multi-source analysis results, especially in the deep information fusion and unified decision-making between multi-task results. Especially when facing the variable target behavior in dynamic scenes, it is difficult to fully integrate spatial information, time sequence features and cross-task semantic clues, which leads to limited performance of the system in comprehensively judging the event type and risk level. SUMMARY
[0004] In view of the above existing problems, the present application is proposed.
[0005] Therefore, the present application provides a security monitoring video intelligent analysis method based on multi-algorithm coordination and unified architecture to solve the problem of inaccurate event recognition and unclear risk judgment caused by the lack of deep fusion and unified decision-making mechanism.
[0006] To solve the above technical problems, the present application provides the following technical solutions:
[0007] In a first aspect, the present application provides a security monitoring video intelligent analysis method based on multi-algorithm coordination and unified architecture, which comprises,
[0008] Collecting security monitoring video for preprocessing and using a CNN-LSTM spatio-temporal fusion engine to perform spatio-temporal feature extraction and fusion, and outputting a fusion feature map;
[0009] Performing target detection, behavior recognition and anomaly detection on the fusion feature map to obtain multi-algorithm analysis results;
[0010] Using a cross-module fusion mechanism based on an attention mechanism to perform weighted integration processing on the multi-algorithm analysis results to obtain a fusion event representation vector;
[0011] Perform event type recognition and risk level assessment on the fusion event feature vector to obtain event classification results and event risk levels;
[0012] Alarm decision is made on the event classification results and event risk levels to obtain alarm information, which is stored and visually displayed to obtain the structured security monitoring video analysis result.
[0013] As a preferred scheme of the security monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture, the security monitoring video includes multiple camera video sources, real-time frame data, timestamps and camera identifiers.
[0014] The preprocessing includes denoising, white balance correction and illumination equalization.
[0015] As a preferred scheme of the security monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture, the output fusion feature map has the following specific steps,
[0016] Based on the preprocessed historical security monitoring video, the CNN-LSTM spatio-temporal fusion engine is jointly fine-tuned by an end-to-end supervised learning method to obtain the trained CNN-LSTM spatio-temporal fusion engine.
[0017] The preprocessed security monitoring video is received by the trained CNN-LSTM spatio-temporal fusion engine, frame rate cutting is performed, and continuous frame sequences are generated.
[0018] Multi-scale convolution and pooling operations are performed on the continuous frame sequences to extract image texture, edge and target shape information, and form spatial feature sequences.
[0019] The inter-frame dependency of the spatial feature sequence is captured through the forget gate and the input gate to extract the time feature sequence of the target motion law and behavior evolution.
[0020] The spatial feature sequence and the time feature sequence are spliced and integrated in the channel dimension to output the fusion feature map.
[0021] As a preferred scheme of the security monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture, the target detection, behavior recognition and anomaly detection are performed on the fusion feature map to obtain multi-algorithm analysis results, and the specific steps are as follows,
[0022] Through the CNN-based target detection method, each target in the fusion feature map is recognized and the boundary box and class are located, and the target detection result is output.
[0023] The target detection result and the fusion feature map are tracked and the action sequence is analyzed by using an LSTM-based behavior recognition method, and a behavior recognition result is output.
[0024] Based on the behavior recognition result and the fusion feature map, the error of the behavior feature is calculated and compared with the preset abnormal threshold, and a multi-algorithm analysis result is output.
[0025] As a preferred scheme of the safety monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture, the multi-algorithm analysis result is weighted and integrated by using a cross-module fusion mechanism based on an attention mechanism to obtain a fusion event representation vector, and the specific steps are as follows,
[0026] The intermediate layer deep semantic features and the output layer log prediction values are spliced to obtain a multi-source perception fusion feature vector;
[0027] The multi-source perception fusion feature vector is subjected to attention weight calculation by using a cross-module fusion mechanism based on an attention mechanism to generate feature dimension attention weights;
[0028] According to the feature dimension attention weights, the multi-source perception fusion feature vector is weighted and summed to obtain a fusion event representation vector.
[0029] As a preferred scheme of the safety monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture, the fusion event representation vector is subjected to event type recognition and risk level evaluation to obtain event classification results and event risk levels, and the specific steps are as follows,
[0030] The high-level semantic features are extracted from the fusion event representation vector as event feature expressions, and the event type recognition and classification are performed on the event feature expressions to output the event classification results;
[0031] The event classification results, the event feature expressions and the historical risk data are combined to perform risk level evaluation, and the event risk levels are output.
[0032] As a preferred scheme of the safety monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture, the event classification results and the event risk levels are subjected to alarm decision to obtain alarm information, and the specific steps are as follows,
[0033] The classification results of the historical events and the historical event risk levels are summarized, compared and abstracted to extract the mapping relationship between the event features and the risk levels, and an alarm decision rule library is constructed;
[0034] Based on the event classification result and the event risk level, a corresponding alarm decision rule in an alarm decision rule library is matched, alarm trigger logic is executed, and alarm information is output.
[0035] As a preferred scheme of the security monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture, the structured security monitoring video analysis result is obtained through the following specific steps,
[0036] The alarm information is formatted to form a structured alarm record, and is stored.
[0037] The structured alarm record is displayed in real time to obtain the structured security monitoring video analysis result.
[0038] As a preferred scheme of the security monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture, the target detection result includes target boundary box coordinates, category labels, category confidence and detection time stamps.
[0039] The behavior recognition result includes behavior category, behavior probability distribution, recognition confidence, action start and end time, and target motion trajectory information.
[0040] As a preferred scheme of the security monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture, the alarm information includes event unique identification, event time information, event location information, event category label, event risk level label and event core feature expression vector.
[0041] In a second aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and wherein the computer program is executed by the processor to implement any step of the security monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture according to the first aspect of the present application.
[0042] In a third aspect, the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement any step of the security monitoring video intelligent analysis method based on the multi-algorithm collaborative and unified architecture according to the first aspect of the present application.
[0043] The application has the beneficial effects that: through multi-scale convolution and long short-term memory inter-frame dependence capture and fusion, the deep joint representation of spatial texture features and time dynamic features in the video frame is realized, the appearance details and motion trajectories of the target can be accurately described, the key behavior information can be reliably captured in the environment with slow motion, partial occlusion or light change, the front-end perception ability of abnormal behavior in complex scenes is improved, the event detection accuracy and anti-interference ability are improved; the multi-source features are integrated according to importance by using the cross-module attention mechanism, the key information of the high-confidence sub-algorithm can be automatically amplified and the noise or redundant results can be suppressed, the false alarm rate is reduced, the waste of computing resources is reduced, and the beneficial effects of enhancing the output explanation are achieved; after the combination of the two, the real-time performance and scalability are improved, and strong technical support is provided for intelligent monitoring in large-scale multi-path control environment. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0045] Fig. 1 The flowchart of the safety monitoring video intelligent analysis method based on the multi-algorithm cooperation and unified architecture.
[0046] Fig. 2 The flowchart of generating a fusion event representation vector.
[0047] Fig. 3 The flowchart of obtaining an event risk level.
[0048] Fig. 4 The flowchart of outputting a structured safety monitoring video analysis result. DETAILED DESCRIPTION
[0049] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the specific embodiments of the application will be described in detail in conjunction with the drawings of the specification.
[0050] In the following description, many specific details are set forth in order to provide a thorough understanding of the application, but the application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the application, therefore the application is not limited to the specific embodiments disclosed below.
[0051] Second, the "one embodiment" or "an embodiment" referred to herein can include a particular feature, structure, or characteristic. The various embodiments appearing at different places in this specification are not necessarily all cumulative or mutually exclusive of each other.
[0052] Referring to Figs. 1-4 For one embodiment of the present application, the embodiment provides a security monitoring video intelligent analysis method based on a multi-algorithm cooperative and unified architecture, including the following steps:
[0053] S1, collecting security monitoring video for preprocessing.
[0054] S1.1, the security monitoring video includes multi-channel camera video source, real-time frame data, timestamp and camera identification.
[0055] Specifically, multi-channel camera video sources are deployed in the security monitoring area, the video data acquisition function inside the multi-channel camera is called, each camera channel is accessed through a unified video data acquisition protocol, and an exemplary acquisition frame rate of 25 frames per second is set. Each image frame is continuously acquired through the acquisition instruction, and the real-time time information corresponding to the image frame is acquired by the camera internal chip or the access acquisition control interface at the same time during the acquisition process and is converted into a unified format timestamp information, such as the format "2025-07-3009: 15: 27.120". The camera identification information is extracted by calling the unique code parameters in the camera hardware inherent identification or network connection address, and each image frame, corresponding timestamp and camera identification are one-to-one corresponding combined to form a complete image frame during the image frame buffering process. The image frame is packed in sequence in the image buffering or transmission protocol, and the security monitoring video containing multi-channel camera video source, real-time frame data, timestamp and camera identification is formed.
[0056] S1.2, preprocessing includes denoising, white balance correction and illumination equalization.
[0057] Specifically, denoising is performed on each real-time frame data of the security monitoring video, based on the non-local mean filtering method, the exemplary search window size is set to 21x21, the exemplary similar window size is set to 7x7, the similarity of the gray mean difference between the current pixel block and the adjacent pixel block is calculated, the pixel block with similarity higher than the exemplary 0.03 is selected to participate in the weighted average, and the pixel block with similarity lower than the exemplary 0.03 is set to 0, and the denoising is completed.
[0058] The white balance correction processing is performed on the denoised image, the pixel mean values of three channels are extracted based on the gray world algorithm, the full image mean values of the red channel, the green channel and the blue channel are obtained respectively, the standard gray value is set as 128, the gain coefficients of each channel are adjusted through the proportional relationship between the channel mean value and the standard gray value, the three channel pixel values are remapped, and the color balanced image frame is obtained;
[0059] The illumination equalization is performed on the color balanced image frame, the gray scale histogram of the luminance channel of the color balanced image frame is extracted based on the histogram equalization method, the number of gray scales is set as 256, the proportion of the cumulative pixel number corresponding to each gray scale value to the total pixel number is obtained by counting the pixel number of each gray scale in the gray scale histogram luminance channel, each gray scale value in the gray scale histogram luminance channel is corresponded to a new gray scale range, the redistribution of the luminance level is realized in the entire gray scale histogram luminance channel, the luminance value is more uniformly distributed in the gray scale histogram, and the illumination equalization is completed.
[0060] S2, spatiotemporal feature extraction and fusion are performed by using a CNN-LSTM spatiotemporal fusion engine, and a fused feature map is output.
[0061] S2.1, based on the preprocessed historical safety monitoring video, a CNN-LSTM spatiotemporal fusion engine is jointly fine-tuned by using an end-to-end supervised learning method, and a trained CNN-LSTM spatiotemporal fusion engine is obtained.
[0062] Specifically, based on the preprocessed historical safety monitoring video, a training sample set containing an input frame sequence and an event label is constructed, each historical safety monitoring video is continuously cut at an exemplary frame rate of 25 frames per second, an image frame sequence with a length of 16 frames is obtained as an input, a behavior category label or an event type label corresponding to the image frame sequence is combined as a supervision signal, a convolutional neural network structure for spatial feature extraction in the CNN-LSTM spatiotemporal fusion engine is called to perform convolution kernel sliding to extract texture, edge and shape features for each image frame, and a spatial feature map sequence with an exemplary dimension of 224x224x64 is generated.
[0063] The long short-term memory neural network structure for time dependence capture in the CNN-LSTM spatio-temporal fusion engine inputs the spatial feature map sequence in frame order, sets the memory cell state update step to an exemplary 1 frame, extracts the inter-frame feature changes and continuous action features, combines the input training sample set labels, and back-propagates the supervision error to the parameters of each layer of the convolutional neural network and the long short-term memory neural network. Based on an exemplary learning rate of 0.001, the parameters are iteratively updated. After each round of training, the training round number is selected as an index based on the validation set loss, and the training is terminated. The joint fine-tuning training of the CNN-LSTM spatio-temporal fusion engine is completed, and the trained CNN-LSTM spatio-temporal fusion engine is obtained.
[0064] S2.2, receiving the pre-processed security monitoring video through the trained CNN-LSTM spatio-temporal fusion engine, performing frame rate cutting, and generating a continuous frame sequence.
[0065] Specifically, the pre-processed security monitoring video is received through the trained CNN-LSTM spatio-temporal fusion engine, the continuous frame image is loaded by calling the image input interface, and the exemplary frame rate is set to 25 frames per second;
[0066] Each segment of the pre-processed security monitoring video is divided into time axes according to the exemplary frame rate, and an image frame sequence with a length of 16 frames is extracted as a continuous frame sequence every 0.64 seconds on the time axis of the security monitoring video;
[0067] The image frame sequence frame number is numbered in time order to ensure that the image frame and the corresponding timestamp are strictly aligned, and the input adaptation interface of the trained CNN-LSTM spatio-temporal fusion engine is called to input each continuous frame sequence to the spatial feature extraction network structure of the trained CNN-LSTM spatio-temporal fusion engine and complete format conversion, and generate a continuous frame sequence.
[0068] S2.3, performing multi-scale convolution and pooling operations on the continuous frame sequence to extract image texture, edge, and target shape information, and forming a spatial feature sequence.
[0069] Specifically, multi-scale convolution and pooling operations are performed on the continuous frame sequence, the spatial feature extraction network structure in the trained CNN-LSTM spatio-temporal fusion engine is called, each image frame is adjusted to an exemplary uniform input size of 224x224 pixels, and three groups of convolution kernel sizes with different receptive fields are set to 3x3, 5x5, and 7x7, respectively.
[0070] The sliding calculation is respectively performed on the image frames to extract local texture, edge and structure information at corresponding scales, the channel independence is maintained in the convolution output, and the scanning is performed in a step of an exemplary 1 pixel; the maximum pooling operation is performed on each group of convolution results, the pooling window size is set to be an exemplary 2*2, the step is an exemplary 2 pixels, the region significant feature response is reserved by the maximum pixel value in the window, the spatial size is reduced and the key region information expression is enhanced, the multi-scale pooling results of each image frame are combined along the channel dimension to form a spatial feature map with an exemplary dimension of 224*224*64, and the above operation is repeated for each image frame in the continuous frame sequence to generate a spatial feature sequence arranged in time sequence.
[0071] S2.4, the inter-frame dependent relationship of the spatial feature sequence is captured through the forgetting gate and the input gate, and a time feature sequence of target motion law and behavior evolution is extracted.
[0072] Specifically, the inter-frame dependent relationship of the spatial feature sequence is captured through the forgetting gate and the input gate, the long short-term memory neural network structure in the trained CNN-LSTM spatio-temporal fusion engine is called, the spatial feature sequence of each image frame is taken as input, the memory cell state update step is set to be an exemplary 1 frame, and the spatial features of each frame input are modeled in the time dimension;
[0073] The current input information weight is controlled through the input gate, the exemplary input gate activation threshold is set to be 0.5, the input spatial feature sequence greater than the input gate activation threshold is weighted and transmitted; the influence of the memory at the last moment is controlled through the forgetting gate, the exemplary forgetting gate threshold is set to be 0.5, the memory information less than the forgetting gate threshold is discarded, the important time dimension information is reserved, the motion mode and behavior evolution between continuous frames are captured; the time feature sequence of target motion law and behavior evolution is extracted to form the time feature sequence;
[0074] It should be noted that the setting process of the input gate activation threshold and the forgetting gate threshold is as follows: an exemplary method is adopted to calculate the frequency distribution curve of the input gate and the forgetting gate corresponding to the activation value in a large number of training sample sets, the threshold value is selected by selecting a demarcation point capable of effectively distinguishing important information and noise, the exemplary 0.5 is taken, that is, the features with the activation value greater than 0.5 are considered as important information to be reserved, and the features lower than the exemplary 0.5 are considered as irrelevant information to be ignored, and the input gate activation threshold and the forgetting gate threshold are determined to be the exemplary 0.5.
[0075] S2.5, the spatial feature sequence and the time feature sequence are spliced and integrated in the channel dimension to output a fusion feature map.
[0076] Specifically, the fusion network structure in the trained CNN-LSTM spatio-temporal fusion engine for integrating multi-source features is called to perform tensor concatenation on the spatial feature sequence with an exemplary dimension of 224x224x64 output by the spatial feature extraction network structure and the temporal feature sequence with an exemplary dimension of 224x224x64 output by the long short-term memory neural network structure in the channel dimension, to generate a spatio-temporal joint feature map sequence with an exemplary dimension of 224x224x128; the convolution kernel size of the fusion network structure is set to be exemplary 3x3, and the step size is exemplary 1 pixel; sliding convolution operation is performed on the spatio-temporal joint feature map sequence after concatenation to extract the coupling information between channels; ReLU is used as the activation operation to enhance the non-linear expression capability; after the convolution output, the maximum pooling operation is performed; the pooling window size is set to be exemplary 2x2, and the step size is exemplary 2 pixels.
[0077] Based on the preprocessed safety monitoring video, a continuous frame sequence in a fixed time window is extracted; in each image frame, the Sobel operator is used to calculate the horizontal and vertical gradient values of each pixel in the image frame to generate an image gray gradient map; the gradient amplitude of each direction in the local region is counted to form a gray gradient direction histogram, which is used as a spatial texture feature.
[0078] Based on the continuous frame images, the dense optical flow method is used to estimate the motion vector of each pixel in adjacent two image frames; the motion vector direction distribution and average amplitude in the local region are extracted as the inter-frame motion feature.
[0079] The continuous frame images are processed; in each image frame, the inter-frame difference method is used to calculate the pixel difference value image between the current frame and the previous frame; by setting a pixel difference value threshold, the significant change region is extracted, and the spatial position and contour information of the target region are retained.
[0080] In each significant change region, the gray gradient direction histogram of the significant change region and the inter-frame motion vector amplitude and direction are calculated.
[0081] The spatial texture feature and the inter-frame motion feature of the significant change region are concatenated to form a local spatio-temporal response vector; all significant change regions are traversed to form a local spatio-temporal feature response sequence in chronological order, and a compressed fusion feature map with an exemplary dimension of 112x112x128 is formed.
[0082] It should be noted that the pixel difference value threshold is set as follows: the distribution of the pixel difference value in the pixel difference value image is obtained; the dividing point that can effectively distinguish the target significant change region from the background noise is selected as the setting basis; the exemplary difference value threshold is exemplary between 0.1 and 0.3.
[0083] It should be noted that the expression for calculating the horizontal and vertical gradient values of each pixel in the image frame using the Sobel operator is as follows:
[0084] Horizontal direction gradient value expression:
[0085] G x = (I x-1,y+1 + 2I x,y+1 + I x+1,y+1 ) - (I x-1,y-1 + 2I x,y-1 + I x+1,y-1 );
[0086] Vertical direction gradient value expression:
[0087] G y = (I x-1,y-1 + 2I x-1,y + I x-1,y+1 ) - (I x+1,y-1 + 2I x+1,y + I x+1,y+1 );
[0088] wherein, G x is a horizontal direction gradient value, G y is a vertical direction gradient value, I x-1,y-1 is a pixel value of the upper left corner of the image frame, I x-1,y is a pixel value of the directly above the image frame, I x-1,y+1 is a pixel value of the upper right corner of the image frame, I x,y-1 is a pixel value of the directly left of the image frame, I x,y+1 is a pixel value of the directly right of the image frame, I x+1,y-1 is a pixel value of the lower left corner of the image frame, I x+1,y is a pixel value of the directly below the image frame, I x+1,y+1 is a pixel value of the lower right corner of the image frame, I is a pixel value of the image frame, x is a horizontal direction coordinate of the image frame, and y is a vertical direction coordinate of the image frame.
[0089] S3, performing target detection, behavior recognition and anomaly detection on the fusion feature map to obtain multi-algorithm analysis results.
[0090] S3.1, identifying each target in the fusion feature map and positioning the boundary box and the category through a CNN-based target detection method, and outputting a target detection result.
[0091] Specifically, based on the compressed fusion feature map with an exemplary dimension of 112x112x128, a CNN-based target detection method is called, the convolution kernel size is set to be exemplary 3x3, the step size is exemplary 1 pixel, sliding convolution calculation is performed on the fusion feature map to extract the feature response of the local region, and candidate box information containing the boundary box coordinates, the category probability and the confidence is output at each sliding window position;
[0092] After generating the candidate box set in all sliding positions, a non-maximum suppression method is used to eliminate the candidate boxes with high overlap degree, the intersection over union of two candidate boxes is calculated and compared with an exemplary overlap degree of 0.5, the low-confidence candidate boxes with an intersection over union greater than the exemplary 0.5 are deleted, and only the candidate box with the highest confidence is retained as the output; for each retained candidate box, the corresponding target bounding box position coordinates and class label are output as the final target detection result;
[0093] It should be noted that the expression for calculating the intersection over union of two candidate boxes is:
[0094]
[0095] where IoU is the intersection over union of the two candidate boxes, w is the intersection area of the predicted box and the real box, h is the height of the overlapping area, S1 is the area of the predicted box, and S2 is the area of the real box.
[0096] S3.2, using the LSTM-based behavior recognition method, the target detection result and the fusion feature map are used to track the target position change and analyze the action sequence, and the behavior recognition result is output.
[0097] Specifically, based on the target detection result, the position coordinates of the target bounding box are sequenced in time order, the bounding box coordinates of the same target in consecutive time frames are associated through spatial overlap matching method, the matching process uses intersection over union calculation, the intersection over union of two bounding boxes is compared with an exemplary 0.3, and the bounding box with an intersection over union greater than the exemplary 0.3 is determined as the continuous position of the same target, and the target position change trajectory sequence is constructed.
[0098] The target position change trajectory sequence and the corresponding fusion feature map are used as input, the LSTM-based behavior recognition method is called, the memory cell state update step is set to an exemplary 1 frame, the target position and feature change are processed through time step recursion, the time sequence dependence of the key moment in the action sequence is captured, the error back propagation is performed combined with the input label information, the LSTM weight parameters are adjusted, and the training of the LSTM-based behavior recognition method is completed.
[0099] The target position change trajectory is input into the LSTM-based behavior recognition method in time sequence, the internal memory unit is used to retain and forget the trajectory information at continuous time, the time dependence and dynamic change characteristics of the action are captured through the gating mechanism, the spatial information in the input fusion feature map is combined to mine the spatio-temporal evolution law of the target behavior, the LSTM model is used to recursively model the time sequence characteristics of the target position change trajectory and the fusion feature map, the behavior evolution pattern is extracted by using the gating mechanism, and the behavior class probability distribution is obtained by using the Softmax function of the output layer. The class whose probability distribution is greater than the exemplary 0.5 is regarded as the effective behavior class, and the behavior recognition result is obtained.
[0100] S3.3, based on the behavior recognition result and the fusion feature map, the error of the behavior feature is calculated and compared with the preset abnormal threshold, and the multi-algorithm analysis result is output.
[0101] Specifically, based on the behavior recognition result and the fusion feature map, a behavior feature vector is extracted, the behavior feature vector includes the action class probability, the action confidence and the spatial and temporal features in the fusion feature map, the mean square error is used as the behavior feature error measurement, and the error value between the current behavior feature vector and the historical normal behavior feature reference vector is obtained.
[0102] The error value is obtained by statistically analyzing a large number of behavior feature error samples, and the frequency intervals are exemplarily divided as [0, 0.1), [0.1, 0.2), [0.2, 0.3), [0.3, 0.4), [0.4, 0.5) and the like, each interval width is exemplarily set as 0.1, the number of error samples in each interval is counted, the frequency proportion of each interval is calculated, and the frequency distribution curve is drawn by combining all the frequency proportions.
[0103] The dividing point which can effectively distinguish the normal and abnormal behaviors is selected as the abnormal threshold, the abnormal threshold is set as the dividing point, the calculated error value and the abnormal threshold are compared, when the calculated error value is greater than the abnormal threshold, the behavior is determined as abnormal, and when the error value is not greater than the abnormal threshold, the behavior is determined as normal, and the multi-algorithm analysis result including the event recognition result and the event feature information is output based on the determination result.
[0104] It should be noted that the historical normal behavior feature reference vector is derived from: after the normal behavior data in a large number of preprocessed historical safety monitoring videos are recognized, the behavior feature vector set is extracted, the representative statistical quantities such as mean value or median value are obtained by statistically analyzing these normal behavior feature vectors, and the stable and representative reference vector is formed.
[0105] S4, a cross-module fusion mechanism based on an attention mechanism is used to weight and integrate the multi-algorithm analysis result to obtain a fusion event representation vector.
[0106] S4.1, splice the intermediate layer deep semantic feature and the logarithmic prediction value of the output layer from the multi-algorithm analysis result to obtain a multi-source perception fusion feature vector.
[0107] Specifically, based on the multi-algorithm analysis result, the deep semantic feature vectors of the intermediate layers are extracted from the CNN target detection method and the LSTM-based behavior recognition method respectively. For the CNN target detection method, the feature maps output by the internal multi-layer convolution layer or the high-dimensional feature vectors of the last fully connected layer are selected. These high-dimensional feature vectors can express the spatial texture, shape and context information of the target. For the LSTM-based behavior recognition method, the time sequence feature representation of the internal LSTM hidden layer is extracted, which reflects the time dynamic evolution characteristics of the target action. For the output layer, the class prediction results of the CNN target detection method and the LSTM-based behavior recognition method are obtained respectively, and the class confidence distribution is logarithmically transformed to generate a logarithmic prediction value, which is used to enhance the numerical stability and discrimination ability of the class probability. The above deep semantic feature vectors and corresponding logarithmic prediction values are respectively subjected to uniform dimension transformation processing to ensure the compatibility of different source features when splicing. Finally, the multi-source perception fusion feature vector with a fusion dimension of 256 is formed by tensor splicing method in the channel dimension.
[0108] The deep semantic feature vectors of the CNN target detection method and the LSTM-based behavior recognition method and the corresponding logarithmic prediction values are respectively subjected to dimension uniform processing. The deep semantic feature vectors and the logarithmic prediction values are spliced in the channel dimension by using the tensor splicing method to generate a multi-source perception fusion feature vector with a fusion dimension of 256. The splicing process maintains the integrity of each feature vector without information loss, which is output as a multi-source perception fusion feature vector.
[0109] S4.2, perform attention weight calculation on the multi-source perception fusion feature vector through the cross-module fusion mechanism based on the attention mechanism to generate a feature dimension attention weight.
[0110] Specifically, the training data of the multi-source perception fusion feature vector and the corresponding event class label are input into the cross-module fusion mechanism based on the attention mechanism in batches with a sample size of 32 as an example;
[0111] For the multi-source perception fusion feature vector in each batch, a query vector is generated by calling the query mapping function in the cross-module fusion mechanism based on the attention mechanism, a key vector is generated by calling the key mapping function, and a value vector is generated by calling the value mapping function.
[0112] The dot product similarity of the query vector and the key vector is calculated, and the similarity is adjusted by using a temperature scaling coefficient 0.07 as an example; the adjusted similarity is subjected to softmax normalization to obtain feature dimension attention weights;
[0113] The attention weights are summed after being multiplied with the value vector by element-by-element multiplication to obtain a weighted feature representation; the weighted feature representation is input into a classification head to output an event category prediction distribution;
[0114] Based on the prediction distribution and the true event category label, a cross-entropy loss function is used to calculate a loss value;
[0115] The loss value is transmitted to the weight parameters of the query mapping function, the key mapping function and the value mapping function through a back propagation algorithm, and the parameters are updated using an Adam optimizer with an example learning rate of 0.0001; the above iteration steps are repeated for example 50 rounds until the validation set accuracy no longer significantly improves; after training is completed, the trained attention mechanism-based cross-module fusion mechanism is saved;
[0116] Based on the trained attention mechanism-based cross-module fusion mechanism, the multi-source perception fusion feature vector is input into the attention mechanism-based cross-module fusion mechanism, the corresponding vectors are generated by using the query, key and value mapping functions, the similarity between the query vector and the key vector is calculated and the attention weights are obtained by softmax normalization, and the value vector is weighted and summed by using the weights to generate feature dimension attention weights;
[0117] It should be noted that the expression for calculating the loss value using the cross-entropy loss function is:
[0118] L = -∑ k T k ×log(P k );
[0119] Where L is the loss value, T k is the value of the kth class in the true target distribution, P k is the prediction probability of the kth class in the probability distribution output by the attention mechanism-based cross-module fusion mechanism, and k is the index variable of the class.
[0120] S4.3, according to the feature dimension attention weights, the multi-source perception fusion feature vector is weighted and summed to obtain a fusion event representation vector.
[0121] Specifically, according to the feature dimension attention weight output by the multi-source perception fusion feature vector and the cross-module fusion mechanism based on the attention mechanism, for each feature channel, the corresponding attention coefficient value in the feature dimension attention weight is read, all channels of the multi-source perception fusion feature vector are traversed, the product operation of all feature values of each channel and the corresponding attention coefficient is performed according to element correspondence, the fusion channel feature weighted by the feature dimension is obtained, and the sum operation is performed on all weighted fusion channel features according to the channel dimension to obtain the fusion event representation vector.
[0122] S5, event type recognition and risk level assessment are performed on the fusion event representation vector to obtain event classification results and event risk levels.
[0123] S5.1, high-level semantic features are extracted from the fusion event representation vector as event feature expressions, event type recognition and classification are performed on the event feature expressions, and event classification results are output.
[0124] Specifically, all feature values are read in dimension order from the fusion event representation vector, and the fusion event representation vector is input into the classification head part of the cross-module fusion mechanism based on the attention mechanism which has been trained, linear transformation is performed on the input fusion event representation vector, and the transformation result is output;
[0125] An activation function is called to perform nonlinear processing on the linear transformation result to generate an intermediate representation, an example dropout operation is used to randomly mask part of the intermediate representation features with a proportion of 0.5 to reduce the risk of overfitting, and the intermediate representation after the dropout operation is input into the linear transformation function again to output the event class prediction distribution, a softmax function is called to perform normalization operation on the prediction distribution to obtain the event classification probability, the maximum value corresponding to the subscript is obtained by traversing the normalized event classification probability vector, and the maximum value corresponding to the subscript is mapped to the event class label one by one, and the event classification result is output.
[0126] S5.2, the event classification result, the event feature expression and the historical risk data are combined to perform risk level assessment, and the event risk level is output.
[0127] Specifically, the class label to which the event belongs is read from the event classification result, the feature values corresponding to all feature dimensions are read from the event feature expression to form an event feature expression vector, and the historical event samples consistent with the event class label are retrieved from the historical risk data according to the event class label;
[0128] obtaining a cosine similarity between the event feature expression vector and each historical event sample feature expression vector, selecting the top 20 historical event samples with the highest similarity as a high correlation sample set, performing statistical operations on the risk level labels corresponding to each sample in the high correlation sample set, counting the frequency of each risk level category in the high correlation sample set, and calculating the frequency distribution of each risk level category, taking the risk level category with the maximum frequency value in the frequency distribution as the event risk level output result of the current event, if there are multiple risk level categories with the same frequency value in the frequency distribution, then selecting the risk level category corresponding to the historical event sample with the highest cosine similarity with the current event feature expression vector as the event risk level;
[0129] It should be noted that the expression for calculating the frequency distribution of each risk level category is:
[0130]
[0131] where f i is the frequency of the i-th risk level category, i is the index variable of the risk level category, n i is the number of samples in the high correlation sample set for the i-th risk level category, and N is the total number of samples in the high correlation sample set.
[0132] S6, making an alarm decision on the event classification result and the event risk level to obtain alarm information.
[0133] S6.1, inducing, comparing and abstracting the classification results of historical events and the risk levels of historical events, extracting the mapping relationship between event features and risk levels, and constructing an alarm decision rule library.
[0134] Specifically, the classification results of historical events and the risk levels of historical events are grouped according to event category labels, all historical event samples under the same event category are sorted into data sets according to event feature expression vectors and corresponding risk level information, a statistical method is used to select and reduce the dimension of event feature expression vectors in each event category, core dimensions of event category features are extracted, different risk level categories are distinguished according to risk level labels, distribution characteristics and boundary ranges of each event risk level category in the core feature space are calculated, based on the feature distribution and the boundary range, an exemplary feature space distance threshold is set to distinguish different event risk level categories, the event category, the core feature dimension range, the risk level and the feature space distance threshold information form a mapping relationship item, all event category mapping relationship items are sorted and summarized as an alarm decision rule library;
[0135] It should be noted that the expression for calculating the distribution characteristics and boundary ranges of each event risk level category in the core feature space is:
[0136]
[0137] wherein, m j is the distribution feature of the jth risk level in the core feature space, j is the risk level category number of the event, M j is the total number of event samples in the jth risk level, is the core feature vector of the nth event belonging to the jth risk level, n is the index of the event sample in the jth risk level;
[0138]
[0139] wherein, d j is the boundary range of the core feature space of the jth risk level.
[0140] S6.2, based on the event classification result and the event risk level, matching the corresponding alarm decision rule in the alarm decision rule library, executing the alarm triggering logic, and outputting the alarm information.
[0141] Specifically, the event classification result and the event risk level are read to obtain the event category label and the event risk level label, the mapping relationship entry of the event category consistent with the event category label is searched from the alarm decision rule library, the feature values corresponding to the core feature dimensions in the event feature expression vector are extracted to form the event core feature expression, the distance between the event core feature expression and the core feature space distribution center corresponding to the different risk level categories in the mapping relationship entry is obtained, and the exemplary feature space distance threshold is taken as the judgment basis to judge which risk level category center has the minimum distance with the event core feature expression and is less than or equal to the feature space distance threshold. If there are multiple risk level categories satisfying the feature space distance threshold, the risk level category corresponding to the minimum value is selected according to the distance from small to large, and the matching result is the corresponding alarm decision rule in the alarm decision rule library. According to the alarm level configuration and the output field definition in the matching result, the alarm information is combined and generated;
[0142] It should be noted that the setting process of the feature space distance threshold is as follows: according to the labeled normal and abnormal events, the distance distribution in the feature space is counted, and the dividing value capable of effectively distinguishing the two types of events is selected as the feature space distance threshold. The exemplary value is 0.15.
[0143] S7, obtaining the structured security monitoring video analysis result.
[0144] S7.1, formatting the alarm information, forming a structured alarm record, and storing it.
[0145] Specifically, the field value in the structured alarm information is filled into the field position information in the structured alarm record format according to the alarm record field order, and the example structured alarm record format includes an event unique identification field, an event time information field, an event location information field, an event category label field, an event risk level label field, an event core feature expression dimension field and a corresponding value field, an alarm level field, and an alarm generation time field. The alarm generation time is recorded in the example timestamp manner and serves as a component of the event unique identification field.
[0146] After the event unique identification field, the event time information field, the event location information field, the event category label field, the event risk level label field, the event core feature expression dimension field and the corresponding value field are spliced into the standard format record structure, the standard format record structure is written into a storage form, and a structured data storage manner is used for disk saving.
[0147] S7.2, real-time display of the structured alarm record is performed to obtain a structured security monitoring video analysis result.
[0148] Specifically, a periodic polling mechanism is configured in a front-end display interface to request a back-end service interface for obtaining the latest structured alarm record once per example time interval of 1 second. In the service interface, a database query statement is called to sort and filter the structured alarm record in the last example one minute in a descending order according to the event time information field, and the query result is encapsulated and returned to the display interface in a JSON format. The display interface parses the returned structured alarm record data to extract the event time information field, the event location information field, the event category label field, the event risk level label field, the event core feature expression dimension field and the corresponding value field, and the alarm level field for structured display rendering. A linkage mechanism is configured in a display area to associate and label the event location information field with an example electronic map component for spatial position display, to associate and label the event time information field with an example video retrieval interface for time index matching and request for a corresponding monitoring video segment, to synchronously load the video content consistent with the alarm event time in the monitoring video display area and highlight the picture position of the event in an example red box manner, and finally to form a structured security monitoring video analysis result including the structured alarm record information and the corresponding monitoring video content in a unified display page.
[0149] S8, the target detection result includes target bounding box coordinates, a category label, a category confidence, and a detection timestamp.
[0150] S8.1, the behavior recognition result includes a behavior category, a behavior probability distribution, a recognition confidence, a motion start time, an end time, and target motion trajectory information.
[0151] S9, the behavior warning information includes event unique identification, event time information, event location information, event category label, event risk level label and event core feature expression vector.
[0152] The embodiment also provides a computer device suitable for the security monitoring video intelligent analysis method based on the multi-algorithm cooperation and unified architecture, which comprises a memory and a processor; the memory is used for storing computer executable instructions, and the processor is used for executing the computer executable instructions to realize the security monitoring video intelligent analysis method based on the multi-algorithm cooperation and unified architecture as proposed in the above embodiment.
[0153] The computer device can be a terminal, which comprises a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. The processor of the computer device is used for providing computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used for wired or wireless communication with external terminals. The wireless communication can be realized through WIFI, an operator network, NFC (Near Field Communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.
[0154] The embodiment also provides a storage medium having a computer program stored thereon, the program being executed by a processor to realize the security monitoring video intelligent analysis method based on the multi-algorithm cooperation and unified architecture as proposed in the above embodiment. The storage medium can be realized by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk or an optical disk.
[0155] To sum up, the application realizes the deep joint representation of spatial texture features and time dynamic features in the video frame by multi-scale convolution and long short-term inter-frame dependence capture and fusion, can accurately depict the appearance details and motion trajectory of the target, ensures that the key behavior information can still be reliably captured in the environment with slow motion, partial occlusion or light change, improves the front-end perception ability of abnormal behavior in complex scenes, and achieves the ability to improve the event detection accuracy and anti-interference; the multi-source features are integrated according to the importance by using the cross-module attention mechanism, which can automatically amplify the key information of the high-confidence sub-algorithm and suppress noise or redundant results, reduce the false alarm rate, reduce the waste of computing resources and enhance the beneficial effects of output explanation; after the combination of the two, the real-time performance and scalability are improved, which provides strong technical support for intelligent monitoring in large-scale multi-path control environment.
[0156] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.
Claims
1. A method for intelligent analysis of security surveillance videos based on multi-algorithm collaboration and a unified architecture, characterized by: include, Security monitoring videos are collected, preprocessed, and then spatiotemporal features are extracted and fused using the CNN-LSTM spatiotemporal fusion engine to output a fused feature map. The fused feature map is subjected to object detection, behavior recognition, and anomaly detection to obtain multi-algorithm analysis results; By utilizing a cross-module fusion mechanism based on attention, the analysis results of multiple algorithms are weighted and integrated to obtain a fused event representation vector; The event type identification and risk level assessment are performed on the fused event representation vector to obtain the event classification results and event risk level. The system makes alarm decisions based on the event classification results and event risk levels, obtains alarm information, stores and visualizes it, and obtains structured security monitoring video analysis results.
2. The intelligent analysis method for security monitoring videos based on multi-algorithm collaboration and unified architecture as described in claim 1, characterized in that: The security monitoring video includes multiple camera video sources, real-time frame data, timestamps, and camera identifiers. The preprocessing includes noise reduction, white balance correction, and illumination equalization.
3. The intelligent analysis method for security monitoring videos based on multi-algorithm collaboration and unified architecture as described in claim 2, characterized in that: The specific steps for generating the fused feature map are as follows: Based on preprocessed historical security surveillance videos, the CNN-LSTM spatiotemporal fusion engine is jointly fine-tuned and trained using an end-to-end supervised learning method to obtain the trained CNN-LSTM spatiotemporal fusion engine. The pre-processed security monitoring video is received by the trained CNN-LSTM spatiotemporal fusion engine, and frame rate segmentation is performed to generate a continuous frame sequence. Multi-scale convolution and pooling operations are performed on consecutive frame sequences to extract image texture, edge and target shape information, forming a spatial feature sequence; By capturing the inter-frame dependencies of spatial feature sequences through forget gates and input gates, the temporal feature sequences of target motion patterns and behavioral evolution are extracted. Spatial feature sequences and temporal feature sequences are concatenated and integrated along the channel dimension to output a fused feature map.
4. The intelligent analysis method for security monitoring videos based on multi-algorithm collaboration and unified architecture as described in claim 3, characterized in that: The process involves performing object detection, behavior recognition, and anomaly detection on the fused feature map to obtain multi-algorithm analysis results. The specific steps are as follows: The object detection method based on CNN identifies each object in the fused feature map, locates the bounding box and category, and outputs the object detection results. Using an LSTM-based behavior recognition method, the target detection results and fused feature map are used to track target position changes and parse action sequences, outputting behavior recognition results. Based on the behavior recognition results and the fused feature map, the error of the behavior features is calculated and compared with the preset anomaly threshold, and the multi-algorithm analysis results are output.
5. The intelligent analysis method for security monitoring videos based on multi-algorithm collaboration and unified architecture as described in claim 4, characterized in that: The method utilizes a cross-module fusion mechanism based on attention to perform weighted integration of the analysis results from multiple algorithms, obtaining a fused event representation vector. The specific steps are as follows: The deep semantic features of the intermediate layer are extracted from the multi-algorithm analysis results and concatenated with the logarithmic prediction value of the output layer to obtain the multi-source perception fusion feature vector. By using a cross-module fusion mechanism based on attention, attention weights are calculated on the multi-source perception fusion feature vectors to generate feature dimension attention weights. Based on the attention weights of the feature dimensions, the multi-source perception fusion feature vectors are weighted and summed to obtain the fusion event representation vector.
6. The intelligent analysis method for security monitoring videos based on multi-algorithm collaboration and unified architecture as described in claim 5, characterized in that: The process of performing event type identification and risk level assessment on the fused event representation vector yields event classification results and event risk levels. The specific steps are as follows: High-level semantic features are extracted from the fused event representation vector as event feature expressions, and event type identification and classification are performed on the event feature expressions to output the event classification results; By combining event classification results, event characteristic expressions, and historical risk data, a risk level assessment is performed, and the event risk level is output.
7. The intelligent analysis method for security monitoring videos based on multi-algorithm collaboration and unified architecture as described in claim 6, characterized in that: The specific steps for making alarm decisions based on event classification results and event risk levels to obtain alarm information are as follows. The classification results and risk levels of historical events are summarized, compared and abstracted to extract the mapping relationship between event characteristics and risk levels, and to construct an alarm decision rule base. Based on the event classification results and event risk level, the corresponding alarm decision rule is matched in the alarm decision rule base, the alarm triggering logic is executed, and the alarm information is output.
8. The intelligent analysis method for security monitoring videos based on multi-algorithm collaboration and unified architecture as described in claim 1, characterized in that: The specific steps for obtaining the structured security surveillance video analysis results are as follows: The alarm information is formatted to form a structured alarm record and then stored. The structured alarm records are displayed in real time, and the results of structured security monitoring video analysis are obtained.
9. The intelligent analysis method for security monitoring videos based on multi-algorithm collaboration and unified architecture as described in claim 4, characterized in that: The target detection results include target bounding box coordinates, category label, category confidence score, and detection timestamp; The behavior recognition results include behavior category, behavior probability distribution, recognition confidence, action start time, end time, and target motion trajectory information.
10. The intelligent analysis method for security monitoring videos based on multi-algorithm collaboration and unified architecture as described in claim 7, characterized in that: The alarm information includes a unique event identifier, event time information, event location information, event category label, event risk level label, and event core feature expression vector.
Citation Information
Patent Citations
Multi-method fusion violence and terrorism event detection alarm method and system and related equipment
CN116797968A
Multi-algorithm intelligent analysis method and device, storage medium and computer equipment
CN116994183A
ResNet and LSTM (Long Short Term Memory)-based multi-dimensional spatio-temporal feature fusion deep counterfeit video detection method and device
CN119339220A
A method to improve event detection accuracy based on multi-model fusion
CN119785162A
Method and system for detecting abnormal traffic behavior
US20250078516A1
Cited By
Remote security alarm system combining video monitoring and intrusion detection identification technology
CN121259956A
Remote security alarm system combining video monitoring and intrusion detection identification technology
CN121259956B
Video content security risk monitoring system based on big data
CN121686335A
Big data-based video content security risk monitoring system
CN121686335B
Network security event association detection method based on big data analysis
CN121690815A