An intelligent video surveillance method and system based on multi-source data fusion

By adopting spatiotemporal feature codec network and feature recognition network in the intelligent video surveillance system, combining the measurement learning strategy of hyperspherical manifold and KL divergence, the problem of difficulty in fusion of multi-source video data in single-modal data analysis is solved, and higher robustness and accuracy are achieved in abnormal event detection and pedestrian re-identification.

CN119169536BActive Publication Date: 2025-05-30QINGDAO BAOSEN NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411657490.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-20
Publication Date
2025-05-30
Estimated Expiration
2044-11-20

AI Technical Summary

Technical Problem

The existing intelligent video surveillance methods mainly focus on the modeling and analysis of single-modal data, and it is difficult to effectively integrate multi-source video data, resulting in the reduction of the accuracy of abnormal event detection and the effect of pedestrian re-identification under conditions such as lighting changes, scene complexity and visual occlusion.

Method used

An intelligent video surveillance method based on multi-source data fusion is proposed. Through spatiotemporal feature encoding and decoding networks and feature recognition networks, combined with the measurement learning strategy of hyperspherical manifold and KL divergence, it realizes feature extraction and abnormal behavior detection of multi-source video data.

Benefits of technology

It significantly improves the robustness and adaptability of the monitoring system, improves the accuracy of abnormal event detection and the effect of pedestrian re-identification, and can promptly identify sudden abnormal events in monitoring scenarios under different scenarios and lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119169536B_ABST
    Figure CN119169536B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent video surveillance method and system based on multi-source data fusion, belonging to the technical field of video data analysis. First, multi-source video data is acquired and preprocessed; the preprocessed video data is input into a spatio-temporal feature encoding and decoding network to obtain a preliminary result of pedestrian image abnormal behavior detection; based on the obtained preliminary detection result, pedestrian images in the abnormal behavior are extracted to obtain abnormal behavior image data; the obtained abnormal behavior image data is input into a feature recognition network to obtain feature vectors; finally, the obtained feature vectors are input into a cross-modal pedestrian feature registration network based on dual-aligned feature embedding to obtain a final result of pedestrian image abnormal behavior detection. An abnormal event detection method based on multi-source data fusion and high-efficiency and stability is provided, which can fuse multi-source surveillance data and adapt to changes in different scenarios and lighting conditions, and timely identify sudden abnormal behaviors in the surveillance scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video data analysis, and particularly relates to an intelligent video monitoring method based on multi-source data fusion. Background Art

[0002] As an important information auxiliary means in the field of daily public safety, the urban-level monitoring camera network is constantly expanding to comprehensively cover every corner of the city and promptly detect and respond to potential security threats. With the rapid increase in the number of monitoring cameras, on the one hand, the spatial distribution of the monitoring camera network is extensive, and there may be diverse lighting conditions and background environments in different scenarios. On the other hand, the images collected by numerous types of cameras at different working time periods exhibit different modal characteristics. The existing technologies have the following limitations:

[0003] Single-modal data analysis: Intelligent video monitoring needs to process multi-source video data, including RGB, grayscale, near-infrared, and thermal infrared, etc. However, the current intelligent video monitoring methods mainly focus on the modeling and analysis of single-modal data, usually only using RGB or grayscale images. This limitation performs poorly in cases such as lighting changes, scene complexity, and visual occlusion, resulting in a decrease in the accuracy of abnormal event detection and the effect of pedestrian re-identification. The information provided by different modal data (such as near-infrared, thermal infrared, etc.) is complementary, and fusing these multi-source data can significantly improve the robustness and adaptability of the monitoring system;

[0004] Mixing of spatial and temporal information: Multi-source video data contains both spatial and temporal information, while the existing research methods have not fully considered the modeling of the temporal-spatial mixed information for the normal video data distribution. This difference may lead to a decrease in the sensitivity of the existing abnormal event detection models when dealing with abnormal events that are not very different from normal events;

[0005] Cross-modal pedestrian re-identification: For the cross-modal pedestrian re-identification problem under multi-source data, the existing work mainly uses the non-linear fitting ability of deep neural networks to transform the features of different modalities into the same feature subspace to extract discriminative pedestrian feature representations. However, this method fails to effectively unify the feature extraction and metric learning strategies of multi-modal video data. In addition, when processing the features of cross-modal pedestrian images with spatial and modal misalignment characteristics, it is often difficult to achieve effective registration, resulting in insufficient robustness in multi-source data scenarios.

[0006] Therefore, there is an urgent need for an abnormal event detection algorithm based on multi-source data fusion and high efficiency and stability, which can fuse multi-source monitoring data and adapt to changes in different scenarios and lighting conditions, promptly identify sudden abnormal events in the monitoring scene and issue an alarm. Summary of the Invention

[0007] In view of the above problems, a first aspect of the present invention provides an intelligent video surveillance method based on multi-source data fusion, including the following steps:

[0008] S1. Obtain multi-source video data and preprocess the data;

[0009] S2. Input the preprocessed video data into a spatio-temporal feature encoding and decoding network to obtain a preliminary result of abnormal behavior detection for pedestrian images;

[0010] S3. Based on the preliminary result of abnormal behavior detection for pedestrian images obtained in S2, extract the pedestrian images in the abnormal behavior to obtain abnormal behavior image data;

[0011] S4. Input the abnormal behavior image data obtained in S3 into a feature recognition network to obtain feature vectors; the feature recognition network includes a feature learning module and a feature mapping module. When a pedestrian image is input into the feature learning module, local features of the image are obtained , local features are input into the feature mapping module to obtain hyper-spherical domain mapping features ; the feature recognition network is trained based on a metric learning strategy of hyper-spherical manifold and KL divergence;

[0012] S5. Input the feature vectors obtained in S4 into a cross-modal pedestrian feature registration network based on double-aligned feature embedding. First, use a backbone network and a global average pooling layer to extract fine-grained features, then use a convolutional layer and a batch normalization layer to reduce the dimension of the fine-grained features, and finally input the dimension-reduced fine-grained features into a fully-connected layer to obtain the final result of abnormal behavior detection for pedestrian images.

[0013] Preferably, the multi-source video data in S1 includes visible light pedestrian video data and infrared light pedestrian video data. The multi-source video data is standardized, denoised, and the spatial resolution and brightness are uniformly adjusted to obtain the preprocessed multi-source video data.

[0014] Preferably, the spatio-temporal feature encoding and decoding network in S2 includes a 3D encoding module, a 2D decoding module, a resampling module, a temporal downsampling module, and a spatio-temporal consistency discriminator. The specific steps of inputting visible light video data and infrared video data into the spatio-temporal feature encoding and decoding network to obtain a preliminary result of abnormal behavior detection are as follows:

[0015] S21. Uniformly sample the pedestrian video and to k frames, and input the video sequence composed of k video frames into the spatio-temporal feature encoding and decoding network, where represents the Frame image;

[0016] S22. The hidden layer feature vector obtained after sequentially connecting the input to 5 3D encoding modules is calculated as follows:

[0017] ;

[0018] ;

[0019] ;

[0020] ;

[0021] ;

[0022] where , , , and represent the first 3D encoding module, the second 3D encoding module, the third 3D encoding module, the fourth 3D encoding module, and the fifth 3D encoding module respectively; , , and are the output features of the first 3D encoding module, the second 3D encoding module, the third 3D encoding module, and the fourth 3D encoding module respectively;

[0023] S23. Input to the resampling module to resample in the hidden layer space where is located to obtain a new feature , and the specific calculation formula is as follows:

[0024] ;

[0025] where is a random vector between zero and one, and are the mean and variance of the standard normal distribution that follows in the high-dimensional space respectively; and are learned from the data distribution of two fully connected layers ;

[0026] Use the Kullback Leibler divergence as the constraint condition for resampling during the sampling process to ensure that the resampled feature and the original feature follow the same distribution, and the calculation formula of the Kullback Leibler divergence is as follows:

[0027] ;

[0028] Among them, represents the Kullback Leibler divergence, and

[0029] S24, after the input passes through five sequentially connected 2D decoding modules in series, it is decoded into the next frame image of the input sequence , and the specific calculation formula is as follows:

[0030] ;

[0031] ;

[0032] ;

[0033] ;

[0034] ;

[0035] Among them, , , , and respectively represent the first 2D decoding module, the second 2D decoding module, the third 2D decoding module, the fourth 2D decoding module, and the fifth 2D decoding module; , , and are respectively the output features of the first 2D decoding module, the second 2D decoding module, the third 2D decoding module, and the fourth 2D decoding module; is the next frame image of the generated input sequence; represents the temporal downsampling module, represents the feature fusion function, which directly merges two tensors with the same height and width in the channel dimension; the convolution operations in the 2D decoding module are all two-dimensional convolutions;

[0036] S25, the and are simultaneously input into the spatio-temporal consistency discriminator to obtain a preliminary abnormal behavior detection result .

[0037] Preferably, the spatio-temporal feature encoding and decoding network is optimized and trained through a spatio-temporal consistency training strategy, and the cost function of the spatio-temporal consistency training strategy consists of a pixel-level mean squared error function and an adversarial loss function;

[0038] The pixel-level mean squared error function , trains the entire spatio-temporal feature encoding and decoding network by minimizing the following cost function:

[0039] ;

[0040] where, are the height and width of the images in the input video sequence, and this function represents calculating the squared error between the generated result and its true value pixel by pixel, is the modulo operation symbol;

[0041] The adversarial loss function consists of a generator loss function and a discriminator loss function;

[0042] The generator loss function is:

[0043] ;

[0044] The discriminator loss function is:

[0045] ;

[0046] where, represents a matrix of all zeros, represents a pedestrian image sample generated by the decoder G in the spatio-temporal feature encoding and decoding network , denotes the concatenation of the time channels, denotes an identity matrix with the same dimension as ; is the actual video frame sequence, and the corresponding true video frame is concatenated with input to the decoder in the spatio-temporal feature encoding and decoding network on the time channel to obtain , is the sampled value of the true video sequence, ;

[0047] The training process of the spatio-temporal consistency enhancement network uses an alternating training method between the decoder in the spatio-temporal feature encoding and decoding network and the spatio-temporal consistency discriminator until both converge.

[0048] Preferably, the feature recognition network in S4 includes a feature learning module and a feature mapping module. The feature learning module consists of a convolutional neural network, a fully connected layer, and a batch normalization layer; the feature mapping module consists of a dropout layer, a fully connected layer, and a second norm normalization layer. The specific steps for the feature recognition network to extract the features of the pedestrian abnormal behavior image are as follows:

[0049] Input the pedestrian image obtained in S3 into the feature learning module to obtain the local features of the image , and the specific calculation formula is as follows:

[0050] ;

[0051] where is the convolutional layer, is the fully connected layer, is the batch normalization layer;

[0052] Input the local features into the feature mapping module to obtain the hypersphere domain mapping features , and the specific calculation formula is as follows:

[0053] ;

[0054] where is the dropout layer, is the second norm normalization layer; Through the mapping of the feature mapping module, the domain-specific features of different modalities are mapped onto the same hypersphere manifold, obtaining the domain-shared features on the hypersphere manifold.

[0055] Preferably, the feature recognition network in S4 is trained based on the metric learning strategy of hypersphere manifold and KL divergence, specifically:

[0056] First, use the hypersphere manifold loss function to train the feature recognition network to accelerate the convergence process of the model. The hypersphere manifold loss function is as follows:

[0057] ;

[0058] where s is the proportional coefficient of the cosine value cos of the included angle, which is used to accelerate the training process of the model, C represents the number of pedestrian abnormal behavior labels, is the number of pedestrian images included in a batch; represents the vector included angle between different behavior categories in different pedestrian images; represents the included angle between the abnormal behavior prediction output vector and the true label vector of the th pedestrian image;

[0059] Then, construct a KL divergence constraint training feature recognition network; that is, use the probability constraint method based on Kullback-Leibler divergence to measure the probability prediction values of different modalities of the same person. and to improve the modality independence of cross-modal features, to the KL divergence value is calculated by the following formula:

[0060] ;

[0061] Since the KL divergence is asymmetric, from the divergence is:

[0062] ;

[0063] where, is the pedestrian behavior prediction result in the visible light modality, is the pedestrian behavior prediction result in the infrared light modality.

[0064] Preferably, the specific processing process of the cross-modal pedestrian feature registration network based on double-aligned feature embedding in S5 is as follows:

[0065] S51, input into the ResNet50 backbone network; the network before the last pooling layer of the traditional ResNet50 is used as the ResNet50 backbone network. Input into the ResNet50 backbone network, and the output is a multi-channel feature map M. The specific calculation formula is as follows:

[0066] ;

[0067] where M is a three-dimensional tensor;

[0068] S52, randomly divide the feature map into K non-overlapping parts along the vertical direction, and then perform global average pooling calculation on each part to obtain the fine-grained feature For the th part the calculation formula is as follows:

[0069] ;

[0070] where is a 2048-dimensional feature vector;

[0071] S53, the fine-grained feature The input convolutional layer and batch normalization layer reduce the dimension to a feature vector with a dimension of 256*K;

[0072] S54, input the K feature vectors of the fine-grained features after dimensionality reduction into K fully connected layers with different weights respectively to obtain the final detection result of abnormal behaviors of pedestrian images.

[0073] Preferably, the cross-modal pedestrian feature registration network based on dual-aligned feature embedding is trained through a modality consistency metric learning strategy, including an intra-class distribution constraint loss function and an inter-class correlation constraint loss function, specifically as follows:

[0074] The intra-class distribution constraint loss function is minimized by minimizing the following formula dispersion to obtain the intra-class distribution constraint loss function :

[0075] ;

[0076] where and are the means of the visible-light modality pedestrian features and the infrared modality pedestrian features respectively, and are the standard deviations of the visible-light modality pedestrian features and the infrared modality pedestrian features, , , and The corresponding calculation formulas are shown as follows:

[0077] ;

[0078] ;

[0079] ;

[0080] ;

[0081] where, represents the feature of the th infrared modality pedestrian image, represents the feature of the th visible-light modality pedestrian image;

[0082] The specific calculation formula of the inter-class correlation constraint loss function is as follows:

[0083] ;

[0084] where The Pearson correlation coefficient matrix for a batch of infrared images The th element in The Pearson correlation coefficient matrix for a batch of visible light images The th element in And The calculation formulas of are as follows:

[0085] ;

[0086] .

[0087] The second aspect of the present invention provides an intelligent video surveillance system based on multi-source data fusion, which is applied to the intelligent video surveillance method as described in the first aspect, and includes: a video acquisition module and a backend server;

[0088] The video acquisition module is used to acquire multi-source video data;

[0089] The backend server internally deploys a video preprocessing module, a spatio-temporal feature encoding and decoding network, an abnormal behavior extraction module, a feature recognition network, and a cross-modal pedestrian feature registration network based on dual-aligned feature embedding;

[0090] The video preprocessing module is used to standardize and denoise the multi-source video, and uniformly adjust the spatial resolution and brightness;

[0091] The spatio-temporal feature encoding and decoding module is used to input the preprocessed video data to obtain a preliminary result of pedestrian image abnormal behavior detection;

[0092] The abnormal behavior extraction module is used to extract the pedestrian images in the abnormal behavior to obtain abnormal behavior image data;

[0093] The feature recognition network is used to obtain feature vectors;

[0094] The cross-modal pedestrian feature registration network based on dual-aligned feature embedding is used to obtain the final detection result of pedestrian image abnormal behavior.

[0095] Compared with the prior art, the present invention has the following beneficial effects:

[0096] 1. Regarding the problem of the mixture of airspace and time-domain information in multi-source data, the present invention proposes an abnormal event detection method based on spatio-temporal consistency enhancement, which improves the spatio-temporal consistency of future frame prediction of video images on the training data set through a spatio-temporal feature encoding and decoding network and a spatio-temporal consistency discrimination network. The difference between abnormal data and normal data in the prediction results is increased, so as to more significantly detect abnormal events in the video sequence. Specifically, the present invention proposes a framework for predicting future frames of an input frame sequence based on a deep convolutional network, and extracts the mixed spatio-temporal domain information of the input video sequence through three-dimensional convolution operations. At the same time, the present invention proposes a spatio-temporal consistency discrimination network to enhance the spatio-temporal consistency between the generated video future frame images and the input video sequence. Combining the above two methods can effectively increase the perturbation of abnormal events to the model output, thereby differentiating the model's response to abnormal and normal data;

[0097] 2. Aiming at the problem that it is difficult to effectively unify the feature extraction and metric learning strategies of multi-modal video data in cross-modal pedestrian re-identification, the present invention proposes a cross-modal pedestrian re-identification method based on hypersphere manifold embedding, which unifies the dimensions of the objective functions in the two tasks of recognition and classification by using hypersphere manifold embedding representation and bidirectional ranking loss, so as to ensure the high coupling of the feature extraction process and the metric constraint while extracting highly discriminative features. In addition, the present invention further analyzes the hypersphere manifold embedding and proposes a decorrelation method based on singular value decomposition to further expand the inter-class distance of features on the hypersphere manifold;

[0098] 3. Aiming at the problem that it is difficult to effectively register the cross-modal pedestrian image features in multi-source video data pedestrian re-identification, the present invention proposes a cross-modal pedestrian feature registration method based on dual-aligned feature embedding, which solves the problems of spatial misalignment and modal misalignment existing in cross-modal pedestrian images by using fine-grained feature extraction and modal consistency constraints, and simultaneously solves the two misalignment problems of cross-modal pedestrian re-identification in a unified framework, enabling the model to have better robustness to pedestrian postures while extracting highly discriminative feature representations with modal invariance. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Figure 1 It is the overall flowchart of the intelligent video monitoring method based on multi-source data fusion of the present invention.

[0100] Figure 2 Schematic diagram of the spatio-temporal feature encoding and decoding network structure of the present invention.

[0101] Figure 3 Schematic diagram of the feature recognition network structure of the present invention.

[0102] Figure 4 Schematic diagram of the cross-modal pedestrian feature registration network structure based on dual-aligned feature embedding of the present invention.

[0103] Figure 5 This is the convergence curve graph of the training loss function during the experiment of the present invention. Specific implementation manner

[0104] The present invention will be further described below in conjunction with specific embodiments.

[0105] The overall logic of the solution of the present invention is as Figure 1 shown. First, multi-source video data is obtained and preprocessed; the preprocessed video data is input into the spatio-temporal feature encoding and decoding network to obtain the preliminary result of pedestrian image abnormal behavior detection; based on the obtained preliminary result of pedestrian image abnormal behavior detection, the pedestrian images in the abnormal behavior are extracted to obtain abnormal behavior image data; the obtained abnormal behavior image data is input into the feature recognition network to obtain feature vectors; the feature recognition network includes a feature learning module and a feature mapping module, the pedestrian image is input into the feature learning module to obtain the local features of the image, and the local features are input into the feature mapping module to obtain the hyperspherical domain mapping features; finally, the obtained feature vectors are input into the cross-modal pedestrian feature registration network based on double-aligned feature embedding to obtain the final result of pedestrian image abnormal behavior detection.

[0106] 1. First, collect visible light pedestrian video data and infrared light pedestrian video data to form multi-source video data; secondly, standardize, denoise the multi-source video data, and uniformly adjust the spatial resolution and brightness to obtain the preprocessed multi-source video data, where the visible light video data is represented by , and the infrared video data is represented by .

[0107] 2. Construct a spatio-temporal feature encoding and decoding network, and input the obtained visible light video data and infrared video data into the spatio-temporal feature encoding and decoding network to obtain the preliminary result of pedestrian image abnormal behavior detection; secondly, propose a spatio-temporal consistency training strategy to train the spatio-temporal feature encoding and decoding network.

[0108] 1. Regarding the spatio-temporal feature encoding and decoding network:

[0109] The structure of the spatio-temporal feature encoding and decoding network is as Figure 2 shown. This network is composed of a 3D encoding module, a 2D decoding module, a resampling module, a temporal downsampling module, and a spatio-temporal consistency discriminator. The specific steps for inputting the visible light video data and infrared video data into the spatio-temporal feature encoding and decoding network to obtain the preliminary result of abnormal behavior detection are as follows:

[0110] For the pedestrian video and Uniformly sample to k frames, and form a video sequence consisting of k video frames Input to the spatiotemporal feature encoding and decoding network, where Indicates the first Frame image;

[0111] Will After inputting 5 3D encoding modules connected in series, we get The hidden feature vector of , the specific calculation formula is as follows:

[0112] ;

[0113] ;

[0114] ;

[0115] ;

[0116] ;

[0117] in , , , and Respectively represent a first 3D encoding module, a second 3D encoding module, a third 3D encoding module, a fourth 3D encoding module and a fifth 3D encoding module; , , and are the output features of the first 3D encoding module, the second 3D encoding module, the third 3D encoding module and the fourth 3D encoding module respectively; the input of the 3D encoding module of the present invention is a five-dimensional tensor, so the spatiotemporal mixed information contained in the pedestrian video can be well captured, that is, the feature Represents the spatiotemporal fusion information contained in pedestrian videos;

[0118] Will Input resampling module, Resample the hidden space to obtain new features , the specific calculation formula is as follows:

[0119] ;

[0120] in, is a random vector between zero and one, and They are The mean and variance of the standard normal distribution in high-dimensional space; and Obtained from the data distribution learned by two fully connected layers ;

[0121] During the sampling process, the Kullback Leibler divergence is used as the constraint condition for resampling to ensure that the features after resampling Follow the same distribution as the original features , and the calculation formula of the Kullback Leibler divergence is as follows:

[0122] ;

[0123] Among them, Represents the Kullback Leibler divergence, Is the value of the Kullback Leibler divergence; since the training process only uses normal video data, this part of the resampling module will only correctly estimate the parameters of the distribution followed by the features containing the spatio-temporal domain information of normal events; through this resampling process, the difficulty of the future frame prediction task is increased. When the learning ability of the model is the same, adding an additional resampling module can effectively reduce the possibility that the model generalizes the future frame prediction ability from normal video data to abnormal video data, thereby enhancing the perturbation of abnormal video data to the model;

[0124] After passing Through five consecutive 2D decoding modules in series, it is decoded into the next frame image of the input sequence , and the specific calculation formula is as follows:

[0125] ;

[0126] ;

[0127] ;

[0128] ;

[0129] ;

[0130] Among them, , , , And Represent the first 2D decoding module, the second 2D decoding module, the third 2D decoding module, the fourth 2D decoding module and the fifth 2D decoding module respectively; , , And The output features of the first 2D decoding module, the second 2D decoding module, the third 2D decoding module, and the fourth 2D decoding module, respectively; The next frame image of the generated input sequence; Represents the time-domain downsampling module, Represents the feature fusion function, which directly merges two tensors with the same height and width in the channel dimension; the convolution operations in the 2D decoding module constructed in the present invention are all two-dimensional convolutions, and the model parameters are shown in Table 1;

[0131] Take and Input into the spatio-temporal consistency discriminator at the same time to obtain the preliminary abnormal behavior detection result .

[0132] Table 1 2D Decoding Module Parameter Table

[0133]

[0134] 2. Regarding the spatio-temporal consistency training strategy:

[0135] The present invention proposes a spatio-temporal consistency training strategy to train the spatio-temporal feature encoding and decoding network, ensuring that the 3D encoding module and the 2D decoding module in the spatio-temporal feature encoding and decoding network can fully learn the spatio-temporal consistent features between each frame image in the pedestrian video. The cost function of the spatio-temporal consistency training strategy consists of a pixel-level mean squared error function and an adversarial loss function, as follows:

[0136] Pixel-level mean squared error function , and train the entire spatio-temporal feature encoding and decoding network by minimizing the following cost function:

[0137] ;

[0138] Among them, is the height and width of the images in the input video sequence. This function represents calculating the mean squared error between the generated result and its true value pixel by pixel, is the modulo calculation symbol;

[0139] Adversarial loss function, consisting of a generator loss function and a discriminator loss function;

[0140] Generator loss function is:

[0141] ;

[0142] Discriminator loss function is:

[0143] ;

[0144] Among them, represents a zero matrix, represents the pedestrian image sample generated by the decoder G in the spatio-temporal feature encoding and decoding network , represents the concatenation of time channels, represents that the dimension is the same as that of the identity matrix; is the actual video frame sequence, and the corresponding real video frame is concatenated with the input to the decoder in the spatio-temporal feature encoding and decoding network on the time channel, and then can be obtained. is the sampling value of the real video sequence, ;

[0145] For the spatio-temporal consistency discriminator, when the input of the spatio-temporal consistency discriminator is , the spatio-temporal consistency discriminator needs to output a feature map with all values being 1, which means that the discriminator determines the input video sequence as a spatio-temporally consistent sequence. When the input is or , the expected output of the discriminator is a feature map with all values being 0, which means that the discriminator determines the input video sequences and as spatio-temporally inconsistent. Therefore, the loss function for training the discriminator can be written as:

[0146] ;

[0147] The training process of the spatio-temporal consistency enhancement network uses the method of alternately training the generator and the discriminator until both converge:

[0148] A: Fix the parameters of the spatio-temporal consistency discriminant network and train the spatio-temporal feature encoding and decoding network. The objective function during training is:

[0149] ;

[0150] Among them, , is the trade-off factor for balancing the multi-task objective function; is the divergence of;

[0151] B: Fix the parameters of the spatio-temporal feature encoding and decoding network and train the spatio-temporal consistency discriminant network. The objective function during training is .

[0152] III. Based on the preliminary results of pedestrian image abnormal behavior detection obtained , the portrait pictures of pedestrians with abnormal behaviors obtained from and are extracted to obtain abnormal behavior image data . First, and are split into frame-by-frame pictures to obtain a pedestrian image dataset; then, according to R, it is determined which images in the pedestrian image dataset contain pedestrians with abnormal behaviors; finally, the images containing pedestrians with abnormal behaviors in the pedestrian image dataset are screened out.

[0153] IV. Construct a feature recognition network. The feature recognition network is as shown in Figure 3 . The obtained abnormal behavior images are input into the feature recognition network to obtain feature vectors; secondly, a metric learning strategy based on hypersphere manifold and KL divergence is proposed to train the feature recognition network.

[0154] 1. Regarding the feature recognition network:

[0155] Construct a feature recognition network to achieve feature extraction of pedestrian abnormal behavior images; the feature recognition network includes a feature learning module and a feature mapping module, where the feature learning module consists of a convolutional neural network, a fully connected layer, and a batch normalization layer; the feature mapping module consists of a dropout layer, a fully connected layer, and a second norm normalization layer; the specific steps for the feature recognition network to achieve feature extraction of pedestrian abnormal behavior images are as follows:

[0156] The obtained pedestrian images are input into the feature learning module to obtain the local features of the images , and the specific calculation formula is as follows:

[0157] ;

[0158] where is the convolutional layer, is the fully connected layer, is the batch normalization layer;

[0159] The local features are input into the feature mapping module to obtain the hypersphere domain mapping features , and the specific calculation formula is as follows:

[0160] ;

[0161] where is the dropout layer, is the second norm normalization layer; through the mapping of the feature mapping module, the domain detailed features of different modalities are mapped onto the same hypersphere manifold to obtain the domain-shared features on the hypersphere manifold;

[0162] 2. Metric learning strategy based on hypersphere manifold and KL divergence:

[0163] The specific steps are as follows:

[0164] The present invention first trains the feature recognition network using the hypersphere manifold loss function to accelerate the convergence process of the model. The hypersphere manifold loss function is as follows:

[0165] ;

[0166] Among them, s is the proportional coefficient of the cosine value cos of the included angle, which is used to accelerate the training process of the model. C represents the number of pedestrian abnormal behavior labels, is the number of pedestrian images;

[0167] Then, a KL divergence constraint is constructed to further train the feature recognition network. The present invention uses a probability constraint method based on Kullback-Leibler divergence to measure the probability prediction values of different modalities of the same person and The similarity between them is used to further improve the modality independence of cross-modal features. It is assumed that the number of samples included in a batch of two different modalities of the input is N. Then to The KL divergence value of can be calculated by the following formula:

[0168] ;

[0169] Since the KL divergence is asymmetric, therefore to The divergence of is:

[0170] ;

[0171] Among them, is the pedestrian behavior prediction result in the visible light modality, is the pedestrian behavior prediction result in the infrared light modality.

[0172] V. Construct a cross-modal pedestrian feature registration network based on dual-aligned feature embedding. The network structure is as Figure 4 shown. The obtained feature vectors are input into the cross-modal pedestrian feature registration network based on dual-aligned feature embedding to obtain the final pedestrian image abnormal behavior detection result. Secondly, a modality consistency metric learning strategy is proposed to train the cross-modal pedestrian feature registration network based on dual-aligned feature embedding.

[0173] 1. Cross-modal pedestrian feature registration network based on dual-aligned feature embedding:

[0174] The present invention constructs a cross-modal pedestrian feature registration network based on dual-alignment feature embedding to obtain the final detection result of abnormal behaviors in pedestrian images. The cross-modal pedestrian feature registration network based on dual-alignment feature embedding includes a ResNet50 backbone network, a global average pooling layer, a convolutional layer, a batch normalization layer, and a fully connected layer. The ResNet50 backbone network and the global average pooling layer are used to extract fine-grained features. Then, the convolutional layer and the batch normalization layer are used to reduce the dimension of the fine-grained features, reducing them to feature vectors with a dimension of 256. Finally, the dimension-reduced fine-grained features are input into the fully connected layer to obtain the final detection result of abnormal behaviors in pedestrian images. The specific steps are as follows:

[0175] 1) Input into the ResNet50 backbone network. In the present invention, the network before the last pooling layer of ResNet50 is used as the ResNet50 backbone network. Input into the ResNet50 backbone network, and the output is a multi-channel feature map M. The specific calculation formula is as follows:

[0176] ;

[0177] where M is a three-dimensional tensor;

[0178] 2) Randomly divide the feature map into K non-overlapping parts along the vertical direction, and then perform global average pooling calculation on each part to obtain the fine-grained feature . For the th part , the calculation formula is as follows:

[0179] ;

[0180] where is a feature vector with a dimension of 2048;

[0181] 3) Input the fine-grained feature into the convolutional layer and the batch normalization layer to reduce the dimension, reducing it to a feature vector with a dimension of 256*K;

[0182] 4) Input the K feature vectors of the dimension-reduced fine-grained features into K fully connected layers with different weights to obtain the detection result of abnormal behaviors in pedestrians.

[0183] 2. Modal consistency metric learning strategy:

[0184] The present invention trains a cross-modal pedestrian feature registration network based on double-aligned feature embedding through a modal consistency metric learning strategy to improve the cross-modal recognition accuracy of the network; the modal consistency metric learning strategy consists of an intra-class distribution constraint loss function and an inter-class correlation constraint loss function, which are specifically as follows:

[0185] The proposed cross-modal alignment is achieved through modal consistency metric constraints and mainly consists of two parts, namely an intra-class distribution constraint loss function and an inter-class correlation constraint loss function;

[0186] 1) Construct the intra-class distribution constraint loss function:

[0187] For pedestrian features in different modalities , they conform to the following distribution representation:

[0188] ;

[0189] Here, V and I are used to represent the visible light modality and the infrared modality respectively. For a pedestrian with a class label of i, since the visible light image and the infrared image of this pedestrian are different forms of the same objective entity in different modal representation spaces, and the features after passing through the feature extraction network should be a high-level cognitive-level feature, it is thus hoped that the feature distribution of the visible light image of the same pedestrian should be as similar as possible to the feature distribution of its infrared image. The Jensen-Shannon divergence is used to measure the similarity between the two distributions and :

[0190] ;

[0191] where is the mixed distribution , represents the Kullback-Leibler divergence. The Jensen Shannon divergence measuring the two distributions can also be calculated through the Jeffreys similarity , then and The similarity between can be expressed as:

[0192] ;

[0193] The above formula can be rewritten as:

[0194] ;

[0195] Directly minimize the following formula to minimize the divergence and obtain the intra-class distribution constraint loss function:

[0196] ;

[0197] where and are the means of features and respectively, and are vectors composed of the diagonal elements of the covariance matrices and respectively. The corresponding calculation formulas are as follows:

[0198] ;

[0199] 2) Construct the inter-class correlation constraint loss function:

[0200] The Pearson correlation coefficient matrix is used as the representation of the correlation of image features in a visible light batch. For two n-dimensional feature vectors and , their Pearson correlation coefficient can be expressed as:

[0201] ;

[0202] Denote the cumulative features in a visible light batch as , and the cumulative features in the corresponding infrared light batch as , where N is the number of pedestrian images in each batch of each modality during training. By calculating the Pearson correlation coefficient between the features of every two images in a visible light batch, the Pearson correlation coefficient matrix of all images in a visible light batch can be obtained:

[0203] ;

[0204] where represents the correlation coefficient between the -th feature and the -th feature in the image features of a visible light batch. Similarly, the correlation coefficient matrix of a batch of infrared images can be obtained:

[0205] ;

[0206] Representing the two correlation coefficient matrices using the column vector groups of the matrices, they can be rewritten in the following format:

[0207] ;

[0208] By constructing a relevant loss function to constrain the matrix and consistency, an inter-class correlation constraint loss function is obtained, and the formula is:

[0209] ;

[0210] where is the Pearson correlation coefficient matrix of a batch of infrared images, the th element in is the Pearson correlation coefficient matrix of a batch of visible light images, the th element in

[0211] VI. Model Deployment: First, deploy the trained spatio-temporal feature encoding and decoding network, feature recognition network, and cross-modal pedestrian feature registration network based on dual-aligned feature embedding to the backend server; second, input the infrared light pedestrian video data collected by the infrared camera and the visible light pedestrian video data collected by the visible light camera into the backend server for video sampling and frame division processing to obtain frame-by-frame pedestrian images; finally, input the pedestrian images into the spatio-temporal feature encoding and decoding network to obtain a preliminary recognition result of pedestrian abnormal behavior, input the images containing pedestrians with abnormal behavior into the feature recognition network to obtain feature vectors, and input the feature vectors into the cross-modal pedestrian feature registration network based on dual-aligned feature embedding to obtain the final pedestrian abnormal behavior detection result.

[0212] VII. Experimental Results of Abnormal Behavior Detection on Multi-source Video Data with Spatio-temporal Consistency Enhancement:

[0213] Experimental Datasets:

[0214] CUHKAvenue: This dataset contains 16 videos for training and 21 videos for testing, and the duration of each video is about 1 minute;

[0215] UCSDPedestrian: This dataset consists of two subsets, namely UCSD Ped1 and UCSD Ped2, and existing methods are independently carried out on these two sets. Among them, UCSD Ped1 includes 34 short video sequences for training and 36 video sequences containing 40 abnormal events for testing, and the length of each video sequence is 200 frames. UCSD Ped2 includes 16 training videos and 12 test videos with 12 abnormal events, and each video has about 170 frames. The normal events in these two subsets are people walking on the sidewalk, and the abnormal events include bicycles, wheelchairs, vehicles, and people playing skateboards in the crowd.

[0216] ShanghaiTech Campus (SHTech): This dataset includes more than 270,000 training frames and more than 42,000 test frames, among which 17,000 frames are abnormal frames. This dataset covers 13 different scenarios.

[0217] Table 2 Abnormal event recognition accuracy on the experimental dataset

[0218]

[0219] As shown in Table 2, on the CUHKAvenue data, the method proposed in the present invention has improved the abnormal time detection accuracy by nearly 6 percentage points compared with the STAE-O method that uses additional optical flow information. On the Ped2 database, the method proposed in the present invention has improved the detection effect by 2 percentage points compared with the previous best method, Frame-Pred. The detection effect has been significantly improved. From the experimental results in the table, the method proposed in the present invention has exceeded the existing best algorithms on three widely used datasets.

[0220] To verify the effectiveness and robustness of the method of the present invention, the present invention selected the ShanghaiTech Campus (SHTech) dataset to conduct 6 pedestrian abnormal behavior detection experiments. Among them, the abscissa is the number of training rounds, and the ordinate represents the value of the loss function. Each curve represents an experiment. The total loss function convergence during the training process of the 6 experiments is as Figure 5 shown. It can be seen from the figure that the method proposed in the present invention has a faster training convergence speed on the six datasets, indicating that the model proposed in the present invention can better learn the fusion features of visible light images and infrared light images, and the training strategy proposed in the present invention can accelerate the model convergence.

[0221] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0222] Although the specific implementation manners of the present invention are described above, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made without creative labor by those skilled in the art are still within the protection scope of the present invention.

Claims

1. An intelligent video monitoring method based on multi-source data fusion, characterized in that: The following steps are involved: S1, acquiring multi-source video data, including visible light pedestrian video data and infrared light pedestrian video data, and preprocessing the data; S2, input the preprocessed video data into the spatiotemporal feature encoding and decoding network to obtain the preliminary results of abnormal behavior detection of pedestrian images; S3, based on the preliminary results of abnormal behavior detection of pedestrian images obtained in S2, extract the pedestrian images in abnormal behavior to obtain abnormal behavior image data; S4, inputting the abnormal behavior image data obtained in S3 into a feature recognition network to obtain a feature vector; the feature recognition network includes a feature learning module and a feature mapping module, and the pedestrian image is input into the feature learning module to obtain the local features of the image , local features Input to the feature mapping module to obtain the hypersphere domain mapping features ; The feature recognition network is trained based on the metric learning strategy of hypersphere manifold and KL divergence; the feature learning module consists of a convolutional neural network, a fully connected layer, and a batch normalization layer; the feature mapping module consists of a random inactivation layer, a fully connected layer, and a two-norm normalization layer; The feature recognition network is trained based on the metric learning strategy of hypersphere manifold and KL divergence, specifically: First, the hypersphere manifold loss function is used to train the feature recognition network to accelerate the convergence process of the model. The hypersphere manifold loss function is as follows: Among them, s is the proportional coefficient of the cosine value of the angle cos, which is used to accelerate the training process of the model, and C represents the number of abnormal behavior labels of pedestrians. is the number of pedestrian images contained in a batch; Represents the vector angle between different behavior categories in different pedestrian images; Indicates The angle between the abnormal behavior prediction output vector of a pedestrian image and the true label vector; Then, a KL divergence constraint training feature recognition network is constructed; that is, a probability constraint method based on Kullback-Leibler discreteness is used to measure the probability prediction value of different modalities of the same person. and The similarity between them improves the modality independence of cross-modal features. arrive The KL divergence value of Calculated by the following formula: Since the KL divergence is asymmetric, arrive Divergence for: in, is the pedestrian behavior prediction result under visible light mode, Pedestrian behavior prediction results in infrared light mode; S5, the feature vector obtained in S4 is input into the cross-modal pedestrian feature registration network based on dual alignment feature embedding. First, the backbone network and global average pooling layer are used to extract The fine-grained features are then used, and the convolution layer and batch normalization layer are used to reduce the dimension of the fine-grained features. Finally, the fine-grained features after dimensionality reduction are input into the fully connected layer to obtain the final pedestrian image abnormal behavior detection results.

2. The intelligent video monitoring method based on multi-source data fusion according to claim 1, characterized in that: The spatiotemporal feature encoding and decoding network in S2 includes a 3D encoding module, a 2D decoding module, a resampling module, a temporal downsampling module and a spatiotemporal consistency discriminator. and infrared video data The specific steps of inputting the spatiotemporal feature encoding and decoding network to obtain the preliminary results of abnormal behavior detection are as follows: S21, Pedestrian Video and Uniformly sample to k frames, and form a video sequence consisting of k video frames Input to the spatiotemporal feature encoding and decoding network, where Indicates the first Frame image; S22, will After inputting 5 3D encoding modules connected in series, we get The hidden feature vector of , the specific calculation formula is as follows: in , , , and Respectively represent a first 3D encoding module, a second 3D encoding module, a third 3D encoding module, a fourth 3D encoding module and a fifth 3D encoding module; , , and are output features of the first 3D encoding module, the second 3D encoding module, the third 3D encoding module, and the fourth 3D encoding module respectively; S23, will Input resampling module, Resample the hidden space to obtain new features , the specific calculation formula is as follows: in, is a random vector between zero and one, and They are The mean and variance of the standard normal distribution in high-dimensional space; and Learned by two fully connected layers The data distribution is obtained; During the sampling process, Kullback Leibler discreteness is used as a constraint for resampling to ensure the characteristics after resampling With the original features They obey the same distribution, where the calculation formula of Kullback Leibler discreteness is as follows: in, represents the Kullback Leibler discreteness, is the Kullback Leibler discrete value; S24, will After inputting the five 2D decoding modules connected in series, it is decoded into the next frame image of the input sequence. , the specific calculation formula is as follows: in, , , , and Respectively represent a first 2D decoding module, a second 2D decoding module, a third 2D decoding module, a fourth 2D decoding module and a fifth 2D decoding module; , , and are output features of the first 2D decoding module, the second 2D decoding module, the third 2D decoding module, and the fourth 2D decoding module respectively; is the next frame image of the generated input sequence; represents the time domain downsampling module, Represents the feature fusion function, which directly merges two tensors with the same height and width in the channel dimension; the convolution operations in the 2D decoding module are all two-dimensional convolutions; S25, and After inputting the spatiotemporal consistency discriminator at the same time, we can obtain the preliminary abnormal behavior detection results. .

3. The intelligent video monitoring method based on multi-source data fusion according to claim 1, characterized in that: The spatiotemporal feature encoding and decoding network is optimized and trained by a spatiotemporal consistency training strategy, and the cost function of the spatiotemporal consistency training strategy is composed of a pixel-level error square function and an adversarial loss function; The pixel-level square error function , the entire spatiotemporal feature encoding and decoding network is trained by minimizing the following cost function: in, is the height and width of the image in the input video sequence. This function calculates the result pixel by pixel. Its true value The square of the error between It is the symbol for modulo calculation; The adversarial loss function is composed of a generation loss function and a discrimination loss function; Generating loss function for: Discriminative loss function for: in, represents a matrix of all zeros, Pedestrian image samples generated by decoder G in the spatiotemporal feature encoding and decoding network , represents the concatenation of time channels, Represents dimension and consistent identity matrix; is the actual video frame sequence, the corresponding real video frame and the decoder input to the spatiotemporal feature encoding and decoding network By cascading on the time channel, we can get , is the sampling value of the real video sequence, ; The training process of the spatiotemporal consistency enhancement network uses the decoder in the spatiotemporal feature encoding and decoding network and the spatiotemporal consistency discriminator to alternately train until both reach convergence.

4. The intelligent video monitoring method based on multi-source data fusion according to claim 1, characterized in that: The specific processing process of the cross-modal pedestrian feature registration network based on dual alignment feature embedding in S5 is: S51, will Input ResNet50 backbone network; the network before the last pooling layer of traditional ResNet50 is used as ResNet50 backbone network. Input the ResNet50 backbone network and output a multi-channel feature map M. The specific calculation formula is as follows: Where M is a three-dimensional tensor; S52, the feature map The image is randomly divided into K non-overlapping parts in the vertical direction, and then global average pooling is performed on each part to obtain fine-grained features. , for the Part , the calculation formula is as follows: in is a feature vector of 2048 dimensions; S53, the fine-grained features The input convolutional layer and batch normalization layer reduce The dimension is reduced to a feature vector with a dimension of 256*K; S54, respectively inputting the K feature vectors of the fine-grained features after dimensionality reduction into K fully connected layers with different weights to obtain the final pedestrian image abnormal behavior detection result.

5. The intelligent video monitoring method based on multi-source data fusion as claimed in claim 4, characterized in that: The cross-modal pedestrian feature registration network based on dual alignment feature embedding is trained through a modality consistency metric learning strategy, including an intra-class distribution constraint loss function and an inter-class correlation constraint loss function, as follows: The intra-class distribution constraint loss function is minimized by minimizing the following formula Discreteness, obtain the intra-class distribution constraint loss function : in and are the means of the visible light modality pedestrian features and the infrared modality pedestrian features, and is the standard deviation of the visible light modality pedestrian features and the infrared modality pedestrian features, , , and The corresponding calculation formulas are as follows: in, Indicates The features of infrared pedestrian images, Indicates Features of visible light modality pedestrian images; The inter-class correlation constraint loss function The specific calculation formula is as follows: in is the Pearson correlation coefficient matrix of a batch of infrared images The elements, is the Pearson correlation coefficient matrix of a batch of visible light images The elements, and The calculation formula is as follows: 。 6. An intelligent video surveillance system based on multi-source data fusion, applied to the intelligent video surveillance method according to any one of claims 1 to 5, characterized in that: include: Video acquisition module and backend server; The video acquisition module is used to acquire multi-source video data; The backend server is internally deployed with a video preprocessing module, a spatiotemporal feature encoding and decoding network, an abnormal behavior extraction module, a feature recognition network, and a cross-modal pedestrian feature registration network based on dual-aligned feature embedding; The video preprocessing module is used to standardize and remove noise from multi-source videos, and uniformly adjust spatial resolution and brightness; The spatiotemporal feature encoding and decoding network is used to input the preprocessed video data to obtain preliminary results of abnormal behavior detection of pedestrian images; An abnormal behavior extraction module is used to extract pedestrian images in abnormal behavior and obtain abnormal behavior image data; The feature recognition network is used to obtain a feature vector; The cross-modal pedestrian feature registration network based on dual-aligned feature embedding is used to obtain the final pedestrian image abnormal behavior detection result.