Road scene recognition method and system based on modal information evaluation
By building a collaborative network with multimodal features and backbone classification network, dynamically adjusting the modal information weight, the problem of low road scene recognition accuracy under complex lighting conditions is solved, and efficient scene recognition under complex lighting conditions is achieved.
Patent Information
- Application Number
- CN202510503972.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art has low road scene recognition accuracy under complex lighting conditions, and the single-modal information recognition algorithm is not robust, and traditional methods cannot effectively deal with noise interference caused by image illumination changes.
The road scene recognition method based on modal information evaluation is adopted, and the collaborative network and backbone classification network are constructed using multimodal features. The modal quality is evaluated through VGG16-places365 and BRISQUE algorithms, and unsupervised joint representation is performed by combining the deep confidence network and the DBN network. The modal information weight is dynamically adjusted, and the relative entropy regularization term is added for sparse constraints.
It improves the robustness and accuracy of road scene recognition, can effectively identify scenes under complex lighting conditions, reduces the risk of overfitting in traditional methods, and improves scene classification performance.
Smart Images

Figure CN120408372A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road scene recognition, and particularly to a road scene recognition method and system based on modal information evaluation. Background Art
[0002] Road scene recognition belongs to the category of scene understanding technology and can provide effective auxiliary semantic information for vehicle violation recognition. Although certain progress has been made in the research on traditional scene recognition, different from general video scene recognition, there is a practical situation where the illumination conditions are complex and variable in outdoor road scenes. Figure 1 The road surveillance videos under normal illuminance, insufficient illuminance, and excessive illuminance are respectively shown. When the video illumination is excessive or insufficient, a large amount of noise will be introduced when representing the visual scene or moving targets such as vehicles and pedestrians. Therefore, complex road scene recognition still faces many challenges.
[0003] Researchers have always used a pre-trained deep learning model to obtain good representation information of video scenes and continuously expanded the research in the field of scene recognition applications. However, the recognition algorithm under the condition of single-modal information is restricted by the information expression ability and generally has the problem of low robustness when applied to actual scenes with complex illuminance.
[0004] In recent years, by means of monitoring audio and video information, the method of using multi-modal information fusion has been proven to be an effective strategy for video scene recognition and video content analysis. For example, in the field of short video scene recognition, Guo et al. proposed a multi-modal enhanced semantic fusion network for scene recognition on the basis of using audio and text as auxiliary modal information of strong modal information of image vision. The network realizes weak modal semantic enhancement by minimizing the loss of the objective function of semantic distance, and realizes the scene classification task of fully supervised deep learning by establishing an adaptive learning of the fusion weights of multi-modalities. Hori et al. extracted image features, motion features and audio features by using VGGNet, C3D and MFCC respectively, and proposed a multi-modal attention mechanism to dynamically use different modal features for classification. Wu et al. addressed the problem of lack of semantic consistency design in multi-modal feature fusion, and used a multi-task learning architecture to complete audio-visual feature fusion, and added a semantic consistency metric loss to obtain constraint conditions. Experiments have proved that this method makes full use of the complementarity between various modal features and has achieved better results in the application of violent video scene recognition. However, the above fusion analysis methods all rely on the assumptions such as clear scene labels and balanced data; and it has also been found in practical applications that traditional strong modal image information is more sensitive to illuminance, and there are large information differences in image convolution features under different illuminance conditions. However, the traditional relative entropy loss cannot enable the model to learn the strong and weak relationship between the image and audio information expressions, and the modal information weights have strong randomness. Therefore, when the image illuminance changes violently, a large amount of irrelevant and random interference noise will be generated in the joint representation of the fusion features, resulting in a decrease in the scene recognition accuracy. Summary of the Invention
[0005] The technical solution of the present invention aims at the technical problem that the existing technical solutions are too single, and provides a solution significantly different from the prior art. It mainly provides a road scene recognition method and system based on modal information evaluation to solve the technical problem that the violent change of image illuminance will lead to a decrease in scene recognition accuracy proposed in the above background art.
[0006] The technical solution adopted by the present invention to solve the above technical problems is as follows:
[0007] A road scene recognition method based on modal information evaluation includes: constructing a scene recognition multi-task model based on modal information evaluation based on the extracted multi-modal features; the multi-modal features include audio features and video image frame features; the scene recognition multi-task model includes a collaborative network and a backbone classification network, and the collaborative network uses VGG16-places365 for visual feature extraction as a modal quality evaluation network, and uses the BRISQUE algorithm based on local normalized luminance coefficients as a teacher model to provide learning reference labels for the modal quality evaluation network; the backbone classification network uses a multi-modal DBN network to perform unsupervised joint representation on visual and audio modal scene information.
[0008] Further, high-dimensional feature extraction and dimensionality reduction are respectively performed on the image frames and audio in the video based on the open-source VGG16-places365 and VGGish networks to form descriptive features for input to the backbone classification network.
[0009] Specifically, the method for extracting audio features is: sampling the video segment, obtaining log mel cepstral features after processing the sampled data, re-grouping the frames and inputting them into the VGGish network model, and each audio segment obtains a 128-dimensional audio feature, and after concatenation, the audio feature output V is obtained. (a) 。
[0010] Specifically, the method for extracting video image frame features includes:
[0011] A. Video segment cutting: For the surveillance video sample under single-shot conditions, the surveillance video sample is cut into several video segments at a specified time interval.
[0012] B. Frame-level feature extraction: Randomly extract any video image frame from the cut video segment, and obtain the output of the first fully connected layer after network convolution and pooling operations and define it as the frame-level feature description matrix M. The row vector and column vector respectively represent the changes of the feature descriptor in the spatial dimension and the time dimension, and are expressed as follows:
[0013]
[0014] Among them, each row vector of the description matrix M is the frame-level feature of each video image frame, and the frame-level feature dimension is 4096; the frame-level feature descriptor of the i-th frame of the image frame in the j-th dimension is represented as F i,j ; n is the number of video segments in the current original video sample;
[0015] C. Video-level feature extraction: Use the method of mean clustering of the frame-level feature sequence to obtain the video-level feature:
[0016] Where
[0017] Among them, V j (f) represents the j-th video-level feature.
[0018] Furthermore, in the collaborative network, the classification features come from the empirical distribution of the locally normalized luminance under the spatial domain statistical model.
[0019] Furthermore, in the collaborative network, the last convolutional operation of the pre-trained model VGG16-places365 is used as the modal quality evaluation feature matrix f v , and the pooling vector η is further calculated as follows:
[0020] η = vec(f v )
[0021] Among them, the vec operation is to expand the matrix into a one-dimensional vector; after spatial mapping and l2 normalization of η, the signed feature η' and the evaluation feature η'' can be obtained respectively in sequence:
[0022]
[0023] Among them, sign is the operation of taking the signed square root, and e is the element-wise multiplication operation; the fused feature is input into the fully connected layer to calculate and generate the output y of the student network; a certain frame of the video segment is randomly selected as the input of the BRISQUE algorithm, and the score a of the teacher model is calculated; through normalization processing, the value range is ensured to be consistent with the predicted output y of the collaborative network; the empirical loss between the prediction and the BRISQUE output score is calculated using the l2 norm:
[0024]
[0025] Among them, N is the number of samples in each training batch; finally, the stochastic gradient descent method is used to complete the fine-tuning of the collaborative learning network;
[0026] Through the learning branch of the collaborative network, VGG16-places365 participates in multi-task learning, and the modal quality evaluation score δ is obtained through inference and prediction. The value range of the normalized modal quality evaluation score δ is [0, 1], and the higher the score, the lower the illuminance of the video sample and the worse the quality of the visual modal information.
[0027] Furthermore, in the backbone classification network, the DBN network model uses the probabilistic generative model to establish a joint distribution between the observed data and the labels, which is expressed as follows:
[0028] p(v, h 1 , h 2 , L, h l ) = p(v|h 1 )p(h 1 |h2 )...p(h l-2 |h l-1 )p(h l-1 ,h l )
[0029] Among them, v represents the observed data, and h represents the label.
[0030] Furthermore, in the backbone classification network, the first layer RBM in the DBN network model adopts the Gaussian-Bernoulli model, and its joint probability p and energy function E are specifically defined as follows:
[0031]
[0032] Among them, Z is the normalization factor obtained by approximate calculation of annealed importance sampling, σ is the variance of the Gaussian unit in the visible layer, b i represents the bias of the visible unit, D represents the number of visible units, F1 represents the number of hidden layer units, a j represents the bias of the hidden layer unit;
[0033] The data types of the hidden layer nodes are all binary structures. Define the hidden layer RBM as the Bernoulli-Bernoulli type, and the output multi-modal joint distribution probability is defined as follows:
[0034]
[0035] Among them, θ represents the DBN model parameter, h f (1) represents the feature extraction hidden layer, the hidden unit for capturing the internal features of a single modality, h a (1) represents the association hidden layer, the hidden unit for modeling cross-modal interactions, h (2) is the high-level abstract hidden layer for integrating the high-order representations of multi-modal information;
[0036] The learning and inference process uses the contrastive divergence CD algorithm to approximate the maximum likelihood solution, so that the hidden layer units can obtain the correlation of the high-order features of the two modalities. The solution process of the posterior probability of the model is as follows:
[0037]
[0038] Among them, h fj (1) represents the encoded single-modal feature, h aj (1) is used to capture cross-modal associations, h k (2)It is used to integrate multimodal information to form high-level representations; the learning process of the CD algorithm is to maximize the variational lower bound under the condition that the parameters θ of the DBN model are fixed for the variational parameter μ; similarly, when the variational parameter μ is fixed, the optimal solution of the network model parameters θ is found through the MCMC stochastic approximation method to complete the inference process.
[0039] Furthermore, in the backbone classification network, a Softmax classifier is cascaded at the end of the DBN network to complete the prediction output of the video scene, and the network weights are supervised and optimized through a loss function that incorporates the modal information evaluation factor δ. The relative entropy loss function is defined as:
[0040]
[0041] where G is the true sample distribution, Q is the posterior probability output of the multimodal DBN model, and ɑ is the modal weight balance factor;
[0042] A regularization term is added to the relative entropy loss function to perform sparse constraints on the hidden layer nodes, and the resulting final logarithmic likelihood cost function is:
[0043]
[0044] where K in the relative entropy regularization term is the number of hidden layer nodes, T is the total number of training samples; γ is the weight factor of the regularization term, and the sparse coefficient ρ is used to control the sparsity of the hidden layer;
[0045] The DBN model is continuously iteratively reduced through the gradient descent method until it finally converges to complete multi-task learning.
[0046] The present invention also provides a road scene recognition system based on modal information evaluation, including a processor and a memory. When the computer program stored in the memory runs in the processor, it executes the above-mentioned road scene recognition method based on modal information evaluation.
[0047] Compared with the prior art, the beneficial effects of the present invention are:
[0048] (1) The collaborative network of the scene classification multi-task model proposed in the present invention is based on a modal information evaluation strategy. This strategy assumes that illumination changes are an important determinant of image quality changes and can cause the exchange of strong and weak properties of heterogeneous modal information. This strategy is based on the BRISQUE algorithm teacher network learning of the normalized brightness coefficient to obtain the modal information score (i.e., the modal information quality evaluation factor), and outputs it to the backbone network loss function as the loss weight of the negative sample that penalizes image quality loss. Compared with the traditional scene recognition model, the present invention links image quality with brightness information and proposes a "modal information quality evaluation factor" to evaluate the quality of video image frames, so that the modal information weight can be dynamically adjusted according to the illumination change, thereby achieving the purpose of robust recognition model.
[0049] (2) This invention utilizes a deep belief network to perform nonlinear feature transformation on high-dimensional features, establishes correlation between the two heterogeneous information through the joint probability distribution of audio and image modal features, and effectively achieves feature complementarity. Compared with traditional scene recognition models, this invention leverages the deep belief network's joint probability associative advantage for heterogeneous information and further improves scene recognition performance through the complementary correlation features of audio and image modalities.
[0050] (3) The present invention adds a relative entropy regularization term to the backbone classification network loss function. This cost function can be used to impose sparse constraints on hidden layer nodes, thereby achieving effective feature space partitioning. Compared with traditional scene recognition models, the present invention emphasizes the sparse nature of high-dimensional scene features, can achieve effective feature space partitioning, and solves the problem of overfitting that is prone to occur during the training process of traditional methods.
[0051] In summary, this paper explores more feature representations adapted to complex scenes from multimodal video features, proposes a multi-task CNN-DBN (Convolutional Deep Belief Network) model for scene recognition based on modal information evaluation, and trains it using a small dataset of scene-labeled videos. Testing on a constructed video surveillance dataset demonstrates that the proposed method outperforms mainstream methods in scene classification accuracy.
[0052] The present invention will be explained in detail below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 Road scenes under different illumination conditions;
[0054] Figure 2 It is a multi-task model based on modal information evaluation;
[0055] Figure 3 This is the VGGish model structure diagram;
[0056] Figure 4It is the feature structure diagram of VGG16;
[0057] Figure 5 It is the structure diagram of DBN (Deep Belief Network);
[0058] Figure 6 It is the video surveillance dataset under different road scenarios in the embodiment;
[0059] Figure 7 It is the spatial visualization distribution map of multi-modal fusion features in the embodiment;
[0060] Figure 8 It is the results before and after ablation of the model scene recognition accuracy rate under the night dataset in the embodiment, where Figure (a) is the result before ablation and Figure (b) is the result after ablation;
[0061] Figure 9 It is the comparison diagram of P-R curves of three different models in the embodiment. Detailed implementation manners
[0062] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant attached drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in different forms and is not limited to the embodiments described in the text. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0063] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used in the present invention includes any and all combinations of one or more of the related listed items.
[0064] Embodiment: A road scene recognition method based on modal information evaluation, comprising:
[0065] 1. Discover more feature representations adapted to complex scenarios from video multi-modal features.
[0066] Based on the open-source VGG16-places365 and VGGish networks, high-dimensional feature extraction and dimensionality reduction are respectively performed on the image frames and audio in the video to form descriptive features for the input of the backbone classification network. This input feature not only maintains the spatial globality of the image feature but also maintains the temporal globality of the audio feature.
[0067] Multi-modal Feature Extraction: By leveraging the excellent feature representation ability of the Convolutional Neural Network (CNN) and its accuracy in solving complex problems, the performance of unsupervised learning and weakly supervised learning has been significantly improved. The core of the convolution algorithm is weight sharing and pooling. The deep learning network model constructed by the convolution operation with sparse interaction and parameter sharing has strong advantages in processing variable-dimensional inputs. For the image input Zl at the (l + 1)-th layer in the network, the output Zl+1 after the convolution operation can be expressed as:
[0068]
[0069] where w, s, b, f, and k are defined as weights, stride, bias, convolutional kernel size, and number of channels respectively; (i, j) and (x, y) represent the pixel positions of the input and output of the pooling layer respectively. It can be seen that the concept of weight sharing makes the convolutional neural network have translational equivariance, so it is widely used in the processing of video image frames. The present invention designs to complete the extraction task of two-modal feature data of audio and video image frames based on the VGG model.
[0070] (1) Audio Feature Extraction
[0071] Since most of the audio scene information is event-driven, VGGish pre-trained on the speech dataset AudioSet can be used as the feature extraction model for audio scene information. In feature preprocessing, only the last 0.96 seconds of each 1-second video segment is resampled at 16KHz. After the sampled data undergoes operations such as fast Fourier transform, framing, and windowing, log mel spectrogram features are obtained. After re-framing, the input to the VGGish network model with a dimension of 96×64 can be obtained. In terms of parameter settings, a frame length of 25 milliseconds, a frame shift of 10 milliseconds, and a Hann window are selected. As Figure 3 shown, the VGGish structure refers to the VGG11 model and deletes the last set of convolution operations. The overall network has only 4 sets of convolutional layers, and the output of the fully connected layer is 128.
[0072] Assuming the original video sample has a duration of n seconds, n audio segments will be obtained during the sampling process. Through VGGish, each audio segment can obtain a 128-dimensional audio feature. After concatenation, the audio feature output V (a) (128×n dimensions) is obtained. Each audio feature representation can generate an audio shared modal codebook, and the shared codebook can be used as an effective information supplement for the deep global features of images to handle scene recognition tasks in complex environments.
[0073] (2) Video Image Frame Feature Extraction
[0074] In the visual feature extraction task, the present invention refers to using VGG16-places365 pre-trained on the scene training set Places365 as the extractor for video image frame features, and performs fine-tuning on this basis. The extractor network consists of 16 convolutional networks with trainable parameters. In terms of parameter selection, a convolutional kernel size of 3×3 and a maximum pooling of 2×2 are adopted. Its structure is as Figure 4 shown.
[0075] It should be noted that in order to extract the texture structure information of video frame images as much as possible, the present invention selects maximum pooling for all pooling operations in the VGG16 structure. In order to reduce the impact caused by jitter on the surveillance video, fully exploit the motion feature information of video temporal frames, and improve the recognition efficiency and robustness of the later fusion model, the present invention designs to use a frame-level feature descriptor to express the scene features of a single-frame image, and uses a temporal aggregation strategy to obtain a video-level feature descriptor for expressing the scene features of a video segment. The detailed method is as follows:
[0076] A. Video segment cutting: For the surveillance video sample under single-shot conditions, the surveillance video sample is cut into several video segments at a specified time interval. Assuming that the original video sample duration is n seconds and the interval duration is m seconds, then n / m video segments can be obtained. Considering the temporal synchronization of the two modal features and referring to the audio feature sampling duration, the present invention adopts a video segment downsampling interval duration of 1 second. Therefore, the original video sample is also cut into n video segments. All frames in the video segment are downsampled in the spatial dimension, and the image frame size is adjusted to 224×224 pixels and used as the input of the VGG16-places365 network.
[0077] B. Frame-level feature extraction: Randomly extract any video image frame from the cut video segment. After network convolution and pooling operations, the output of the first fully connected layer is obtained and defined as the frame-level feature description matrix M. The row vector and column vector respectively represent the changes of the feature descriptor in the spatial dimension and the temporal dimension. The expression is as formula 2:
[0078]
[0079] Among them, each row vector of the description matrix M is the frame-level feature of each video image frame, and the frame-level feature dimension is 4096. The frame-level feature descriptor of the i-th frame of the image frame in the j-th dimension is represented as F i,j . n is the number of video segments in the current original video sample.
[0080] C. Video-level feature extraction: Since the video samples of scene categories are all single-shot, most of the scene information and moving target texture features are in a single video frame image, which can be regarded as a static background in a certain temporal sense. Different from applications such as behavior classification, event detection, and target classification, since the contribution of the temporal linear features in the video frame sequence to the scene classification task is relatively low, the temporal features of consecutive video frames are not learned here. At the same time, by observing the description matrix, it can be found that the video scene feature descriptor F i,j changes more in the spatial dimension than in the temporal dimension. Selecting the static statistical metric method can better obtain the computational efficiency. Therefore, the present invention uses the method of clustering the mean of the frame-level feature sequence to obtain the video-level features.
[0081] Among them
[0082] Among them, V j (f) represents the j-th video-level feature.
[0083] Each video sample will correspond to 1 clustering feature, that is, the video image frame feature V (f) , and the feature dimension is still 4096.
[0084] 2. Propose a multi-task model for scene recognition based on modal information evaluation.
[0085] As Figure 2 shown, it includes a collaborative network and a backbone classification network.
[0086] Collaborative network: To address the problem of non-uniform distortion caused by illumination changes in real scenes, the VGG16-places365 for visual feature extraction is selected as the modal quality evaluation network to generate modal quality evaluation features. At the same time, to better describe the mapping relationship between the luminance coefficient and the image quality, the BRISQUE (blind / referenceless image spatial quality evaluator) algorithm based on the locally normalized luminance coefficient is designed to be used as the teacher model to provide learning reference labels for the evaluation network. The present invention assumes that the change in image quality will cause a change in the strong weight attribute of the visual modal information, and the illumination change is an important determinant of the change in image quality. For example, in the surveillance video scene with poor lighting, the extracted visual high-dimensional features can be regarded as negative sample noise interference, and more dependence on the audio modal information features is required for the inference of the scene category.
[0087] Specifically, as a no-reference image quality assessment algorithm, the classification features of BRISQUE come from the empirical distribution of locally normalized luminance under the spatial domain statistical model, and can guide the visual representations extracted by VGG16-places365 to be more sensitive to contrast information through transfer learning. In the present invention, the last convolutional operation of the pre-trained model VGG16-places365 is used as the modal quality evaluation feature matrix f v , and further calculate to obtain the pooling vector η:
[0088] η = vec(f v ) (5)
[0089] where the vec operation is to expand the matrix into a one-dimensional vector. After spatial mapping and l2 normalization of η, the signed feature η' and the evaluation feature η'' can be obtained respectively in sequence:
[0090]
[0091]
[0092] where sign is the operation of taking the signed square root, and e is the element-wise multiplication operation. The fused feature is input into the fully connected layer to calculate and generate the output y of the student network. Assume that the illuminance does not change within any video segment, so a certain frame of the video segment is randomly selected as the input of the BRISQUE algorithm, and the score a of the teacher model is calculated. Through normalization processing, ensure that the value range is consistent with the collaborative network prediction output y. Use the l2 norm to calculate the empirical loss between the prediction and the BRISQUE output score:
[0093]
[0094] where N is the number of samples in each training batch. Finally, the stochastic gradient descent method SGD is used to complete the fine-tuning of the collaborative learning network. Through the learning branch of the collaborative network, VGG16-places365 participates in multi-task learning, and the modal quality evaluation factor δ is inferred and predicted, which provides a loss penalty factor for the loss function of the backbone classification network, so as to achieve the purpose of affecting the fusion weight of negative illuminance samples. After normalization, the value range of the evaluation score δ is [0,1], and the higher the score, the lower the illuminance of the video sample and the worse the quality of the visual modal information.
[0095] Trunk classification network: A two-stream multi-modal DBN network is constructed to receive the input of high-dimensional audio-visual features, and at the last hidden layer, the real-valued input is mapped to a binary feature space to form a feature fusion representation. The relative entropy loss is calculated based on the softmax output probability and the co-network modal quality evaluation score, and the classification network is fine-tuned by feedback. The present invention designs to use the multi-modal DBN network to perform unsupervised joint representation on the visual and audio modal scene information. Through the joint representation, the RBM hidden layer maps the global representation of the video frame image and the audio event to a latent semantic representation space, and at the same time effectively performs non-linear dimensionality reduction on the high-dimensional video scene feature vector to obtain a low-dimensional representation rich in scene structure information. Finally, a Softmax classifier is cascaded and supervised backfine-tuning is performed.
[0096] As a relatively common deep network model, DBN is widely used in various unsupervised training and supervised training, and can effectively complete data dimensionality reduction of high-dimensional features and perform classification tasks. Structurally, DBN is stacked by multiple restricted Boltzmann machines RBM, and adopts a pre-training strategy based on a greedy approach, which simplifies the training process while reducing the reconstruction error. Since each RBM has only one hidden layer, the stacked DBN network consists of one visible layer (input layer), several hidden layers and one output layer. The network structure is as Figure 5 shown:
[0097] It can be seen from the structure diagram that the nodes between the layers of DBN are fully connected, while there is no connection between the nodes of the same layer, and the hidden layer nodes are binary neural units. From the perspective of data representation, DBN, as a probability generation model and discriminant model, is used to describe the internal relationship between the observed data v and the label h, and evaluates p(v|h) and p(h|p) in the form of probability. Using the probability generation model to establish a joint distribution between the observed data and the label can be expressed as follows:
[0098] p(v,h 1 ,h 2 ,L,h l )=p(v|h 1 )p(h 1 |h 2 )...p(h l-2 |h l-1 )p(h l-1 |h l ) (9)
[0099] It can be seen that in the multi-modal input classification network constructed by the present invention, this structural relationship and data representation are conducive to establishing a unified learning model, mapping the image and audio modal information to independent feature spaces, and learning the joint representation between the two. Therefore, in the above-obtained visual modal information V (f)With the audio modality information V (a) Based on the constructed scene features, the backbone classification network will learn the joint probability distribution p(v f , v a ) of the two modality features. Among them, an unsupervised training is carried out through an unlabeled video training set to obtain an initial network weight close to the global optimum, and then a supervised learning is implemented using a labeled scene video set.
[0100] (1) Joint distribution learning of multi-modal features
[0101] Both the visual modality features and the audio modality features extracted by the deep learning model are of real number type. Therefore, the first layer RBM in the DBN model adopts a Gaussian-Bernoulli model, and its joint probability p and energy function E are specifically defined as follows:
[0102]
[0103]
[0104] Among them, Z is the normalization factor obtained by approximate calculation of annealed importance sampling (AIS), σ is the variance of the Gaussian unit of the visible layer, b i represents the bias of the visible unit, D represents the number of visible units, F1 represents the number of hidden layer units, a j represents the bias of the hidden layer unit. The data type of the hidden layer nodes is all binary structure. The hidden layer RBM is defined as Bernoulli-Bernoulli type, and the output multi-modal joint distribution probability is defined as follows:
[0105]
[0106] Among them, θ represents the DBN model parameters, h f (1) represents the feature extraction hidden layer, the hidden unit for capturing the internal features of a single modality, h a (1) represents the associated hidden layer, the hidden unit for modeling cross-modal interactions, h (2) is the high-level abstract hidden layer for integrating the high-order representation of multi-modal information.
[0107] The learning inference process uses the contrastive divergence CD algorithm to approximate the maximum likelihood solution, so that the hidden layer units can obtain the correlation of the high-order features of the two modalities. The solution process of the posterior probability of the model is as follows:
[0108]
[0109]
[0110] Among them, h fj (1)Denote the encoded unimodal feature, h aj (1) Used to capture cross-modal associations, h k (2) Used to integrate multimodal information to form high-level representations. As an algorithm that combines variational inference and the stochastic approximation method of MCMC (Markov Chain Monte Carlo), the learning process of the CD algorithm is to maximize the variational lower bound with the DBN model parameter θ fixed for the variational parameter μ. Similarly, when the variational parameter μ is fixed, the optimal solution of the network model parameter θ is found through the MCMC stochastic approximation method to complete the inference process.
[0111] (2) Classification loss function
[0112] Such as Figure 2 As shown, a Softmax classifier is cascaded at the end of the DBN network to complete the prediction output of the video scene. A loss function with a modal quality evaluation factor δ added is proposed to supervise and optimize the network weights. The relative entropy loss function is defined as:
[0113]
[0114] where G is the true sample distribution, Q is the posterior probability output of the multimodal DBN model, and ɑ is the modal weight balance factor. Considering the contribution rate of visual modal information in scene information classification, the present invention sets ɑ to 0.85. The influence factor is assigned and constrained within the interval (1, δmax] through the exponential function, approaching 1 under good lighting conditions. When the visual strong modal information has a high influence factor score due to poor illumination, the relative entropy loss will be further strengthened, which can be used to alleviate the atypical outlier features caused by the weakening of the expression ability of the strong modal information, and at the same time will guide the DBN network to adjust the modal information weights in the negative samples. At the same time, it is found during debugging that different scene structure features with small inter-class distances can lead to a high activation probability of hidden layer nodes. Therefore, the present invention adds a regularization term to the relative entropy loss function to perform sparse constraints on the hidden layer nodes, and the final logarithmic likelihood cost function obtained is:
[0115]
[0116] where K in the relative entropy regularization term is the number of hidden layer nodes, T is the total number of training samples. γ is the weight factor of the regularization term, and the sparse coefficient ρ is used to control the sparsity of the hidden layer. After cross-validation, ρ = 0.2 is set. The DBN model is continuously iteratively reduced by the gradient descent method until it finally converges to complete multi-task learning.
[0117] The CNN-DBN model uses an information quality evaluation strategy to quantify the strength of modal information, and utilizes the image quality loss to enable the network to automatically learn the modal information weights in negative samples. At the same time, the DBN network is used to establish a joint probability distribution to describe the correlation between two heterogeneous information, so it can also complete scene classification with high accuracy under complex illumination conditions.
[0118] Embodiment: (1) Road video dataset
[0119] The dataset involved in this embodiment can be divided into a pre-training dataset and a test dataset. Among them, the pre-training dataset Places365, as a subset of the large-scale image dataset Place, contains 1.8 million pictures, covering 365 different types of scenes, providing data support for the baseline CNN machine learning training of video scene classification and recognition applications. In terms of the audio dataset, AudioSet is the audio version of the ImageNet dataset provided by Google in 2017, consisting of 2 million 10-second voice segments, and a total of 527 audio label class indices, especially having good coverage in scene audio events. In addition, this embodiment conducts real-scene experiments on a self-built video surveillance test dataset. The samples of this dataset come from the public security video private network in Jiangsu Province, with a total of 1,951 samples, including 8 types of road scenes such as highways, urban roads, parking lots, tunnels, bus stops, bridges, intersections, and internal roads, as Figure 6 shown.
[0120] Among them, about 831 samples have been manually labeled, and some of the data samples are given the above scene labels, with a duration of 15 seconds each. In this embodiment, according to the splitting ratio of 6:2:2, the video set samples are sequentially selected as the training set (1,170), the validation set (380), and the test set (401). To reflect the non-uniform distortion in the real scene, in the experiments of this embodiment, no brightness adjustment is used to synthesize and expand the data samples. The test set is further divided according to the video acquisition time, that is, 248 video samples collected during the time periods of 20:00 - 24:00 and 00:00 - 4:00 are selected to generate a night test video set, and the remaining are 153 video sample sets for other time periods.
[0121] (2) Training details
[0122] In the backbone network DBN network structure, the number of visible layer units designed to receive visual modal information is 4,096, which is (f) consistent with the dimension of visual modal information V; the number of visible layer units designed to receive audio modal information is 1,920 (128×15), which is (a) consistent with the dimension of audio modal information V. In terms of the structural level, it is selected that the modal information passes through a hidden layer h with 1,024 nodes respectively (1)After data reconstruction, a combined representation hidden layer h of 2048 nodes is formed. (2) The combined representation features are then input into the softmax classifier after dimensionality reduction operations in sequence.
[0123] Regarding the training parameter settings of the collaborative network, the Adam gradient descent method is used for optimization. The batch size is 16, and the learning rate adopts a logarithmic decay strategy at intervals within the range of [10-2, 10-4]. A total of 5 rounds of network training are carried out. The training uses the GPU NVIDIA Tesla P40 to accelerate the network training.
[0124] (3) Evaluation criteria
[0125] In this embodiment, Precision (P) and Recall (R) are selected to evaluate the algorithm performance.
[0126] P = TP / (TP + FP) (17)
[0127] R = TP / (TP + FN) (18)
[0128] Among them, TP represents the number of video samples that can be correctly recognized in the current scenario, FP represents the number of video samples that are misrecognized in the current scenario, and FN represents the number of video samples recognized as the current scenario in other scenarios.
[0129] Next, ablation experiment analysis is carried out.
[0130] (1) Feature availability analysis
[0131] To verify the learning ability of the CNN-DBN model proposed in this embodiment on multi-modal data, the high-dimensional features of the samples output by the backbone classification network are unfolded in a low-dimensional manner in three-dimensional space, and the spatial visualization distribution is as Figure 7 shown:
[0132] Although the visualization in three-dimensional space is only an approximate projection of the high-dimensional feature space, it can be seen from the above figure that the 8 types of scenarios to be classified in the experiment show a clustering effect in the feature space. Therefore, the multi-modal posterior probability features obtained by the CNN-DBN network constructed in this embodiment using non-linear feature transformation can be used for scene classification.
[0133] (2) Ablation experiment on the modal quality influence factor
[0134] In response to poor illuminance, this embodiment proposes a modal information quality evaluation strategy. To verify the usability of this strategy, this embodiment completes the ablation of the modal evaluation function by controlling the modal information quality evaluation factors. The scene recognition model before and after ablation is used to test 248 samples of the night-time dataset respectively, and the scene recognition accuracy rate in low-illuminance video samples is statistically calculated. The obtained classification confusion matrix is as Figure 8 shown.
[0135] Figure 8 In the classification results of each scene monitoring video in [the figure], the row data is the reference label, and the column data is the predicted result. It can be seen from the comparison that the recognition accuracy rate drops significantly in scenes with drastic illuminance changes such as urban roads and intersections. Based on the statistical data of the above confusion matrix, the following conclusion can be drawn: Since the information quality evaluation strategy enhances the weight of audio modal information, the recognition accuracy rate of scene categories with strong audio feature specificity has been significantly improved. The recognition accuracy rate of the model after ablation is only 78.63%. After adding the modal quality impact factor, the scene recognition accuracy of the model under night-time illuminance conditions has been improved by 9.68%.
[0136] (3) Modal information masking experiment
[0137] To verify the complementarity between the two modal information, on the basis of the recognition model after modal evaluation ablation, visual feature masks and audio feature masks are respectively implemented on the extracted features to obtain the input of the backbone network under each single-modal feature condition. The column data of TD-all, TD-day, and TD-night respectively represent the scene recognition accuracy rates obtained on the entire test set (401), the daytime test set (153), and the night-time test set (248). The results are shown in Table 1. On the one hand, it can be seen from the data in the above table that compared with single-modal features, the scene classification accuracy rate of multi-modal features is the best, indicating that the correlation between heterogeneous information under multi-modal features is effective in improving road scene classification. On the other hand, the visual feature classification accuracy rate is slightly lower than that of multi-modal features and much higher than that of audio features, verifying the inference that visual features have strong modal information attributes in scene representation. The comparison of audio features in the two test sets is not obvious, verifying the hypothesis that audio features have strong robustness to illuminance. At the same time, the classification accuracy rate of visual features on the night-time test set also reaches a certain recognition accuracy, which is due to the association between modal features realized by the joint probability distribution established by the deep belief network DBN.
[0138] Table 1 Comparison of scene recognition accuracy rates under different modal features (%)
[0139]
[0140] The following is a performance comparison and analysis:
[0141] To further verify the recognition performance of the method proposed in this embodiment in complex road scenarios, this embodiment selects relevant methods in the field of scene recognition in recent years for comparison with the method of this embodiment, including: Weakly Supervised PatchNets (WSPN), Mixup-Based Multi-Channel Convolutional (MBMCC), Multi-Modal Enhancement Semantic Learning (MESL), Attention-Based Multimodal Fusion (ABMF), and Semantic Correspondence-Based Multitask Learning (SCBML). Table 2 shows the scene recognition accuracy of road monitoring data samples under different models.
[0142] Table 2 Comparison of recognition performance between CNN-DBN and related methods (%)
[0143]
[0144] It can be seen from the experimental results that for the dataset containing low-illumination night video samples, the CNN-DBN model of the method in this embodiment achieved the best experimental results. Although the ABMF method based on the attention mechanism utilizes motion information as another modal information, it is not robust for scene recognition of video samples under complex illuminations, which also verifies the effectiveness of the modal quality impact evaluation strategy in the method of this embodiment.
[0145] For the recognition application of road traffic monitoring videos, this embodiment further statistically analyzes the precision rate and recall rate of the proposed method and the Top2 (MESL, SCBML) methods with similar classification accuracies. The specific approach includes converting the output of the backbone network into a binary classification model, that is, classifying urban roads and intersections as positive examples, and classifying the remaining 6 types of scenes as negative examples. By adjusting the "positive example" threshold of the softmax classifier output, the Figure 9 shown Precision-Recall Curve (P-R curve) is obtained.
[0146] As can be seen from the above figure, compared with the baseline method: (1) the CNN-DBN model has the largest area under the P-R curve; (2) when P = R, the CNN-DBN model obtains the largest Break-Even Point (BEP) value. Therefore, according to the above two observations, it can be concluded that the method of this embodiment has the best performance in road scene classification. After analyzing the false positives and false negatives, it is found that although the ABMF model dynamically adjusts the contributions of different modal information to the classification output through the attention mechanism, due to its sensitivity to illumination changes, the recognition performance of some night scene samples is poor. At the same time, although the multi-modal semantic enhancement strategy introduced by the MESL model greatly improves the performance, it does not learn the image quality changes caused by illumination, so its robustness is not strong under complex illumination conditions.
[0147] Compared with the existing algorithms, the method and system proposed in the present invention have achieved the best recognition effect in the recognition of road scenes with brightness changes. Through the research on scene recognition technology, the automatic scene labeling of different surveillance videos is realized, and semantic constraint conditions are established for the subsequent tasks of marking detection.
[0148] The above description of the present invention with reference to the accompanying drawings is exemplary. Obviously, the specific implementation of the present invention is not limited by the above methods. As long as such non-substantial improvements are made by adopting the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A road scene recognition method based on modal information evaluation, characterized in that: Based on the extracted multi-modal features, a multi-task model for scene recognition based on modal information evaluation is constructed; the multi-modal features include audio features and video image frame features; the scene recognition multi-task model includes a collaborative network and a backbone classification network. The collaborative network uses VGG16-places365 for visual feature extraction as a modal quality evaluation network, and uses the BRISQUE algorithm based on the local normalized luminance coefficient as a teacher model to provide learning reference labels for the modal quality evaluation network; the backbone classification network uses a multi-modal DBN network for unsupervised joint representation of visual and audio modal scene information.
2. The road scene recognition method based on modal information evaluation according to claim 1, wherein: Based on the open-source VGG16-places365 and VGGish networks, high-dimensional feature extraction and dimensionality reduction are respectively performed on the image frames and audio in the video to form descriptive features for input to the backbone classification network.
3. The road scene recognition method based on modal information evaluation according to claim 2, wherein: The method for extracting audio features is as follows: sample the video clip, obtain the log Mel cepstral features after processing the sampled data, input them into the VGGish network model after re-framing, each audio clip obtains a 128-dimensional audio feature, and the audio feature output V is obtained after concatenation (a) 。 4. The road scene recognition method based on modal information evaluation according to claim 2, characterized in that: The method for extracting video image frame features includes: A. Video segment cutting: For a surveillance video sample under single-shot conditions, the surveillance video sample is cut into several video segments at specified time intervals. B. Frame-level feature extraction: Randomly extract any video image frame from the cut video segments. After network convolution and pooling operations, the output of the first fully connected layer is obtained and defined as the frame-level feature description matrix M. The row vector and column vector respectively represent the changes of the feature descriptors in the spatial dimension and the time dimension, and are expressed as follows: Among them, each row vector of the description matrix M is the frame-level feature of each video image frame, and the dimension of the frame-level feature is 4096; the frame-level feature descriptor of the i-th frame of the image frame in the j-th dimension is represented as F i,j ; n is the number of video segments in the current original video sample; C. Video-level feature extraction: The video-level features are obtained by using the method of mean clustering of the frame-level feature sequence. wherein Among them, V j (f) represents the j-th video-level feature.
5. The road scene recognition method based on modal information evaluation according to claim 1, characterized in that: In the collaborative network, the classification features come from the empirical distribution of the local normalized luminance under the spatial domain statistical model.
6. The road scene recognition method based on modal information evaluation according to claim 5, wherein: In the collaborative network, the last convolutional operation of the pre-trained model VGG16-places365 is used as the modal quality evaluation feature matrix f v , and the pooling vector η is further calculated as follows: η = vec(f v ) Among them, the vec operation expands the matrix into a one-dimensional vector; after spatial mapping and l2 normalization of η, the signed feature η′ and the evaluation feature η″ can be obtained respectively in sequence: Among them, sign is the operation of finding the signed square root, and e is the element-wise multiplication operation; the fused feature is input into the fully connected layer to calculate and generate the output y of the student network; a random video segment frame is taken as the input of the BRISQUE algorithm, and the score a of the teacher model is calculated; through normalization processing, the value range is ensured to be consistent with the predicted output y of the collaborative network; the empirical loss between the prediction and the BRISQUE output score is calculated using the l2 norm: Among them, N is the number of samples in each training batch; finally, the random gradient descent method is used to complete the fine-tuning of the collaborative learning network. Through the learning branch of the collaborative network, VGG16-places365 participates in multi-task learning, and the modal quality evaluation score δ is obtained through inference and prediction. After normalization, the value range of the modal quality evaluation score δ is [0,1]. The higher the score, the lower the illuminance of the video sample and the worse the quality of the visual modal information.
7. A road scene recognition method based on modal information evaluation according to claim 1, characterized in that: In the backbone classification network, the DBN network model uses a probabilistic generative model to establish a joint distribution between the observed data and the labels, which is expressed as follows: p(v,h 1 ,h 2 ,L,h l ) = p(v|h 1 )p(h 1 |h 2 )...p(h l-2 |h l-1 )p(h l-1 ,h l ) Among them, v represents the observed data, and h represents the label.
8. A road scene recognition method based on modal information evaluation according to claim 7, characterized in that: In the backbone classification network, the first layer RBM in the DBN network model uses a Gaussian-Bernoulli model, and its joint probability p and energy function E are specifically defined as follows: Among them, Z is the normalization factor obtained by approximate calculation of annealing importance sampling, σ is the variance of the Gaussian unit in the visible layer, b i represents the bias of the visible unit, D represents the number of visible units, F1 represents the number of hidden layer units, a j represents the bias of the hidden layer unit; The data types of the hidden layer nodes are all binary structures. The hidden layer RBM is defined as a Bernoulli-Bernoulli type, and the multi-modal joint distribution probability of the output is defined as follows: where θ represents the DBN model parameters, and h f (1) represents the feature extraction hidden layer, which is the hidden unit for capturing the internal features of a single modality, and h a (1) represents the association hidden layer, which is the hidden unit for modeling cross-modal interactions, and h (2) is the high-level abstract hidden layer for integrating the high-order representations of multi-modal information; The contrastive divergence CD algorithm is used in the learning and inference process to approximate the maximum likelihood, so that the hidden layer units can obtain the correlation of the high-order features of the two modalities. The solution process of the posterior probability of the model is as follows: Among them, h fj (1) represents the encoded unimodal features, and h aj (1) is used to capture cross-modal correlations, and h k (2) is used to integrate multi-modal information to form high-level representations; the learning process of the CD algorithm is the variational parameter μ that maximizes the variational lower bound under the condition that the DBN model parameter θ is fixed; similarly, when the variational parameter μ is fixed, the optimal solution of the network model parameter θ is found through the MCMC stochastic approximation method to complete the inference process.
9. The road scene recognition method based on modal information evaluation according to claim 8, characterized in that: In the backbone classification network, a Softmax classifier is cascaded at the end of the DBN network to complete the prediction output of the video scene. The network weights are supervised and optimized through a loss function that incorporates a modality quality evaluation factor δ. The relative entropy loss function is defined as: Among them, G is the true sample distribution, Q is the posterior probability output of the multi-modal DBN model, and ɑ is the modality weight balance factor; A regularization term is added to the relative entropy loss function to perform sparse constraints on the hidden layer nodes, and the final logarithmic likelihood cost function obtained is: Among them, K in the relative entropy regularization term is the number of hidden layer nodes, and T is the total number of training samples; γ is the weight factor of the regularization term, and the sparse coefficient ρ is used to control the sparsity of the hidden layer; The DBN model is continuously iteratively reduced by the gradient descent method until it finally converges to complete multi-task learning.
10. A road scene recognition system based on modal information evaluation, characterized in that: It includes a processor and a memory. When the computer program stored in the memory runs in the processor, it executes the road scene recognition method based on modality information evaluation according to any one of claims 1-9.
Citation Information
Cited By
Bicycle accounting evaluation system and method and electronic equipment
CN120746806A
Ceramic packaging substrate image super-resolution method based on mixed space interaction
CN121746185A
A Super-Resolution Method for Ceramic Packaging Substrates Based on Hybrid Spatial Interaction
CN121746185B
Multi-mode collaborative video quality evaluation network based on graph structure
CN121788524A
Road scene target detection method, system and equipment based on Transform and cross mode and medium
CN122435554A