A Multi-Model Pedestrian Recognition Method for Highway Videos Based on Convolutional Neural Networks

By adopting multi-environment compensation preprocessing and multi-model feature fusion technology in the pedestrian recognition system, combined with adaptive feature weighting and multi-level feature analysis, the problem of low recognition accuracy in complex environments is solved, and higher recognition accuracy and robustness are achieved.

CN119785300BActive Publication Date: 2025-05-30HANGZHOU HUIJING TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510280129.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-30
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The accuracy of the existing pedestrian identification method in complex and changing highway environments has significantly decreased, and it is difficult to adapt to environmental interference and space-time consistency analysis, resulting in frequent false alarms and missed detection phenomena.

Method used

A multi-model pedestrian recognition method based on convolutional neural network is adopted to obtain video stream data through a surveillance camera and perform multi-environment compensation preprocessing, including adaptive histogram equalization, raindrop detection and elimination, and image defogging processing. Then, the preprocessed features are input into three convolutional neural network models in parallel, and spatial, motion and context-related features are extracted respectively, and multi-model feature fusion and adaptive feature weighting are performed through the dynamic feature fusion module. Finally, multi-level feature analysis and spatiotemporal consistency verification are performed through the cascade classifier and time series verification module.

Benefits of technology

It significantly improves the accuracy and robustness of pedestrian recognition in complex environments, enhances image quality, eliminates rain and fog interference, adapts to light changes, reduces false alarms and missed detection, and meets the practical application needs of highway video surveillance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785300B_ABST
    Figure CN119785300B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-model pedestrian recognition method for highway videos based on a convolutional neural network, which includes obtaining a video stream through a monitoring camera and intercepting consecutive frame images; performing multi-environment compensation preprocessing on the images to generate a group of standardized feature maps; inputting the group of feature maps into three parallel convolutional neural network models to respectively extract spatial, motion, and context feature matrices; fusing the feature matrices through a dynamic feature fusion module to obtain a comprehensive feature tensor; adaptively weighting the comprehensive feature tensor based on the real-time environmental parameters output by an environmental perception module to generate an environment compensation feature tensor; using a cascade classifier to perform multi-level analysis on the feature tensor to output an initial recognition result; and finally, performing spatio-temporal consistency verification on the recognition results of consecutive frames through a time series verification module to generate a final recognition result. The present invention can improve the accuracy and robustness of pedestrian recognition in highway videos under complex environments and enhance the reliability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and intelligent transportation. More specifically, the present invention relates to a multi-model pedestrian recognition method for highway videos based on a convolutional neural network. Background Art

[0002] In intelligent transportation systems, highway video surveillance technology has been widely applied, mainly for traffic flow monitoring, accident warning, and pedestrian detection, etc. Existing pedestrian recognition methods are mostly based on a single model and rely on traditional image processing techniques or simple neural network architectures. These methods can achieve certain results in ideal environments, but in the complex and changeable highway environment, such as rainy and foggy weather, light changes, or background interference, the recognition accuracy drops significantly. In addition, traditional methods often ignore the spatio-temporal continuity and environmental adaptability of pedestrian features, resulting in false alarms or missed detections in dynamic scenes.

[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: a single model is difficult to cope with complex and changeable environmental conditions, lacks spatio-temporal consistency analysis of dynamic scenes, and has insufficient adaptability to environmental interference, resulting in the accuracy and robustness of pedestrian recognition being difficult to meet the actual application requirements. Summary of the Invention

[0004] The present invention provides a multi-model pedestrian recognition method for highway videos based on a convolutional neural network, including: S1. Real-time acquiring highway video stream data through a monitoring camera, and intercepting continuous frame images to form an input sequence, where .

[0005] S2. Performing multi-environment compensation preprocessing on the input sequence to generate a group of standardized feature maps.

[0006] S3. Inputting the group of standardized feature maps into three convolutional neural network models set in parallel to respectively obtain a spatial feature matrix , a motion feature matrix , and a context correlation matrix .

[0007] S4. Performing multi-model feature fusion on the three feature matrices through a dynamic feature fusion module, and calculating to obtain a comprehensive feature tensor , whose dimension is .

[0008] S5. Based on the real-time environmental parameters output by the environmental perception module, performing adaptive feature weighting on the comprehensive feature tensor to generate an environment compensation feature tensor .

[0009] S6. Use a cascade classifier to perform multi-level feature analysis and output an initial recognition result set .

[0010] S7. Use a time series verification module to perform spatio-temporal consistency verification on consecutive frames of and generate a final recognition result , where .

[0011] Furthermore, the specific steps of S2 include: S21. Process each frame of image using the adaptive histogram equalization algorithm, and its enhancement function is: ; where is the number of gray levels, is the occurrence probability of the -th gray level, is the pixel coordinate.

[0012] S22. Identify and eliminate rain streak interference through a rain drop detection network, and the network structure includes 5 residual blocks and 3 attention modules.

[0013] S23. Perform image dehazing processing based on the dark channel prior, and the restoration model is: ; where is the observed image, is the restored image, is the atmospheric light value, is the transmittance.

[0014] Furthermore, the dynamic feature fusion module in S4 performs the following operations: ; where represents the channel dimension concatenation operation, is the dynamic weight coefficient, and its calculation formula is: ; in the formula is the Sigmoid function, is the learnable parameter matrix, is the bias term.

[0015] Furthermore, the adaptive feature weighting process in S5 includes: ; where represents element-wise multiplication, is the environmental compensation matrix, and its calculation formula is: ; in the formula is the real-time rainfall intensity coefficient, is the fog concentration coefficient, is the light intensity coefficient, is the compensation coefficient.

[0016] Further, the cascade classifier of S6 includes: S61. The first-level classifier uses a depthwise separable convolutional layer to extract local features and output a preliminary classification confidence .

[0017] S62. The second-level classifier expands the receptive field through an atrous convolutional layer and calculates the global correlation confidence .

[0018] S63. The final classification decision function is: ; where is the temporal correlation confidence, is a fixed weight parameter.

[0019] Further, the calculation method of the temporal correlation confidence is: ; where is the temporal consistency function, defined as follows: ; in the formula is the target center coordinate, is the scale parameter.

[0020] Further, the time series verification module of S7 executes: ; where, is the verification score of the target in the th frame, is the verification threshold, and the verification score calculation formula is: ; in the formula is the detection confidence, is the trajectory continuity score, is the appearance similarity, .

[0021] Further, it also includes: S8. Executing false alarm filtering on the targets in , and the filtering condition is: ; where is the height of the video frame, and when the target size satisfies both conditions, it is determined as a false alarm.

[0022] Further, the environment perception module calculates the environment parameters through the following formula: rainfall intensity coefficient: ; fog concentration coefficient: ; in the formula is the gradient magnitude map, is the historical average gradient value, is the mean value of the red channel of the current frame, is the maximum reference value.

[0023] Furthermore, the three convolutional neural network models adopt a differentiated structure: the spatial feature model adopts the ResNet50 architecture; the motion feature model includes 3D convolutional layers and LSTM layers; the context model adopts the Nonlocal network structure, and its correlation calculation formula is: ; where is a similarity function, is a feature transformation function, is a normalization factor.

[0024] According to the above embodiments of the present invention, it has at least the following beneficial effects: Through multi-model feature fusion and dynamic feature weighting, the present invention can effectively improve the accuracy and robustness of pedestrian recognition in complex environments. The multi-environment compensation preprocessing can enhance the image quality, eliminate rain and fog interference, and adapt to light changes, thereby providing high-quality input data for subsequent recognition. The dynamic feature fusion module combines spatial, motion, and context correlation features, which can make full use of information in different dimensions and further improve the richness and distinctiveness of feature expression. The adaptive feature weighting adjusts the feature tensor according to real-time environmental parameters, enhances the model's adaptability to environmental changes, and ensures stable recognition performance under different weather and lighting conditions.

[0025] In addition, the present invention adopts a cascade classifier and a time series verification module, which can perform multi-level analysis and spatio-temporal consistency verification on pedestrian features. The cascade classifier can accurately distinguish pedestrians from non-pedestrian targets through multi-level feature extraction and confidence fusion, reducing false alarms and missed detections. The time series verification module further verifies the recognition results of consecutive frames to ensure the continuity and appearance consistency of the target trajectory, thereby further improving the reliability and accuracy of pedestrian recognition. Finally, through the false alarm filtering mechanism, misdetected targets that do not conform to the pedestrian size characteristics can be eliminated, further optimizing the recognition results and meeting the actual application requirements of highway video surveillance. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understandable. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner.

[0027] Among them: Figure 1 is a schematic flowchart of a multi-model pedestrian recognition method for highway video based on a convolutional neural network provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are given only to enable those skilled in the art to better understand and then implement the present invention, rather than limiting the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to be able to fully convey the scope of the present invention to those skilled in the art.

[0029] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present invention can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0030] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0031] The following reference Figure 1 , Figure 1 is a schematic flowchart of a multi-model pedestrian recognition method for highway videos based on a convolutional neural network provided for an embodiment of the present invention. As Figure 1 shown, a multi-model pedestrian recognition method 100 for highway videos based on a convolutional neural network includes: S1. Real-time acquisition of highway video stream data through a monitoring camera, and intercepting continuous frames of images to form an input sequence, where .

[0032] S2. Perform multi-environment compensation preprocessing on the input sequence to generate a group of normalized feature maps.

[0033] S3. Input the group of normalized feature maps into three convolutional neural network models set in parallel, and respectively obtain a spatial feature matrix , a motion feature matrix , and a context correlation matrix .

[0034] S4. Perform multi-model feature fusion on the three feature matrices through a dynamic feature fusion module, and calculate to obtain a comprehensive feature tensor , whose dimension is .

[0035] S5. Based on the real-time environmental parameters output by the environmental perception module, perform adaptive feature weighting on the comprehensive feature tensor to generate an environment compensation feature tensor .

[0036] S6. Use a cascade classifier to perform multi-level feature analysis on and output an initial recognition result set 。

[0037] S7. Perform spatio-temporal consistency verification on consecutive frames through the time series verification module to generate the final recognition result , where 。

[0038] It should be noted that the present invention obtains highway video stream data in real time through a monitoring camera, and intercepts consecutive frame images to form an input sequence. The monitoring camera refers to a video monitoring device installed on the highway for real-time capturing of road conditions. The input sequence refers to consecutive image frames intercepted from the video stream, which are usually used for subsequent image processing and analysis, and the number of frames is not less than 5 frames.

[0039] Specifically, perform multi-environment compensation preprocessing on the input sequence to generate a set of standardized feature maps. The multi-environment compensation preprocessing includes adaptive histogram equalization, raindrop detection and elimination, and image defogging processing based on dark channel prior. Adaptive histogram equalization is used to enhance image contrast. Raindrop detection and elimination identify and remove rain streak interference through residual blocks and attention modules. Image defogging processing removes the haze effect through a restoration model. The set of standardized feature maps refers to a set of image data with unified features generated after preprocessing.

[0040] Preferably, input the set of standardized feature maps into three parallel convolutional neural network models to obtain a spatial feature matrix, a motion feature matrix, and a context correlation matrix respectively. The spatial feature model adopts the ResNet50 architecture. The motion feature model includes 3D convolutional layers and LSTM layers. The context model adopts the Nonlocal network structure. Perform multi-model feature fusion on the three feature matrices through a dynamic feature fusion module to calculate a comprehensive feature tensor. The dynamic feature fusion module realizes the fusion of different features through channel dimension concatenation operations and dynamic weight coefficient calculations. Based on the real-time environmental parameters output by the environmental perception module, perform adaptive feature weighting on the comprehensive feature tensor to generate an environment compensation feature tensor. The environmental perception module adjusts the weight of the feature tensor by calculating the rainfall intensity coefficient, fog concentration coefficient, and light intensity coefficient. Use a cascade classifier to perform multi-level feature analysis on the environment compensation feature tensor and output an initial recognition result set. The cascade classifier extracts features and calculates classification confidence through 3×3 depthwise separable convolutional layers and dilated convolutional layers. Perform spatio-temporal consistency verification on the initial recognition results of consecutive frames through the time series verification module to generate the final recognition result. The time series verification module ensures the accuracy and reliability of the recognition results through verification score calculation and target size filtering.

[0041] In some embodiments, S2 specifically includes: S21. Processing each frame of image by using the adaptive histogram equalization algorithm, and its enhancement function is: ; where is the number of gray levels, is the occurrence probability of the -th gray level, are pixel coordinates.

[0042] S22. Identifying and eliminating rain streak interference through a rain drop detection network, and the network structure includes 5 residual blocks and 3 attention modules.

[0043] S23. Performing image defogging processing based on the dark channel prior, and the restoration model is: ; where is the observed image, is the restored image, is the atmospheric light value, is the transmittance.

[0044] It should be noted that the multi-environment compensation preprocessing mentioned in the present invention is one of the key steps to achieve accurate pedestrian recognition. Specifically, this step includes adaptive histogram equalization, rain drop detection and elimination, and image defogging processing. Adaptive histogram equalization is an image enhancement technology that adjusts the gray distribution of an image, enhances the contrast of the image, and makes it more suitable for subsequent feature extraction and analysis. The rain drop detection network is a deep learning model used to identify and eliminate rain streak interference in an image, thereby reducing the impact on pedestrian recognition in rainy environments. Image defogging processing is based on the dark channel prior theory, and removes the fog component in the image through the restoration model to improve the clarity and visibility of the image.

[0045] Specifically, the enhancement function of the adaptive histogram equalization algorithm is: where L is the number of gray levels, is the -th gray level occurrence probability, and are pixel coordinates. The structure of the rain drop detection network includes 5 residual blocks and 3 attention modules. The residual blocks are used to learn the residual features in the image, thereby improving the network's ability to detect rain drops; the attention modules are used to enhance the network's attention to the rain drop area and improve the elimination effect. The restoration model of the image defogging processing is: ; where, is the observed image, is the restored image, is the atmospheric light value, is the transmittance, and this model restores the clarity of the image by estimating the atmospheric light value and the transmittance.

[0046] Preferably, in order to further optimize the effect of multi-environment compensation preprocessing, relevant parameters can be adjusted. For example, in adaptive histogram equalization, the number of gray levels L can be adjusted according to the gray distribution of the actual image to better meet the image enhancement requirements under different lighting conditions. For the raindrop detection network, the detection ability for complex raindrop patterns can be further improved by increasing the number of residual blocks or adjusting the weights of the attention module. In image dehazing, more prior knowledge can be introduced, such as the prior information of the sky region, to more accurately estimate the atmospheric light value and transmittance. In addition, other dehazing algorithms can also be combined, such as the dehazing method based on physical models, combined with the dark channel prior method, to further enhance the dehazing effect.

[0047] In some embodiments, the dynamic feature fusion module in S4 performs the following operations: ; where represents the channel dimension concatenation operation, is the dynamic weight coefficient, and the calculation formula is: ; in the formula is the Sigmoid function, is the learnable parameter matrix, is the bias term.

[0048] It should be noted that the dynamic feature fusion module is a key part of the present invention for fusing the feature matrices extracted by different convolutional neural network models. Through a specific fusion strategy, this module integrates the spatial feature matrix, the motion feature matrix, and the context correlation matrix to generate a comprehensive feature tensor. This process not only considers the complementarity of different feature matrices but also enables the fusion process to adaptively adjust the weight distribution according to the characteristics of the input features through the calculation of dynamic weight coefficients. Among them, the calculation of the dynamic weight coefficient is based on the learnable parameter matrix and the bias term, and is normalized through the Sigmoid function to ensure the rationality of the weights.

[0049] Specifically, the operation formula of the dynamic feature fusion module is: ; where, represents the channel dimension concatenation operation, and are the dynamic weight coefficients. The calculation formulas of these weight coefficients are: ; where, is the Sigmoid function used to normalize the weight value to the interval (0, 1); and are the learnable parameter matrices, and are the bias terms. These parameters are optimized through backpropagation during the training process to adapt to different input features and scenarios.

[0050] Preferably, in order to further optimize the performance of the dynamic feature fusion module, it can be refined or replaced in the following aspects. First, a more complex weight calculation mechanism can be introduced, such as introducing a multi-layer perceptron (MLP) to replace the linear transformation, so as to better capture the complex relationships between feature matrices. Second, an attention mechanism can be considered to enable the model to focus more on important features, thereby improving the fusion effect. In addition, for the channel dimension concatenation operation, other fusion strategies can be tried, such as weighted summation or feature crossing, to explore a better fusion method. Finally, during the training process, overfitting can be prevented through regularization techniques (such as Dropout or L2 regularization) to further improve the generalization ability of the model.

[0051] In some embodiments, the adaptive feature weighting process in S5 includes: ; where represents element-wise multiplication, is the environmental compensation matrix, and its calculation formula is: ; in the formula is the real-time rainfall intensity coefficient, is the fog concentration coefficient, is the light intensity coefficient, is the compensation coefficient.

[0052] It should be noted that the adaptive feature weighting process is a key link in the present invention for dynamically adjusting the weights of feature tensors according to environmental parameters. This process performs element-wise weighting on the comprehensive feature tensor through the environmental compensation matrix, enhancing the adaptability of the model to different environmental conditions. Among them, the calculation of the environmental compensation matrix is based on the real-time rainfall intensity coefficient, fog concentration coefficient, and light intensity coefficient, which reflect the impact degree of the current environment on the image quality. By multiplying and adding with the compensation coefficient, an environmental compensation matrix for weighting is generated.

[0053] Specifically, the formula for adaptive feature weighting is: ; where, represents the element-wise multiplication operation, is the environmental compensation matrix, and the calculation formula is: ; where, is the real-time rainfall intensity coefficient, is the fog concentration coefficient, is the light intensity coefficient. The compensation coefficients and are set according to experimental experience to balance the influence of different environmental factors on the feature tensor.

[0054] Preferably, to further optimize the effect of adaptive feature weighting, refinement or substitution can be carried out in the following aspects. First, a dynamic adjustment mechanism can be introduced to adaptively adjust the compensation coefficient according to the changes in real-time environmental parameters and , rather than a fixed value, so as to better adapt to extreme environmental conditions. Second, a non-linear function (such as ReLU or Tanh) can be considered to replace the linear weighting to enhance the expressive ability of the environmental compensation matrix. In addition, for the acquisition of environmental parameters, multiple sensor data (such as weather station data or the output of the environmental perception module) can be combined to improve the accuracy and reliability of the environmental parameters. Finally, the optimal combination of compensation coefficients under different environmental conditions can be verified through experiments and used as the hyperparameters of the model for optimization to further improve the robustness of the model in complex environments.

[0055] In some embodiments, the cascaded classifier of S6 includes: S61. The first-level classifier uses a depthwise separable convolutional layer to extract local features and output a preliminary classification confidence .

[0056] S62. The second-level classifier expands the receptive field through a dilated convolutional layer and calculates the global correlation confidence .

[0057] S63. The final classification decision function is: ; where is the temporal correlation confidence, is a fixed weight parameter.

[0058] It should be noted that the cascaded classifier is a key part for multi-level feature analysis and pedestrian recognition in the present invention. It gradually extracts features through multiple convolutional layers and combines the confidences at different levels for classification decisions, thereby improving the accuracy and reliability of recognition. Among them, the first-level classifier uses a depthwise separable convolutional layer to extract local features, the second-level classifier expands the receptive field through a dilated convolutional layer to calculate the global correlation confidence, and the final classification decision function combines the local features, global correlation, and temporal correlation confidence to output the pedestrian recognition result.

[0059] Specifically, the first-level classifier uses a 3×3 depthwise separable convolutional layer. This convolutional layer separates the standard convolution operation into two parts: channel-wise convolution and point-wise convolution, which can reduce the computational amount while maintaining the effectiveness of feature extraction. Its output is the preliminary classification confidence , which is used to evaluate the local feature similarity of whether the target is a pedestrian. The second-level classifier uses a dilated convolutional layer. By introducing a dilation rate, the receptive field of the convolutional kernel is expanded, so as to capture more extensive context information and calculate the global correlation confidence 。The final classification decision function is as follows: ; where is the temporal correlation confidence, which is used to evaluate the continuity of the target in the time series. The fixed weight parameters and are set according to the experimental results and are used to balance the contributions of different confidences to the classification decision.

[0060] Preferably, to further optimize the performance of the cascade classifier, refinement or substitution can be carried out in the following aspects. First, more convolutional layers or different types of convolutional operations (such as deformable convolutions) can be introduced to enhance the feature extraction ability. Second, the weight parameters and can be dynamically adjusted to adaptively change according to the quality of the input features and environmental conditions instead of fixed values.

[0061] Furthermore, an attention mechanism can be introduced to enable the classifier to focus more on the key regions of pedestrian features, thereby improving the recognition accuracy. Finally, for the calculation of the temporal correlation confidence, a more complex temporal model (such as Transformer) can be considered to better capture the dynamic changes of the target in the time series, further enhancing the accuracy and robustness of pedestrian recognition.

[0062] In some embodiments, the calculation method of the temporal correlation confidence is as follows: ; where is the temporal consistency function, which is defined as follows: ; in the formula is the center coordinate of the target, is the scale parameter.

[0063] It should be noted that the calculation of the temporal correlation confidence is an important part of the present invention for evaluating the consistency and continuity of pedestrian targets in the time series. It calculates the similarity score of the target in consecutive frames to ensure the spatio-temporal consistency of the recognition result, thereby improving the accuracy and reliability of pedestrian recognition. The calculation of the temporal correlation confidence is based on the temporal consistency function, which evaluates the stability of the target in the time series by measuring the distance between the center coordinates of the target.

[0064] Specifically, the calculation formula of the temporal correlation confidence is: ; where N is the number of frames of the target in the time series, is the temporal consistency function, which is used to measure the similarity of the target in two consecutive frames. The temporal consistency function is defined as: ; where and are the center coordinates of the target in the i-th frame and the (i + 1)-th frame respectively, is a scale parameter used to control the sensitivity of distance. In the present invention, is set to 5. This parameter value is obtained through experimental verification, aiming to balance the distance difference between the target center coordinates and the stability of the similarity score.

[0065] Preferably, in order to further optimize the calculation effect of the temporal correlation confidence, refinement or substitution can be carried out in the following aspects. First, more feature information (such as the appearance features or motion trajectories of the target) can be introduced to enhance the expression ability of the temporal consistency function, rather than relying solely on the target center coordinates. Second, the scale parameter can be dynamically adjusted so that it adaptively changes according to the motion speed of the target or the scene complexity, rather than a fixed value. In addition, more complex temporal models (such as recurrent neural networks or Transformer architectures) can be considered to replace the temporal consistency function to better capture the dynamic changes of the target in the time series. Finally, regularization techniques (such as Dropout) can be introduced to prevent overfitting and further improve the robustness of the model in complex scenarios.

[0066] In some embodiments, the time series verification module of S7 performs: ; where is the verification score of the target in the th frame, is the verification threshold, and the verification score calculation formula is: ; in the formula is the detection confidence, is the trajectory continuity score, is the appearance similarity, .

[0067] It should be noted that the final confirmation of the pedestrian detection result is to be carried out here, using the previously obtained spatial correlation confidence , feature correlation confidence and temporal correlation confidence , and the comprehensive correlation confidence P is calculated through the formula . Among them, is the corresponding weighting coefficient, used to adjust the importance of different correlation confidences in the comprehensive result. If P is greater than the set threshold T, the detected target is confirmed as a pedestrian.

[0068] Specifically, the weighting coefficient satisfies , and , this set of coefficients is obtained through a large number of experiments and can better balance the influence of the correlation information in the three aspects of space, features, and time series on the final result. The threshold T is set to 0.7, which is an empirical value and is used as the critical value for judging whether the target is a pedestrian. The spatial correlation confidence measures the degree of association of the target in terms of spatial position, and the feature correlation confidence reflects the similarity of the target features, and the temporal correlation confidence reflects the consistency of the target in the time series.

[0069] Preferably, in practical applications, the weighting coefficients can be dynamically adjusted according to different scenarios. For example, in a monitoring scenario, if the movement speed of the target is relatively fast, then the value of can be appropriately increased to pay more attention to the temporal consistency of the target; if the feature differences of the targets in the scenario are relatively large, then the value of is increased. For the threshold T, an adaptive adjustment method can also be adopted. According to the accuracy of the detection results within a certain period of time, a machine learning algorithm is used to dynamically adjust the threshold to improve the detection accuracy. In addition, when calculating the comprehensive correlation confidence P, some penalty terms can be added. For example, when a certain correlation confidence value is too low, additional penalties are imposed on it to further improve the reliability of the detection.

[0070] In some embodiments, it further includes: S8. Perform false alarm filtering on the targets, and the filtering condition is: ; where is the height of the video frame, and when the target size satisfies both conditions, it is determined as a false alarm.

[0071] It should be noted that here, feature fusion is performed on the multi-modal information in the multi-modal pedestrian detection system. The multi-modal information used includes visual image information and millimeter-wave radar information. The feature fusion operation is to integrate these two different types of information to obtain a unified feature representation. The feature fusion method adopted is bilinear pooling, and its purpose is to enhance the feature expression ability in this way so that pedestrians can be detected more accurately in the follow-up.

[0072] Specifically, the visual image information is the scene picture information collected by the camera, which contains rich visual features such as the appearance and posture of pedestrians. The millimeter-wave radar information is the information of the distance, speed, angle, etc. of the target obtained after the millimeter-wave radar transmits and receives millimeter-wave signals. When the bilinear pooling method performs feature fusion, it will first extract feature vectors from the visual image information and the millimeter-wave radar information respectively, then perform an outer product operation on these two feature vectors, and finally perform some normalization operations to obtain the fused features. In this process, L2 normalization can be used for normalization, and its function is to normalize the norm of the feature vector to 1 to avoid the influence of too large or too small feature values on subsequent detections.

[0073] Preferably, when extracting the feature vectors of the visual image information and the millimeter-wave radar information, different deep learning network structures can be used. For the visual image information, the ResNet network can be selected, which has a very deep network structure and can extract more advanced semantic features. For the millimeter-wave radar information, a convolutional neural network specifically designed for radar data can be used to better adapt to the characteristics of radar data. When performing the outer product operation, in order to reduce the amount of calculation, a low-rank approximation method can be used, such as projecting the feature vector into a low-dimensional subspace and then performing the outer product operation. In addition, in addition to L2 normalization, the Batch Normalization method can also be tried, which can make the feature distribution more stable during the training process and accelerate the convergence speed of the model.

[0074] In some embodiments, the environment perception module calculates the environmental parameters through the following formula: rainfall intensity coefficient: ; fog concentration coefficient: ; where is the gradient magnitude map, is the historical average gradient value, is the mean value of the red channel of the current frame, is the maximum reference value.

[0075] It should be noted that the process of pedestrian detection and classification using the fused multi-modal features in the multi-modal pedestrian detection system is described here. The support vector machine (SVM) classifier is used, which will judge whether the target is a pedestrian based on the fused features obtained through bilinear pooling before. SVM is a supervised machine learning algorithm, and its core is to find an optimal hyperplane to separate samples of different classes, which is to distinguish pedestrians from non-pedestrians here.

[0076] Specifically, the fused multi-modal features are the unified feature representations obtained after bilinear pooling operations, which integrate the features of visual image information and millimeter-wave radar information. When training a support vector machine classifier, a large amount of sample data with labels (pedestrians or non-pedestrians) is required. The radial basis function (RBF) is selected as the kernel function, which can map the data in the low-dimensional space to the high-dimensional space, making it easier to find the hyperplane that separates different categories of samples. For the radial basis function, there is an important parameter \(\gamma\), which is set to 0.1 here. It controls the width of the kernel function and affects the generalization ability and classification accuracy of the classifier.

[0077] Preferably, when selecting the kernel function, in addition to the radial basis function, the polynomial kernel function can also be tried. The form of the polynomial kernel function is ; where \(r\) and \(d\) are parameters that need to be adjusted. When the data has a certain polynomial distribution law in the low-dimensional space, the polynomial kernel function may achieve better classification results. When training a support vector machine classifier, the cross-validation method can be used to select the optimal parameters. For example, the training data is divided into several parts, and one part is used as the validation set in turn, and the rest are used as the training set. By comparing the classification accuracies on the validation set under different parameter combinations, the optimal parameter settings are found. In addition, dimensionality reduction processing can also be performed on the fused multi-modal features, such as using the principal component analysis (PCA) method to reduce the dimensionality of the features, reduce the computational amount, and may also improve the performance of the classifier.

[0078] In some embodiments, the three convolutional neural network models adopt a differentiated structure:

[0079] The spatial feature model adopts the ResNet50 architecture; the motion feature model includes 3D convolutional layers and LSTM layers; the context model adopts the Nonlocal network structure, and its association calculation formula is: ; where is the similarity function, is the feature transformation function, is the normalization factor.

[0080] It should be noted that this step is to evaluate the multi-modal pedestrian detection system, and the evaluation metric used is the mean average precision (mAP). The mean average precision is an important metric for measuring the performance of object detection algorithms, which comprehensively considers the accuracy and recall rate of the detection results. By calculating this metric, a quantitative evaluation of the performance of the entire multi-modal pedestrian detection system can be obtained, so as to judge the advantages and disadvantages of the system and make improvements.

[0081] Specifically, when calculating the mean average precision (mAP), the average precision (AP) for each class needs to be calculated first. For pedestrian detection, the class is the pedestrian. The calculation of the average precision is based on the precision-recall curve (P-R curve), which describes the change of precision at different recall rates. Integrating the area under the P-R curve gives the average precision (AP). The mean average precision (mAP) is the average of the average precisions of all classes. Although there is only one class, the pedestrian, here, the calculation method principle is the same. During the calculation process, the detection results need to be compared with the ground truth annotations to determine whether the detected pedestrians are correct, so as to calculate the precision and recall.

[0082] Preferably, when calculating the P-R curve, interpolation can be used to smooth the curve, making the calculation of the average precision more accurate. For example, for each recall value, take the maximum value of all subsequent precisions as the precision at this point, which can avoid the influence of curve fluctuations on the calculation of the average precision. In addition, in addition to the mean average precision (mAP), other evaluation metrics can be introduced, such as the F1 score. The F1 score is the harmonic mean of precision and recall, and the formula is: 。

[0083] It can more comprehensively reflect the balance between the accuracy and recall rate of the detection system. At the same time, when comparing the detection results with the ground truth annotations, different IoU (Intersection over Union) thresholds can be set to calculate the mAP under different thresholds, so as to more carefully evaluate the performance of the system under different matching accuracy requirements.

[0084] The above embodiments of the present invention have the following beneficial effects: Through multi-model feature fusion and dynamic feature weighting, the present invention can significantly improve the accuracy and robustness of pedestrian recognition in complex environments. Multi-environment compensation preprocessing can enhance the image quality, eliminate rain and fog interference, and adapt to light changes, thus providing high-quality input data for subsequent recognition. The dynamic feature fusion module combines spatial, motion, and context-related features, which can make full use of information in different dimensions and further improve the richness and distinctiveness of feature expression. Adaptive feature weighting adjusts the feature tensor according to real-time environmental parameters, enhancing the model's adaptability to environmental changes and ensuring stable recognition performance under different weather and lighting conditions.

[0085] In addition, the present invention adopts a cascade classifier and a time series verification module, which can perform multi-level analysis and spatio-temporal consistency verification on pedestrian features. The cascade classifier can accurately distinguish pedestrians from non-pedestrian targets through multi-level feature extraction and confidence fusion, reducing false alarms and missed detections. The time series verification module further verifies the recognition results of consecutive frames to ensure the continuity and appearance consistency of the target trajectory, thereby further improving the reliability and accuracy of pedestrian recognition. Finally, through the false alarm filtering mechanism, the false detection targets that do not conform to the pedestrian size characteristics can be eliminated, further optimizing the recognition results and meeting the actual application requirements of highway video surveillance.

[0086] Furthermore, the storage medium of the embodiment of the present application stores program instructions capable of implementing all the above methods. Among them, the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0087] The above description is only some preferred embodiments of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.

Claims

1. A multi-model pedestrian recognition method for highway videos based on convolutional neural networks, characterized in that The following steps are involved: S1. Obtain highway video stream data in real time through surveillance cameras and capture continuous The frame images constitute the input sequence, where ; S2. Perform multi-environment compensation preprocessing on the input sequence to generate a standardized feature map group; S3. Input the standardized feature map group into three parallel convolutional neural network models to obtain spatial feature matrices respectively , motion feature matrix , context association matrix ; S4. The dynamic feature fusion module performs multi-model feature fusion on the three feature matrices and calculates the comprehensive feature tensor , whose dimensions are ; S5. Based on the real-time environmental parameters output by the environmental perception module, perform adaptive feature weighting on the comprehensive feature tensor to generate an environmental compensation feature tensor ; S6. Using cascade classifier Perform multi-level feature analysis and output the initial recognition result set ; S7. Through the time series verification module, continuous Frame Perform spatiotemporal consistency verification to generate the final recognition result ,in ; The dynamic feature fusion module in S4 performs the following operations: ;in, represents the channel dimension concatenation operation, is the dynamic weight coefficient, and the calculation formula is: ; In the formula is the Sigmoid function, is the learnable parameter matrix, is the bias term; The adaptive feature weighting process in S5 includes: ;in represents element-wise multiplication, is the environmental compensation matrix, and its calculation formula is: ; In the formula is the real-time rainfall intensity coefficient, is the fog concentration coefficient, is the light intensity coefficient, is the compensation coefficient; The cascade classifier of S6 comprises: S61. The first level classifier adopts The depth-wise separable convolutional layer extracts local features and outputs preliminary classification confidence ; S62. The second-level classifier expands the receptive field through the hole convolution layer and calculates the global association confidence ; S63. The final classification decision function is: ;in is the confidence of temporal association, is a fixed weight parameter.

2. The method according to claim 1, characterized in that The S2 specifically includes: S21. Using an adaptive histogram equalization algorithm to process each frame image, the enhancement function is: ;in, is the gray level, For the The probability of grayscale occurrence, is the pixel coordinate; S22. Identify and eliminate rain streak interference through the raindrop detection network. The network structure includes 5 residual blocks and 3 attention modules; S23. Perform image defogging based on dark channel prior. The restoration model is: ;in To observe the image, To restore the image, is the atmospheric light value, is the transmittance.

3. The method according to claim 1, characterized in that The timing correlation confidence The calculation method is: ; in is the temporal consistency function, defined as follows: ; In the formula is the target center coordinate, is the scale parameter.

4. The method according to claim 1, characterized in that The S7 time series validation module performs: ;in For the goal In the The verification score of the frame, is the verification threshold, and the verification score calculation formula is: ; In the formula To test the confidence level, is the trajectory continuity score, is the appearance similarity, .

5. The method according to claim 1, characterized in that Also includes: S8.Yes Targets in the filter are subject to false positive filtering, and the filtering conditions are: ;in is the video frame height. When the target size meets both conditions, it is considered a false positive.

6. The method according to claim 1, characterized in that The environmental perception module calculates environmental parameters through the following formula: Rainfall intensity coefficient: ; Fog concentration coefficient: ; In the formula is the gradient magnitude map, is the historical average gradient value, is the mean value of the red channel of the current frame, is the maximum reference value.

7. The method according to claim 1, characterized in that The three convolutional neural network models adopt differentiated structures: the spatial feature model adopts the ResNet50 architecture; The motion feature model includes a 3D convolution layer and an LSTM layer; the context model uses a Nonlocal network structure, and its associated calculation formula is: ;in is the similarity function, is the feature transformation function, is the normalization factor.

Citation Information

Patent Citations

  • Human body behavior recognition method and device, electronic equipment and storage medium

    CN112418032A

  • Building deformation analysis method and system based on three-dimensional laser scanning and deep learning

    CN119197368A