Vehicle-to-person detection method based on YOLOv4-Tiny improved by domain adaptive algorithm of class alignment
By improving the YOLOv4-Tiny algorithm and combining the KCF tracking algorithm, the existing vehicle detection method has solved the problem of low recognition rate and high error detection rate in low visibility environments, and real-time and accurate detection of whether the vehicle is polite to pedestrians.
Patent Information
- Application Number
- CN202311670976.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-06
AI Technical Summary
The existing vehicle detection methods have low recognition rate, high error detection rate in low visibility environments, and insufficient real-time performance, making it difficult to accurately detect whether the vehicle gives way to pedestrians on the zebra crossing.
The YOLOv4-Tiny algorithm is used to improve the YOLOv4-Tiny algorithm, combined with the KCF tracking algorithm, obtain video information through the camera, detect the location of the vehicle and pedestrians, and determine whether the vehicle gives way to pedestrians.
The detection rate and accuracy of vehicles and pedestrians are improved, and real-time and accurate detection of whether vehicles give way to pedestrians in foggy conditions is achieved, reducing the false detection rate and reducing the calculation needs.
Smart Images

Figure CN120107993A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a vehicle giving way to pedestrians detection method of YOLOv4-Tiny improved by a domain adaptive algorithm based on class alignment, which belongs to the field of target detection in deep learning and is mainly applied to the scenario of whether vehicles give way to pedestrians on a crosswalk. Background Art
[0002] In urban traffic, pedestrians are particularly vulnerable road users in low-visibility environments such as foggy days due to the limited visibility of drivers. Road traffic accidents caused by foggy days often lead to serious injuries or even deaths due to unclear vision, accounting for a large proportion of traffic accidents. According to the "China Road Traffic Accident Statistics Yearbook", in traffic accidents caused by such low-visibility environments, the proportion of pedestrian deaths is as high as 26%. Therefore, designing a vehicle-passenger detection algorithm for foggy days is particularly critical to ensure road traffic safety.
[0003] Traditional methods of detecting vehicles yielding to pedestrians mainly rely on direct observation by traffic police on the scene or using traditional cameras to capture images to find out whether vehicles are yielding to pedestrians. However, the on-site judgment method of traffic police is inefficient, consumes human resources, and is easily affected by personal subjective judgment, which is prone to missed detection and false detection; while the conventional camera capture method is limited by the installation position and angle of the camera, and has problems such as low recognition rate, poor system stability and high computing power requirements.
[0004] The vehicle-to-passenger detection method proposed in the present invention is specially designed for low visibility conditions such as foggy days. It can detect the positions of pedestrians and vehicles from the video obtained by the camera in real time and accurately, and then determine whether the vehicle gives way to pedestrians. Currently common vehicle-to-passenger detection methods include: object-based detection and relationship-based detection. For example, Chinese patent CN104680133A discloses a relationship-based detection method, which acquires image and video information, monitors vehicle speed and tracks trajectory, and then determines the relative position relationship between the vehicle and pedestrians, and finally determines whether the vehicle gives way to pedestrians.
[0005] Although the above-mentioned vehicle-passing-pedestrian detection algorithm can process video information to obtain vehicle behavior results, it cannot guarantee the same viewing angle during camera installation. Once the camera angle changes, it may lead to errors in the recognition of vehicle and pedestrian positions. In addition, under poor weather conditions such as foggy days, the false detection rate will also increase greatly. At the same time, the existing algorithms are not real-time enough and require high computing power, which makes it difficult for them to accurately detect in real time whether vehicles are giving way to pedestrians on zebra crossings in foggy conditions. Summary of the invention
[0006] In order to overcome the deficiencies of the above-mentioned prior art, the present invention proposes a vehicle-and-pedestrian detection method based on a YOLOv4-Tiny algorithm improved by a domain adaptive algorithm based on class alignment. On the one hand, in order to improve the detection rate, the general target detection framework YOLOv4-Tiny in deep learning is optimized and designed, and a domain adaptive algorithm based on class alignment is introduced, thereby improving the detection rate of vehicles and pedestrians and improving the detection accuracy under foggy conditions. On the other hand, the method of drawing frames and labeling the detection area can greatly improve the detection accuracy even when the camera viewing angle changes, and realize real-time and accurate detection of whether the vehicle is courteous to pedestrians.
[0007] The technical solution adopted by the present invention is:
[0008] A vehicle-passenger detection method,
[0009] Step 1: Improved design of YOLOv4-Tiny algorithm
[0010] Step S1, YOLOv4-Tiny is a forward network composed of 9 convolutional layers and 6 maximum pooling layers alternately. The overall network structure is end-to-end. The CSPDarknet53 network is used as the backbone network.
[0011] Step S2, add upsampling layer in CSPDarknet53 network.
[0012] In step S3, the convolution layer is first deconvolved, then added to the feature layer, and finally convolved again.
[0013] Step S4, adding a horizontal line to normalize the final scale.
[0014] In step S5, a domain alignment module with a DANN network is added to reduce the distribution difference in feature space between the sunny pedestrian dataset and the foggy pedestrian dataset.
[0015] Step S6, constructing a hierarchical domain classifier in multiple intermediate layers of the backbone network, and training the domain classifier.
[0016] Step 2: Design of vehicle-passenger detection algorithm
[0017] Step S10: obtaining a regional video of the current area through a camera, processing the video into frame information, dividing the lane area in the frame, and naming the area ID;
[0018] Step S20: Perform preliminary processing on the information in the video, detect pedestrian information, use the improved YOLOv4-Tiny algorithm to identify the pedestrian in the current frame, and name the pedestrian ID;
[0019] Step S30: Detect vehicle information, use the improved YOLOv4-Tiny based on class alignment domain adaptation algorithm to identify the license plate number of the current vehicle, the status of the vehicle driver, and name the vehicle ID;
[0020] Step S40: Detect the position of pedestrians, use the KCF tracking algorithm to detect pedestrians and analyze the current position of pedestrians, and obtain the pedestrian position ID.
[0021] Step S50: Detect the vehicle position, use the KCF tracking algorithm to calculate the current position of the vehicle, and obtain the vehicle position ID.
[0022] Step S60: Assign area IDs to vehicles and pedestrians, and determine whether the vehicle gives way to pedestrians based on whether the pedestrian position ID and vehicle position ID in the previous and next frame detection areas are similar. If yes, return to step 1; if no, track and identify the license plate, record the license plate, and transmit the data back to the traffic management department.
[0023] The above technical solution determines whether a vehicle gives way to pedestrians based on regional information, which effectively solves the problem of low recognition rate or even recognition errors caused by the camera's viewing angle, making the vehicle detection results obtained based on the acquired vehicle image more accurate. At the same time, the YOLOv4-Tiny algorithm improved by the domain adaptive algorithm based on class alignment makes its recognition accuracy higher in foggy days, the recognition speed faster, and the delay of the detection system smaller. The above technical solution can obtain accurate video vehicle detection results even when the camera viewing angle is not good, and there is no need to use other complex road condition analysis equipment to analyze whether the vehicle gives way to pedestrians. It greatly improves the detection efficiency of whether the vehicle gives way to pedestrians, and also provides convenience for later system maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 : A step diagram of a vehicle-passenger detection method provided by an example of the present invention.
[0025] Figure 2 It is: YOLOv4-Tiny algorithm improvement flow chart.
[0026] Figure 3 For: Network diagram of target detection algorithm based on category domain adaptation. DETAILED DESCRIPTION
[0027] The present invention is further described below with reference to the accompanying drawings. An embodiment of the present invention provides a vehicle-passenger detection method in which multiple intermediate layers of a backbone network construct a hierarchical domain classifier, comprising:
[0028] Step 1: Improved design of YOLOv4-Tiny algorithm
[0029] Step S1 uses the CSPDarknet53-tiny network as the backbone network in YOLOv4-Tiny and adds an upsampling layer to the CSPDarknet53-tiny network. The added upsampling layer is integrated with the features of the previous convolutional layer to improve the network's feature learning of the target.
[0030] Step S2, first deconvolution operation is performed on the convolution layer, then it is added to the feature layer, and finally convolution operation is performed again. The 13*13 convolution layer is deconvolved twice, then it is added to the corresponding pixels of the previous feature layer 52*52, and then convolution operation with a step size of 2 is performed.
[0031] Step S3, add a horizontal line to normalize the final scale. Add a horizontal line (that is, fuse the features after the convolution operation in the previous step with the 26*26 feature layer), and finally normalize the final feature scale to 13*13 size for the final prediction, thereby improving the detection accuracy and making the detection faster.
[0032] In step S4, a domain alignment module with a DANN network is added to reduce the distribution difference in feature space between the sunny pedestrian dataset and the foggy pedestrian dataset.
[0033] Step S5, construct a hierarchical domain classifier in multiple intermediate layers of the backbone network, and train the domain classifier.
[0034] Step S6, further adjusting the alignment strategy and model parameters based on the evaluation results on the pedestrian dataset in a foggy environment to improve the performance of the model on the pedestrian dataset in a foggy environment.
[0035] Step 2: Design of vehicle-passenger detection algorithm
[0036] Step S11, acquiring a video image through a camera, and performing specific analysis on the video image to obtain key frame information therein.
[0037] Step S12, manually divide the lane information in the image frame information into 1, 2, 3, ...,
[0038] Step S21: Pedestrians are detected by using the YOLOv4-Tiny algorithm improved by the domain adaptation algorithm based on class alignment to detect pedestrians in the image frame and name the pedestrian ID.
[0039] The specific process is as follows;
[0040] Using the improved YOLOv4-Tiny algorithm, first resize the image to 448*448, then send the image to the VGG16 convolutional network for feature extraction, use the convolution filter to predict the category and coordinate position of the target, and set different predictors for different recognition objects, and finally perform non-maximum suppression to obtain the result. In pedestrian detection, the candidate box is extracted from the input video or image, and it is determined whether it contains pedestrians. If it is included, the position in the image is given. At the same time, the pedestrian ID is named.
[0041] Conventional
[0042] Step S31, detect vehicle information, use the improved YOLOv4-Tiny algorithm based on class alignment domain adaptation algorithm to detect the vehicle in the image frame, name it as vehicle ID, and detect the vehicle license plate information at the same time.
[0043] The process of detecting vehicles is similar to the above steps. Candidate boxes are extracted from the input video or image and it is determined whether the box contains a vehicle. If it does, the location of the image is given and the vehicle ID is named.
[0044] The license plate detection process is as follows:
[0045] Step S311 uses a window with a pixel size of 400x400 px to completely cover the license plate label, and adds a random offset of the window center coordinates so that the offset window still meets the goal of completely covering the license plate. Assume that the window size is Wwin*hwin, the license plate frame size is Wlabel*hlabel, and the license plate frame center position is (X o ,Y o ), then the random offset range of the x-axis of the window center position [20 is [Xc-(Wwin-Wlabe) / 2,Xc+(Wwin-Wlabe) / 2], and the random offset range of the y-axis is [Yc-(Wwin-Wlabe) / 2,Yc+(Wwin-Wlabe) / 2].
[0046] Step S312 uses a bilinear interpolation method to reduce the cropped window size and combines the original window size as a set of low-resolution and high-resolution training pairs. In order to enhance the network's ability to learn license plate information, randomness is added to the scaling factor, and all scaled images are reset to 200*200px as low-resolution input.
[0047] Step S313 uses a deconvolution layer to expand the size of the original image, and then generates a high-resolution result through a convolution layer to obtain the license plate information of the vehicle in the image.
[0048] Step S41 detects the position of pedestrians, uses the KCF tracking algorithm to detect pedestrians and analyze the current position of pedestrians, and obtains the pedestrian position ID. The main process is as follows:
[0049] First, the image is input, then it is tracked by Kalman filter and Hungarian matching, and finally the data is obtained to obtain the ID of the pedestrian in the image. Kalman filter tracking refers to the target information of a certain target in the previous frame obtained by the Kalman filter according to the target detection algorithm, and it is tracked and detected to obtain the specific information of the target in the next frame. After being processed by the Kalman filter tracking module, the information contained in each frame of the video is not only the target information obtained by the target detection algorithm, but also the target tracking information obtained by the tracking algorithm. Therefore, it is necessary to match the two types of information through the Hungarian matching module to obtain the final detection result, thereby avoiding the tracking target from being detected multiple times and reducing the performance of the algorithm. Kalman filter tracking actually consists of two processes: prediction process and correction process. The prediction process is that the filter uses the estimation of the previous state to make a prediction of the current state. The correction process is that the filter uses the observation value of the current state to correct the predicted value in the prediction stage, so as to obtain a new estimated value that is closer to the true value, thereby obtaining the location ID of the pedestrian.
[0050] Step S51 detects the vehicle position, uses the KCF tracking algorithm to calculate the current position of the vehicle, and obtains the vehicle position ID.
[0051] Similar to the above step 106, the image information is output, tracked by Kalman filter, and then Hungarian matching to finally obtain the location ID of the vehicle.
[0052] After obtaining the position IDs of the vehicle and pedestrian in the image, step S61 checks whether the two IDs are the same. If they are not the same, return to step 101; if they are the same, the vehicle did not give way to the pedestrian, and the vehicle's license plate is transmitted back to the database system.
[0053] The embodiment of the present invention provides a vehicle-passing pedestrian detection method based on the YOLOV4 algorithm. On the one hand, the improved YOLOv4-Tiny algorithm can increase the rate to 120 frames per second on the embedded development board using Nano, so as to accurately count the traffic conditions of pedestrian crossings in foggy environments. On the other hand, the method of whether the regional ID overlaps is used to determine whether the vehicle is courteous to pedestrians, which overcomes the existing technical defects and provides a vehicle-passing pedestrian detection method that does not require the camera installation position, has a fast detection rate, and has a low false detection rate. In addition, the calculation requirements of the entire vehicle-passing pedestrian detection system are greatly reduced, the vehicle-passing pedestrian detection efficiency is improved, and the implementation is simple while having strong independence.
[0054] The vehicle-passing-pedestrian detection method provided by the embodiment of the present invention can be applied in a big data traffic system. The detection cameras are installed under traffic light poles at different heights, thereby reducing the requirements for camera installation. At the same time, the YOLOv4-Tiny algorithm improved by the domain adaptation algorithm based on class alignment is used to realize real-time and accurate detection of vehicle-passing-pedestrian.
[0055] The above description is only a specific implementation mode of the present invention. Any feature disclosed in this specification, unless otherwise stated, can be replaced by other equivalent or alternative features with similar purposes; all disclosed features, or all methods or steps in the process, except for mutually exclusive features and / or steps, can be combined in any way.
Claims
1. A vehicle-passing-people detection method based on a domain adaptation algorithm improved by class alignment using YOLOv4-Tiny. It is characterized in that Includes the following sections: Step 1: Improved design of YOLOv4-Tiny algorithm Step S1 uses the CSPDarknet53-tiny network as the backbone network in YOLOv4-Tiny and adds an upsampling layer to the CSPDarknet53-tiny network. The added upsampling layer is integrated with the features of the previous convolutional layer to improve the network's feature learning of the target. Step S2 is one of the core steps of the patent. First, the convolution layer is deconvolved, then added to the feature layer, and finally convolution is performed again. The 13*13 convolution layer is deconvolved twice, then added to the corresponding pixels of the previous feature layer 52*52, and then convolution is performed with a step size of 2. Finally, a domain alignment module with a DANN network is added and a hierarchical domain classifier is constructed in multiple intermediate layers of the backbone network. Step S3, add a horizontal line to normalize the final scale. Add a horizontal line (that is, fuse the features after the convolution operation in the previous step with the 26*26 feature layer), and finally normalize the final feature scale to 13*13 size for the final prediction, thereby improving the detection accuracy and making the detection faster. In step S4, a domain alignment module with a DANN network is added to reduce the distribution difference in feature space between the sunny pedestrian dataset and the foggy pedestrian dataset. Step S5, construct a hierarchical domain classifier in multiple intermediate layers of the backbone network, and train the domain classifier. Step S6, further adjusting the alignment strategy and model parameters based on the evaluation results on the pedestrian dataset in a foggy environment to improve the performance of the model on the pedestrian dataset in a foggy environment. Step 2: Design of vehicle-passenger detection algorithm Step S11, acquiring a video image through a camera, and performing specific analysis on the video image to obtain key frame information therein. Step S12, which is the second core step of the patent, manually divides the lane information in the image frame information into 1, 2, 3, ..., Step S21: Pedestrians are detected by using the YOLOv4-Tiny algorithm improved by the domain adaptation algorithm based on class alignment to detect pedestrians in the image frame and name the pedestrian ID. Step S31, detect vehicle information, use the improved YOLOv4-Tiny algorithm based on class alignment domain adaptation algorithm to detect the vehicle in the image frame, name it as vehicle ID, and detect the vehicle license plate information at the same time. Step S41 detects the position of pedestrians, uses the KCF tracking algorithm to detect pedestrians and analyze the current position of pedestrians, and obtains the pedestrian position ID. The main process is as follows: Step S51 detects the vehicle position, uses the KCF tracking algorithm to calculate the current position of the vehicle, and obtains the vehicle position ID. Step S61 is the third core step of the patent. After obtaining the position ID of the vehicle and pedestrian in the image, check whether the two IDs are the same. If not, return to step 101; if they are the same, the vehicle did not give way to pedestrians, and the vehicle's license plate is transmitted back to the database system.
Citation Information
Patent Citations
Real-time detecting method for pedestrian avoidance behavior of illegal vehicle
CN104680133A