A Multi-Model Detection Method and System for Intelligent Transportation
By adopting a multi-model detection method in the field of smart transportation, replacing the loss function and backbone network of YOLO v5 and deploying secondary networks, the shortcomings of single model detection in real-time, accuracy and deployment difficulty are solved, and higher detection accuracy and faster detection speed are achieved.
Patent Information
- Application Number
- CN202310291427.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-03-23
AI Technical Summary
In the field of smart transportation, the existing single model detection method has shortcomings in real-time, detection accuracy and embedded platform deployment difficulty, resulting in slow detection speed and deployment difficulties.
The multi-model detection method is adopted, and the loss function of YOLO v5 is replaced by SIOU and the backbone network as GhostNet, and the secondary networks MTCNN, LPRNet and ResNet are deployed in the system to perform primary and secondary network collaborative detection.
It improves the accuracy of main target detection and the effect of multi-object tracking in traffic scenarios, reduces single-frame detection time, avoids lag, reduces the computing power demand for embedded devices, and improves the overall computing speed.
Smart Images

Figure CN116311100B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of object detection and intelligent transportation, and mainly relates to a multi-model detection method and system for intelligent transportation. Background Technique
[0002] Object detection is an important and popular topic in the field of computer vision. Due to its high efficiency and scalability, object detection has high practical value in today's increasingly complex traffic environment. However, there are still problems in terms of accuracy, tracking accuracy, difficulty in deploying on embedded platforms, and real-time performance, which have hindered its large-scale deployment and scenario applications in intelligent transportation. Previous researchers have used networks such as MobileNet for depthwise separable convolution classification to identify vehicles and pedestrians. Its characteristic is relatively fast speed, but its detection accuracy is not very high, and the model is not lightweight enough, with high requirements for deployment to embedded devices, and there is still much room for improvement. Moreover, when conventional single-model detection is used in the field of intelligent transportation, the real-time detection and computing capabilities are poor, and lag will occur when the amount of data obtained and processed in a single-frame detection is too large. This problem is undoubtedly fatal in fields that highly value real-time performance and detection accuracy. Therefore, single-model detection alone cannot achieve parallel detection of primary and secondary targets within a single frame, and solving this problem is crucial for the application of object detection in the field of intelligent transportation. Summary of the Invention
[0003] Object of the Invention: The present invention helps the lightweight deployment of object detection in the field of intelligent transportation by proposing a multi-model detection method and system for intelligent transportation. It replaces the original single detection model with multi-model detection, improves the real-time detection frame rate, shortens the single-frame detection time, effectively avoids lag caused by detecting and processing a large number of targets within a frame, and reduces the computing power requirements for embedded devices. Based on the YOLO v5 object detection main detection network, it is improved by replacing the original loss function and the original backbone network with SIOU and GhostNet. On the basis of meeting real-time performance, it improves the accuracy of detecting primary targets and the effect of multi-target tracking in traffic scenarios. At the same time, MTCNN and ResNet are used for secondary target auxiliary detection to improve the detection speed. And a visualization interface is developed, providing a software platform for the further implementation of this invention.
[0004] Technical Solution: To achieve the above object, the technical solution adopted by the present invention is as follows:
[0005] A multi-model detection method for intelligent transportation, comprising the following steps:
[0006] Step S1, taking yolov5 as the main network and replacing its original loss function with SIOU;
[0007] Step S2: On the basis of Step S1, simultaneously replace the C3 module in the backbone (the fundamental structure in the neural network) of the yolov5 network structure with GhostNet (a method for extracting feature maps with low cost and high efficiency).
[0008] Step S3: On the basis of Step S2, simultaneously process the training set samples using 9-mosaic (a data augmentation method), and then send them into the model for training to obtain the optimal weight file.
[0009] Step S4: While using yolov5 as the main detection network of the system, deploy secondary networks MTCNN, LPRNet, and ResNet inside the system. The secondary network MTCNN is used to detect the license plate area, the secondary network LPRNet is used for license plate character recognition, and the secondary network ResNet is used for traffic light color recognition.
[0010] Step S5: Read the weight file obtained in Step S3 in the system, set each system parameter, calibrate the zebra crossing, traffic lights, lane line class markers, and then detect vehicles and pedestrians.
[0011] Step S6: According to the detection results of the main network yolov5, extract and save the vehicle area and pedestrian area; slice the extracted vehicle area and send it into the secondary network MTCNN for secondary recognition to obtain the license plate area; if the license plate area is detected, slice this area again and send it into the secondary network LPRNet for license plate character recognition, and then return the recognition result and save it.
[0012] Step S7: Parallelly extract and save the detection results of the secondary network ResNet for the calibrated traffic light area.
[0013] Step S8: Combine the detection results obtained from all primary and secondary networks with the calibrated area coordinates and frame-by-frame coordinate displacement data for behavior judgment.
[0014] Furthermore, in Step S1, the loss function SIOU cited is defined as follows:
[0015]
[0016]
[0017]
[0018]
[0019] where Δ is the distance loss; ρ x is the ratio of the width difference between the center points of the ground truth box and the predicted box to the square of the width of the minimum bounding rectangle, ρ yIt is the ratio of the height difference between the center points of the ground truth box and the predicted box to the square of the height of the minimum bounding rectangle; e is the natural constant; γ is 2 - Λ, where Λ is the angular loss; Ω is the shape loss; θ controls the degree of attention to the shape loss; w is the width of the predicted box, h is the height of the predicted box; w t is the ratio of the absolute value of the difference between the width (height) of the predicted box and the width (height) of the ground truth box to the maximum value of the width (height) of the predicted box and the width (height) of the ground truth box; IOU is the intersection over union, B is the predicted box, B GT is the ground truth box; L box is the bounding box regression loss.
[0020] Furthermore, in the step S2, the design theory of the GhostNet module is as follows:
[0021] Step S2.1: First, generate a set of basic feature maps through the main convolution: X is the input feature map, where the size parameters are width h im , length w im , and the number of input channels is c; f’ is the convolutional layer, where the size parameters are the convolutional kernel size k and the number of output channels is m; the input feature map X outputs a set of feature maps Y’ after convolution with the convolutional layer f’, where the size parameters are width h’ and length w’; R is the set of feature maps;
[0022] Y'∈R h'×w'×m , f'∈R c×k×k×m
[0023] Y' = X * f';
[0024] Step S2.2: Perform a linear transformation on the output set of basic feature maps Y’ to generate Ghost feature maps; where y i ' is the i-th feature map in the set of basic feature maps Y’, and is the linear transformation performed by the j-th generated Ghost feature map. Here, the operation means generating s Ghost feature maps through a linear operation for each feature map in the set of basic feature maps, and the last linear operation is the identity mapping to retain the inherent features;
[0025]
[0026] Y' = [y1, y2,... y num
[0027] Y = [y 11 , y 12 ,... y 1s ,.... y nums ;
[0028] where both num and s are constants; y num is the num-th feature map; Φi,,j is a linear transformation;
[0029] Step S2.3: Output num×s feature maps as the final output feature map set, that is, Y (the set of y in the above formula). i,j ).
[0030] Furthermore, in the above-mentioned step S3, the 9-Mosaic data augmentation method is used; this method randomly crops, scales 9 images, and then randomly arranges and splices them to form one image, which enriches the data set, increases small-sample targets, ensures high recognition rate at close range, improves the road long-distance detection ability of the model, and improves the training speed of the network; when performing the normalization operation, the data of 9 images is calculated at one time, so the memory requirement of the model is reduced.
[0031] Furthermore, in the above-mentioned step S4, the activation function adopted by the convolutional network in the MTCNN model (secondary network MTCNN) is PReLU (Parametric Rectified Linear Unit, parametric Relu), and its formula is as follows:
[0032] f(yi) = max(0, yi) + aimin(0, yi)
[0033] where yi is the input of the non-linear activation function f in the i-th channel; f(yi) is the output value of the activation function; ai is responsible for controlling the slope of the negative half-axis; different channels are allowed to have different activation functions; when ai = 0, PReLU (Parametric Rectified Linear Unit, parametric Relu) becomes ReLU (activation function), and ai is a learnable parameter.
[0034] Furthermore, in step S4, the secondary network MTCNN abandons the conventional cascaded third-layer O-Net network and merges the core of the cascaded second-layer R-Net structure to the end of the P-Net network.
[0035] Furthermore, in step S6, the extracted region format is (x1, y1, x2, y2, cls), which are the upper-left coordinates (x1, y1), lower-right coordinates (x2, y2), and category respectively; this format is the general format for subsequent region slicing; the operations for feeding into the secondary network MTCNN are as follows:
[0036] Step S6.1: First, perform transformations on the input image at different scales to construct an image pyramid;
[0037] Step S6.2: Input the P-Network to quickly generate candidate windows. For the image pyramid constructed in Step S6.1, perform feature extraction and calibration of the bounding box through a Fully Convolutional Networks (FCN), and perform Bounding-Box Regression to adjust the window and Non-Maximum Suppression (NMS) to filter the windows, and then output the results.
[0038] Further, for the result output by the MTCNN network, before the image slice is fed into the LPRNet, it first undergoes perspective transformation for skew correction. The general formula for perspective transformation is as follows:
[0039]
[0040] where a ij represents the 9 elements of the perspective transformation matrix; u, v are the original positions, d is a constant 1; x', y' are the positions after transformation, and d' is a high-dimensional parameter;
[0041] The coordinates obtained after perspective transformation are:
[0042]
[0043]
[0044] This operation improves the character recognition accuracy of LPR_detect without significantly increasing the model complexity.
[0045] Further, in Step S7, the ResNet residual network is ResNet-18, and the residual structure it adopts is BasicBlock (a linear code sequence of ours); when inputting data, its steps are as follows:
[0046] Step S7.1: Conv1, the first layer of convolution, this convolutional layer does not come with a shortcut mechanism (a method in the CNN model to solve the problem of gradient divergence caused by increasing the network depth); the calculation formula for this layer is:
[0047]
[0048] where n out is the size of the output image of the convolution; n in is the size of the input image of the convolution; p represents the number of zero paddings; k is the size of the convolution kernel; st is the stride; the output of this layer is 64×112×112;
[0049] Step S7.2, Conv2: The first residual block, there are 2 in total, and the output is 64×56×56;
[0050] Step S7.3, Conv3: The second residual block, there are 2 in total, and the output is 128×28×28;
[0051] Step S7.4, Conv4: The third residual block, there are 2 in total, and the output is 256×14×14.
[0052] Step S7.5, Conv5: The fourth residual block, there are 2 in total, and the output is 512×7×7;
[0053] Step S7.6, fc: The fully connected layer, and the output is 512×1×1;
[0054] The ResNet residual network also includes a dilated convolution module for increasing the receptive field of the convolution kernel for the image to be recognized.
[0055] The present invention also discloses a system for a multi-model detection method for intelligent transportation. This system is a multi-model detection system for intelligent transportation, which performs weight file selection, detection device selection, and intersection over union (IoU) threshold setting; the system performs road marker calibration, calculates behavior judgment synchronously during detection, and records, saves, and displays traffic violation behaviors in real time on the system visualization interface.
[0056] Beneficial effects:
[0057] The present invention helps the lightweight deployment of object detection in the field of intelligent transportation by proposing a multi-model detection method and system for intelligent transportation. Its core idea is to replace the original single detection model with an optimized multi-model detection. For the main detection model, SIOU and GhostNet are used to replace the original loss function and the original backbone network. On the basis of meeting real-time requirements, it can improve the accuracy of detecting main objects and the effect of multi-object tracking in traffic scenarios. At the same time, the secondary models MTCNN and ResNet are used to perform parallel auxiliary detection on secondary objects, detecting and calculating secondary objects without occupying the detection and calculation time of main objects, which can greatly improve the overall operation speed and reduce the computing power occupancy requirements within a single thread. The introduction of the multi-model idea enables the original network to improve the detection ability for a large number of and various types of objects, which is crucial especially for objects that originally require a large amount of calculation after detection. Therefore, this detection method can well solve problems such as the slow detection rate and difficult deployment of a single model in the field of intelligent transportation. Description of the drawings
[0058] Figure 1 is the flowchart of the multi-model detection method and system for intelligent transportation provided by the present invention. Specific embodiments
[0059] The present invention will be further described below with reference to the accompanying drawings.
[0060] As Figure 1 shown, a multi-model detection method and system for intelligent transportation, characterized by including the following steps:
[0061] Step S1: Take yolov5 as the main network and replace its original loss function with SIOU, which is defined as follows:
[0062]
[0063]
[0064]
[0065]
[0066] In formula (1), Δ is the distance loss; ρ x is the ratio of the width difference between the center points of the ground truth box and the predicted box to the square of the width of the minimum bounding rectangle, and ρ y is the ratio of the height difference between the center points of the ground truth box and the predicted box to the square of the height of the minimum bounding rectangle; e is the natural constant; γ is 2 - Λ (angle loss).
[0067] In formula (2), Ω is the shape loss; θ controls the degree of attention to the shape loss, and the general parameter range is [2, 6]. In this paper, the parameter value is 4; w and h are the width and height of the predicted box; w t is the ratio of the absolute value of the width (height) of the predicted box - the width (height) of the ground truth box to the maximum value of the width (height) of the predicted box and the width (height) of the ground truth box.
[0068] In formula (3), IOU is the intersection over union, B is the predicted box, and B GT is the ground truth box.
[0069] In formula (4), L box is the bounding box regression loss.
[0070] Step S2: On the basis of step S1, simultaneously replace the C3 module in the backbone of the yolov5s network structure with GhostNet; the design theory of the GhostNet module is as follows:
[0071] First, generate a set of basic feature maps through the main convolution: X is the input feature map, where the size parameters are width h im , length w im, the number of input channels is c; f’ is a convolutional layer, where the size parameter is the convolutional kernel size k, and the number of output channels is m; the input feature map X outputs a set of feature maps Y’ after convolution with the convolutional layer f’, where the size parameters are h’ (width) and w’ (length); R is a set of feature maps.
[0072] Y' ∈ R h'×w'×m , X ∈ R h×w×c , f' ∈ R c×k×k×m
[0073] Y' = X * f'
[0074] Then, a linear transformation is performed on the output set of basic feature maps Y’ to generate Ghost feature maps; where y i ' is the i-th feature map in the set of basic feature maps Y’, and it is the linear transformation performed by the j-th generated Ghost feature map. Here, the operation means that each feature map in the corresponding set of basic feature maps generates s Ghost feature maps through a linear operation, and the last linear operation is an identity mapping to retain the inherent features;
[0075]
[0076] Y' = [y1, y2,... y num
[0077] Y = [y 11 , y 12 ,... y 1s ,.... y nums ;
[0078] Where both num and s are constants; y num is the num-th feature map; Ф i,j is a linear transformation;
[0079] Finally, num × s feature maps are output as the final output set of feature maps, that is, Y (the set of y i,j in the above formula).
[0080] Step S3: On the basis of step S2, use 9-mosaic processing on the training set samples at the same time, and then send them into the model for training to obtain the optimal weight file; the main idea of the 9-Mosaic data augmentation method is to randomly crop, scale 9 images, and then randomly arrange and splice them to form an image. While enriching the dataset, it increases small-sample targets, ensures high recognition rate at close range, and improves the road long-range detection ability of the model more than 4-Mosaic, and improves the training speed of the network. When performing the normalization operation, the data of 9 images will be calculated at one time, so the memory requirement of the model is slightly reduced.
[0081] Step S4: While using yolov5 as the main detection network of the system, deploy secondary networks MTCNN, LPRNet, and ResNet inside the system, which are used to detect license plate regions, recognize license plate characters, and recognize traffic light colors respectively; the activation function used in the convolutional network of the MTCNN model is PReLu, and its formula is as follows:
[0082] f(yi)=max(0,yi)+ai min(0,yi)
[0083] where yi is the input of the non-linear activation function f in the i-th channel, and ai is responsible for controlling the slope of the negative half-axis. Here, we allow the activation functions of different channels to be different. When ai = 0, PReLu becomes ReLu, and ai is a parameter that can be learned.
[0084] Moreover, on the premise of not requiring high precision for the return value of this network, in order to improve the real-time detection speed of multiple models, this MTCNN discards the conventional cascaded third-layer O-Net network, and simply merges the core of the cascaded second-layer R-Net structure to the end of the P-Net network.
[0085] Step S5: Read the weight file obtained in Step S3 in the system, set each system parameter, calibrate various markers such as zebra crossings, traffic lights, and lane lines, and then detect vehicles and pedestrians;
[0086] Step S6: According to the detection results of the main network yolov5, extract and save the vehicle region and pedestrian region; slice the extracted vehicle region and send it into MTCNN for secondary recognition to obtain the license plate region; if the license plate region is detected, slice this region again and send it into LPRNet for license plate character recognition, and then return the recognition result and save it; the format of the above-extracted region is (x1, y1, x2, y2, cls), which are the upper-left, lower-right coordinates and category respectively. This format is the general format for subsequent region slicing. The operation of sending data into MTCNN is as follows:
[0087] First, perform transformations of different scales on the input image to construct an image pyramid.
[0088] Then input it into the P-Net network to quickly generate candidate windows. For the image pyramid constructed in the previous step, perform feature extraction and calibration of the border through a FCN, and perform Bounding-Box Regression (border regression) to adjust the window and filter most of the windows through NMS, and output the result.
[0089] The image slices sent into LPRNet later will first undergo perspective transformation for skew correction, and the general formula for perspective transformation is as follows:
[0090]
[0091] Among them, a ij represents 9 elements of the perspective transformation matrix; u and v are the original positions, d is the constant 1; x' and y' are the positions after transformation, and d' is a high-dimensional parameter.
[0092] The coordinates obtained after perspective transformation are:
[0093]
[0094]
[0095] This operation can significantly improve the character recognition accuracy of LPR_detect without greatly increasing the model complexity.
[0096] Step S7: Parallelly extract the detection results of the secondary network ResNet for the calibrated traffic light area and save them; among them, the ResNet residual network is ResNet-18, and the residual structure it adopts is BasicBlock. When the input data is 3×224×224, its steps are as follows:
[0097] Step S7.1: Conv1, the first convolutional layer, which does not have a shortcut mechanism. The calculation formula of this layer is:
[0098]
[0099] In the formula, n out is the size of the convolutional output image; n in is the size of the convolutional input image; p represents the number of zero padding; k is the size of the convolutional kernel; st is the stride; the output of this layer is 64×112×112;
[0100] Step S7.2: Conv2: The first residual block, there are 2 in total, and the output is 64×56×56.
[0101] Step S7.3: Conv3: The second residual block, there are 2 in total, and the output is 128×28×28.
[0102] Step S7.4: Conv4: The third residual block, there are 2 in total, and the output is 256×14×14.
[0103] Step S7.5: Conv5: The fourth residual block, there are 2 in total, and the output is 512×7×7.
[0104] Step S7.6: fc: The fully connected layer, and the output is 512×1×1.
[0105] The traffic light recognition model further includes an atrous convolution module for increasing the receptive field of the convolution kernel for the image to be recognized.
[0106] Step S8: Combine the results obtained from all the primary-secondary network detections with data such as the calibration area coordinates and the inter-frame coordinate displacement to make a behavior judgment.
[0107] A multi-model detection system for intelligent transportation can perform other parameter settings such as weight file selection and detection device selection; it can perform road marker calibration, calculate behavior judgment synchronously during detection, and can record and save traffic violation behaviors in real time and then display them on the visualization interface of the system.
[0108] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A multi-model detection method for intelligent transportation, characterized in that, It includes the following steps: Step S1: Take yolov5 as the main network and replace its original loss function with SIOU; Step S2: On the basis of Step S1, simultaneously replace the C3 module in the backbone of the yolov5 network structure with GhostNet; Step S3: On the basis of Step S2, simultaneously perform 9-mosaic processing on the training set samples, and then send them into the model for training to obtain the optimal weight file; Step S4: While taking yolov5 as the main detection network of the system, deploy secondary networks MTCNN, LPRNet, and ResNet inside the system. The secondary network MTCNN is used to detect the license plate area, the secondary network LPRNet is used for license plate character recognition, and the secondary network ResNet is used for traffic light color recognition; Step S5: Read the weight file obtained in Step S3 in the system, set the system parameters, calibrate the zebra crossing, traffic lights, lane line class markers, and then detect vehicles and pedestrians; Step S6: According to the detection results of the main network yolov5, extract and save the vehicle area and pedestrian area; Slice the extracted vehicle area and send it into the secondary network MTCNN for secondary recognition to obtain the license plate area; If the license plate area is detected, slice the area again and send it into the secondary network LPRNet for license plate character recognition, and then return the recognition result and save it; Step S7: Parallelly extract the detection results of the secondary network ResNet for the calibrated traffic light area and save them; Step S8: Combine the detection results of all primary and secondary networks with the calibrated area coordinates and the frame-by-frame coordinate displacement data for behavior judgment.
2. The multi-model detection method for intelligent transportation according to claim 1, wherein, In the said Step S1, the referenced loss function SIOU is defined as follows: where Δ is the distance loss; ρ x is the ratio of the width difference between the centers of the ground truth box and the predicted box to the square of the width of the minimum bounding rectangle, and ρ y is the ratio of the height difference between the centers of the ground truth box and the predicted box to the square of the height of the minimum bounding rectangle; e is the natural constant; γ is 2 - Λ, where Λ is the angular loss; Ω is the shape loss; θ controls the degree of attention to the shape loss; w is the width of the predicted box, and h is the height of the predicted box; w t=w is the ratio of the absolute value of the difference between the width of the predicted box and the width of the ground truth box to the maximum value of the width of the predicted box and the width of the ground truth box; w t=x is the ratio of the absolute value of the difference between the height of the predicted box and the height of the ground truth box to the maximum value of the height of the predicted box and the height of the ground truth box; IOU is the intersection over union, B is the predicted box, and B GT is the ground truth box; L box is the bounding box regression loss.
3. A multi-model detection method for intelligent transportation according to claim 1, characterized in that In the said Step S2, the design of the GhostNet module is as follows: Step S2.1: First, generate a set of basic feature maps through the main convolution: X is the input feature map, where the size parameter is the width h im , the length w im , and the number of input channels is c; f’ is a convolutional layer, where the size parameter is the convolutional kernel size k and the number of output channels is m; the input feature map X outputs the feature map set Y’ after convolution with the convolutional layer f’, where the size parameters are the width h’ and the length w’; R is the feature map set; Y' ∈ R h'×w'×m , f' ∈ R c×k×k×m Y' = X * f'; Step S2.2: Perform a linear transformation on the output set of basic feature maps Y' to generate Ghost feature maps; where y i ' is the i-th feature map in the set of basic feature maps Y', and it is the linear transformation performed on the j-th generated Ghost feature map. Here, the operation means generating s Ghost feature maps for each feature map in the set of basic feature maps through a linear operation, and the last linear operation is an identity mapping to retain the inherent features; Y' = [y1, y2,... y num Y = [y 11 , y 12 ,... y 1s ,.... y nums ; where num and s are both constants; y num is the num-th feature map; Φ i,j is a linear transformation; Step S2.3: Output num × s feature maps as the final output feature map set, that is, Y.
4. A multi-model detection method for intelligent transportation according to claim 1, characterized in that, In the said Step S3, the 9-Mosaic data augmentation method is used; this method randomly crops, scales 9 pictures, and then randomly arranges and stitches them together to form one picture, which enriches the dataset and increases small-sample targets; when performing the normalization operation, the data of 9 pictures will be calculated at one time.
5. A multi-model detection method for intelligent transportation according to claim 1, characterized in that In the said Step S4, the activation function adopted by the convolutional network is PRelu, and its formula is as follows: f(yi) = max(0, yi) + ai min(0, yi) where yi is the input of the non-linear activation function f in the i-th channel; f(yi) is the output value of the activation function; ai is responsible for controlling the slope of the negative half-axis; different channels are allowed to have different activation functions; when ai = 0, PReLu becomes ReLu, and ai is a parameter that can be learned.
6. The multi-model detection method for intelligent transportation according to claim 5, wherein In step S4, the secondary network MTCNN discards the conventional cascaded third-layer O-Network and merges the core of the cascaded second-layer R-Net structure to the end of the P-Network.
7. A multi-model detection method for intelligent transportation according to claim 1, characterized in that, In step S6, the extracted region format is (x1, y1, x2, y2, cls), which are the upper-left coordinates (x1, y1), the lower-right coordinates (x2, y2), and the category respectively; this format is the general format for subsequent region slicing; the operations for feeding into the secondary network MTCNN are as follows: Step S6.1: First, transform the input image at different scales to construct an image pyramid. Step S6.2: Input it into the P-Network to quickly generate candidate windows; for the image pyramid constructed in step S6.1, perform feature extraction and calibration of the bounding box through a fully convolutional network (FCN), and perform bounding box regression, adjust the windows, and filter the windows through non-maximum suppression (NMS), and output the results.
8. A multi-model detection method for intelligent transportation according to claim 7, characterized in that, The results output by the MTCNN network, before the image slices are fed into LPRNet, first undergo perspective transformation for skew correction. The general formula for perspective transformation is as follows: where a ij represents nine elements of the perspective transformation matrix; u and v are the original positions, d is the constant 1; x' and y' are the transformed positions, and d' is a high-dimensional parameter; The coordinates obtained after perspective transformation are:
9. A multi-model detection method for intelligent transportation according to claim 1, characterized in that, In step S7, the ResNet residual network is ResNet-18, and the residual structure it adopts is BasicBlock; when inputting data, its steps are as follows: Step S7.1: Conv1, the first layer of convolution, and this convolutional layer does not come with a shortcut mechanism; the calculation formula for this layer is: where n out is the size of the convolutional output image; n in is the size of the convolutional input image; p represents the number of zero paddings; k is the size of the convolutional kernel; st is the stride; the output of this layer is 64×112×112; Step S7.2: Conv2: The first residual block, with a total of 2, and the output is 64×56×56. Step S7.3: Conv3: The second residual block, with a total of 2, and the output is 128×28×28. Step S7.4: Conv4: The third residual block, with a total of 2, and the output is 256×14×14. Step S7.5: Conv5: The fourth residual block, with a total of 2, and the output is 512×7×7. Step S7.6: fc: The fully connected layer, and the output is 512×1×1. This ResNet residual network also includes a dilated convolution module, which is used to increase the receptive field of the convolutional kernel for the image to be recognized.
10. A system for a multi-model detection method for intelligent transportation according to any one of claims 1-9, which performs weight file selection, detection device selection, and intersection over union (IoU) threshold setting; the system performs calibration of road markers, calculates behavior judgment synchronously during detection, and records, saves, and displays traffic violation behaviors in real time on the system visualization interface.
Citation Information
Patent Citations
License plate detection method based on multi-task cascade convolutional neural network
CN110033002A
Collaborative deep network model method for pedestrian detection
WO2018107760A1