End-to-end shunting direction traffic flow density identification method and device
Through the end-to-end diversion to traffic flow density identification method, the traffic density is directly identified from the road monitoring image, solving the problems of high complexity and low efficiency of the existing methods, and achieving faster and more accurate traffic density identification.
Patent Information
- Application Number
- CN202510188483.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-13
AI Technical Summary
The existing lane traffic density identification methods are complex and inefficient, making it difficult to quickly and accurately identify traffic density.
The end-to-end diversion direction traffic density recognition method is adopted, and the road monitoring image is obtained, lane detection and arrow recognition are carried out, single lane images are generated, and the trained lane traffic density recognition model is used to directly identify the traffic density, and the average traffic density in each direction is calculated based on the arrow direction.
It realizes faster and less network complexity traffic density recognition, saves computing resources, and is suitable for automatic control systems for variable lanes.
Smart Images

Figure CN120147981A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to an end-to-end method and device for identifying the traffic flow density in a divided flow direction. Background Art
[0002] In some application scenarios, such as in an automatic control system for variable lanes, it is necessary to identify the traffic flow density of a target lane in a specific flow direction. The current mainstream method of obtaining traffic flow density information by capturing images through a camera is to first detect vehicles and then calculate the traffic flow density by the ratio of the area occupied by the detection box to the road area. This method has the disadvantages of consuming computing resources, high network complexity, and slow inference speed. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies in the prior art and provide an end-to-end method and device for identifying the traffic flow density in a divided flow direction, so as to solve the technical problems of high complexity and low efficiency existing in the existing method for identifying the traffic flow density of lanes.
[0004] To achieve the above object, the present invention is implemented by the following technical solutions: In a first aspect, the present invention provides an end-to-end method for identifying the traffic flow density in a divided flow direction, including: Obtaining a road surveillance image of a traffic intersection as an image to be recognized; Performing lane detection and arrow recognition on the image to be recognized to obtain a lane detection result and arrows on the lane; Cropping the image to be recognized according to the lane detection result to generate a single-lane image; Identifying the single-lane image through a trained lane traffic flow density recognition model to obtain the lane traffic flow density; Calculating the average traffic flow density of each direction according to the arrows on the lane and the lane traffic flow density.
[0005] Optionally, the training process of the lane traffic flow density recognition model includes: Obtaining a road surveillance image of a traffic intersection as a training sample image; Performing lane detection and vehicle detection on the training sample image to obtain a lane detection result and vehicle detection frames on the lane; Cropping the training sample image according to the lane detection result and the vehicle detection frames on the lane to generate a single-lane image with the vehicle detection frames retained; Performing an inverse perspective transformation on the single-lane image with the vehicle detection frames retained to generate a top-down view single-lane image; Calculating the lane traffic flow density according to the top-down view single-lane image and labeling the lane traffic flow density to the corresponding single-lane image; A lane traffic density recognition model is constructed, and the lane traffic density recognition model is trained using the labeled single lane image.
[0006] Optionally, the step of cropping the image to be identified according to the lane detection result to generate a single lane image includes: Crop each single lane into a single lane image according to the lane detection results; In each single lane image, the portion that does not belong to the single lane is filled with pure white.
[0007] Optionally, calculating lane traffic density according to the top-view single-lane image includes: Obtain the length of each vehicle detection frame and the total length of the single lane in the top-down single lane image, and calculate the lane traffic density : ; In the formula, For the The traffic density of a single lane is The top-down view of the single lane image The length of the vehicle detection frame and the total length of a single lane.
[0008] Optionally, the training of the lane traffic density recognition model by using the labeled single lane image includes: Initializing model parameters of the lane traffic density recognition model, wherein the model parameters include parameters of a feature extraction network and a multi-layer perceptron; Repeat the following steps until the loss function value is lower than the preset threshold: The annotated single lane image is input into the feature extraction network for feature extraction, and the feature extraction result is input into the multi-layer perceptron to predict the lane traffic density; The loss function is calculated by predicting the traffic density of the lane and the traffic density of the marked lane; Based on the loss function value, gradient backpropagation is used to optimize the parameters of the feature extraction network and the multi-layer perceptron through the optimizer.
[0009] In a second aspect, the present invention provides an end-to-end diversion direction vehicle flow density identification device, comprising: An image acquisition module is configured to acquire a road monitoring image of a traffic intersection as an image to be recognized; A lane arrow recognition module is configured to perform lane detection and arrow recognition on the image to be recognized, and obtain a lane detection result and an arrow on the lane; An image cropping module is configured to crop the image to be identified according to the lane detection result to generate a single lane image; The traffic flow density recognition module is configured to recognize the single-lane image through a trained lane traffic flow density recognition model to obtain the lane traffic flow density; The sub-flow direction average value module is configured to calculate the average value of the traffic flow density in each direction according to the lane arrow and the lane traffic flow density on the lane.
[0010] Optionally, the training process of the lane traffic flow density recognition model includes: Obtain the road monitoring image of the traffic intersection as the training sample image; Perform lane detection and vehicle detection on the training sample image to obtain the lane detection result and the vehicle detection frame on the lane; Crop the training sample image according to the lane detection result and the vehicle detection frame on the lane to generate a single-lane image with the vehicle detection frame retained; Perform inverse perspective transformation on the single-lane image with the vehicle detection frame retained to generate a single-lane image from a bird's-eye view; Calculate the lane traffic flow density according to the single-lane image from the bird's-eye view, and label the lane traffic flow density to the corresponding single-lane image; Construct a lane traffic flow density recognition model, and train the lane traffic flow density recognition model through the labeled single-lane image.
[0011] In a third aspect, the present invention provides an electronic device, including a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the above method.
[0012] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0013] In a fifth aspect, the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0014] Compared with the prior art, the beneficial effects achieved by the present invention: The end-to-end sub-flow direction traffic flow density recognition method and device provided by the present invention bypass vehicle detection and directly recognize the traffic flow density of the lane from the lane picture. Combining with the lane arrow direction, the traffic flow density in each direction is counted, and the result can be used as a reference for the automatic control of the variable lane. Under the condition of ensuring the detection accuracy rate, compared with the traditional object detection method, it has a faster detection speed and less network complexity. It has the remarkable advantages of saving computing resources and being easy to deploy. Description of the Drawings
[0015] Figure 1 It is a schematic flowchart of the end-to-end split flow direction traffic flow density recognition method provided by an embodiment of the present invention; Figure 2 It is an example diagram of a road monitoring image of a traffic intersection provided by an embodiment of the present invention; Figure 3 It is an example diagram of a lane detection result provided by an embodiment of the present invention; Figure 4 It is a schematic structural diagram of a lane traffic flow density recognition model provided by an embodiment of the present invention; Figure 5 It is a schematic flowchart of the training process of the lane traffic flow density recognition model provided by an embodiment of the present invention; Figure 6 It is an example diagram of a vehicle detection result provided by an embodiment of the present invention; Figure 7 It is an example diagram of a single-lane image with vehicle detection frames retained provided by an embodiment of the present invention; Figure 8 It is an example diagram of a single-lane image from a top-down perspective provided by an embodiment of the present invention; Figure 9 It is a schematic logical diagram of the end-to-end split flow direction traffic flow density recognition method provided by an embodiment of the present invention. Detailed implementation manners
[0016] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0017] Embodiment 1:
[0018] As Figure 1 shown, an embodiment of the present invention provides an end-to-end split flow direction traffic flow density recognition method, including the following steps: Step S1, obtain a road monitoring image of a traffic intersection as an image to be recognized.
[0019] Collect the road monitoring video of the traffic intersection. As Figure 2 shown, extract one frame of image from the road monitoring video at intervals of a period of time t as the road monitoring image. t can be set flexibly and is default to 1 second. The obtained road monitoring images can be stored in the form of a picture set.
[0020] Step S2, perform lane detection and arrow recognition on the image to be recognized to obtain a lane detection result and an arrow on the lane.
[0021] Both lane detection and arrow recognition adopt the solutions in the prior art. For example, the patent with the publication number CN110852228B describes a method and system for extracting dynamic background and detecting foreground objects in a surveillance video, which can perform road background extraction and lane detection, and its detection results are as Figure 3 shown. For example, the patent with the publication number CN113947763A describes a method and device for road arrow recognition based on template self-supervision, which can perform arrow recognition. There are also other technical solutions for lane detection and arrow recognition in the prior art, which will not be listed one by one in this embodiment.
[0022] Step S3: Crop the image to be recognized according to the lane detection result to generate a single-lane image.
[0023] Each single lane is cropped into a single-lane image according to the lane detection result. For example, if a lane contains 4 single lanes, 4 single-lane images need to be cropped. To facilitate subsequent recognition operations, in each single-lane image, the parts that do not belong to the single lane are filled with pure white.
[0024] Step S4: Recognize the single-lane image through the trained lane traffic flow density recognition model to obtain the lane traffic flow density.
[0025] Specifically, in this embodiment, as Figure 4 shown, the lane traffic flow density recognition model includes a feature extraction network and a multi-layer perceptron. The feature extraction network adopts the ResNet50 network, including Stages 1 to 4. Stage 1 contains 3 convolutional layers and 2 pooling layers, which are mainly responsible for extracting features from the input image. Stage 2 contains 4 convolutional layers and 2 pooling layers, which are mainly responsible for further deepening and enhancing the features extracted in the first stage. Stage 3 contains 6 convolutional layers and 2 pooling layers, which are mainly responsible for further deepening and enhancing the features extracted in the second stage. Stage 4 contains 3 convolutional layers and 1 pooling layer, which are mainly responsible for further deepening and enhancing the features extracted in the third stage. The multi-layer perceptron MLP makes a prediction and recognition based on the features output by the ResNet50 network.
[0026] Step S5: Calculate the average traffic flow density in each direction according to the arrows on the lane and the lane traffic flow density.
[0027] In the above steps, obtaining the trained lane traffic flow density recognition model is an important link. As Figure 5 shown, in this embodiment, the training process of the lane traffic flow density recognition model includes the following steps: Step S11: Obtain the road surveillance images at the traffic intersection as training sample images.
[0028] Step S12: Perform lane detection and vehicle detection on the training sample images to obtain the lane detection results and vehicle detection boxes on the lanes.
[0029] For vehicle detection, the solutions in the prior art are adopted. For example, the improvement of vehicle detection algorithm for intelligent traffic guidance published in the 9th issue of 2022 of Computer Technology and Development can perform vehicle detection on each image and generate vehicle detection boxes for each vehicle in the image as Figure 6 shown. There are also other technical solutions for vehicle detection in the prior art, which will not be listed one by one in this embodiment.
[0030] Step S13: Crop the training sample images according to the lane detection results and vehicle detection boxes on the lanes to generate single-lane images with vehicle detection boxes retained.
[0031] Similarly, in each single-lane image, the parts that do not belong to the single lane are filled with pure white, and the result is as Figure 7 shown.
[0032] Step S14: Perform inverse perspective transformation on the single-lane images with vehicle detection boxes retained to generate top-down view single-lane images.
[0033] Perspective transformation and inverse perspective transformation are important techniques in the fields of computer vision and image processing. Inverse perspective transformation is the inverse process of perspective transformation, which restores a perspectively transformed image to its original view or plane. For example, when you take a photo of a slanted table, the tabletop may look trapezoidal, and through inverse perspective transformation, it can be restored to a rectangle. By performing inverse perspective transformation on the single-lane images with vehicle detection boxes retained, top-down view single-lane images are generated, as Figure 8 shown.
[0034] Calculating the lane traffic flow density based on the top-down view single-lane images includes: Obtain the lengths of each vehicle detection box and the total length of the single lane in the top-down view single-lane image, and calculate the lane traffic flow density : ; In the formula, is the lane traffic flow density of the th single lane, is the length of the th vehicle detection box and the total length of the single lane in the top-down view single-lane image.
[0035] Step S15: Calculate the lane traffic flow density based on the top-down view single-lane images, and label the lane traffic flow density to the corresponding single-lane images.
[0036] Step S16: Construct a lane traffic flow density recognition model, and train the lane traffic flow density recognition model with the labeled single-lane images. Specifically, it includes: Step S21: Initialize the model parameters of the lane traffic flow density recognition model. The model parameters include the parameters of the feature extraction network and the multi-layer perceptron.
[0037] Step S22: Repeat steps S23 - S25 until the value of the loss function is lower than the preset threshold.
[0038] Step S23: Input the labeled single-lane image into the feature extraction network for feature extraction, and input the feature extraction result into the multi-layer perceptron to predict the lane traffic flow density.
[0039] Step S24: Calculate the loss function based on the lane traffic flow density obtained by prediction and the labeled lane traffic flow density.
[0040] In machine learning, the loss function is an important tool for measuring the difference between the model prediction and the actual label. Common loss functions include Mean Squared Error (MSE), Cross-Entropy Loss, etc., and each has its unique advantages and disadvantages. Specifically, in this embodiment, the mean squared error is used as the loss function for calculation.
[0041] Step S25: Based on the value of the loss function, perform backpropagation using the gradient, and optimize the parameters of the feature extraction network and the multi-layer perceptron through the optimizer.
[0042] Specifically, in this embodiment, the initial learning rate is 0.001, the weight decay is set to 0.0005, compare that the loss function descent rate is lower than the set threshold, and obtain a trained specific lane traffic flow density detection network. The threshold is set to 0.01.
[0043] As Figure 9 shown, an end-to-end traffic flow density detection method for different directions proposed in the embodiment of the present invention bypasses vehicle detection, directly recognizes the traffic flow density of the lane from the lane picture, and combines with the direction of the lane arrow to count the traffic flow density in each direction. The result can be used as a reference for the automatic control of variable lanes.
[0044] Extract pure road surface using road surface background extraction technology, process the images extracted from road monitoring videos using lane detection and inverse perspective transformation, retain a single lane, and calculate the traffic flow density information of this lane obtained by the vehicle detection method. Use the processed data to train a neural network so that the network can capture the traffic flow density information in the images. Use a CNN-based feature extraction network to extract the traffic flow density information in the images and convert it into a feature vector, use a multi-layer perceptron to output the feature vector as a regression value to represent the traffic flow density of the lane in the image, and return the traffic flow direction information of the lane by identifying the road surface arrows. The present invention trains and uses an end-to-end network to directly capture images through a camera and output the traffic flow density of the lane, greatly reducing the network complexity and greatly improving the inference speed of the network.
[0045] Embodiment 2:
[0046] The embodiment of the present invention provides an end-to-end traffic flow density recognition device with split flow directions, including: An image acquisition module, configured to acquire a road monitoring image of a traffic intersection as an image to be recognized; A lane arrow recognition module, configured to perform lane detection and arrow recognition on the image to be recognized to obtain a lane detection result and the arrows on the lane; An image cropping module, configured to crop the image to be recognized according to the lane detection result to generate a single-lane image; A traffic flow density recognition module, configured to recognize the single-lane image through a trained lane traffic flow density recognition model to obtain the traffic flow density of the lane; A split-flow mean value module, configured to calculate the mean value of the traffic flow density in each direction according to the arrows on the lane and the traffic flow density of the lane.
[0047] Specifically, the training process of the lane traffic flow density recognition model includes: Acquire a road monitoring image of a traffic intersection as a training sample image; Perform lane detection and vehicle detection on the training sample image to obtain a lane detection result and vehicle detection frames on the lane; Crop the training sample image according to the lane detection result and the vehicle detection frames on the lane to generate a single-lane image retaining the vehicle detection frames; Perform inverse perspective transformation on the single-lane image retaining the vehicle detection frames to generate a single-lane image with a bird's-eye view; Calculate the traffic flow density of the lane according to the single-lane image with a bird's-eye view, and label the traffic flow density to the corresponding single-lane image; Construct a lane traffic flow density recognition model, and train the lane traffic flow density recognition model through the labeled single-lane images.
[0048] Embodiment 3:
[0049] Based on the end-to-end traffic flow diversion and traffic density recognition method provided in the first embodiment, an embodiment of the present invention provides an electronic device, including a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the above method.
[0050] Embodiment Four:
[0051] Based on the end-to-end traffic flow diversion and traffic density recognition method provided in the first embodiment, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0052] Embodiment Five:
[0053] Based on the end-to-end traffic flow diversion and traffic density recognition method provided in the first embodiment, an embodiment of the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0054] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0055] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0056] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions in the processFigure 1 one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.
[0057] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple processes and / or the functions specified in one block or multiple blocks.
[0058] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. An end-to-end diversion traffic density identification method, characterized in that: include: Acquire a road monitoring image of a traffic intersection as an image to be recognized; Performing lane detection and arrow recognition on the image to be recognized to obtain a lane detection result and an arrow on the lane; Cropping the image to be identified according to the lane detection result to generate a single lane image; The single lane image is recognized by using a trained lane traffic density recognition model to obtain lane traffic density; The mean value of the traffic density in each direction is calculated according to the arrows on the lane and the traffic density of the lane.
2. The end-to-end diversion direction traffic density identification method according to claim 1 is characterized in that: The training process of the lane traffic density recognition model includes: Obtain road monitoring images of traffic intersections as training sample images; Performing lane detection and vehicle detection on the training sample image to obtain a lane detection result and a vehicle detection frame on the lane; The training sample image is cropped according to the lane detection result and the vehicle detection frame on the lane to generate a single lane image retaining the vehicle detection frame; Performing an inverse perspective transformation on the single-lane image retaining the vehicle detection frame to generate a single-lane image from a bird's-eye view; Calculating lane traffic density according to the top-view single-lane image, and marking the lane traffic density to the corresponding single-lane image; A lane traffic density recognition model is constructed, and the lane traffic density recognition model is trained using the labeled single lane image.
3. The end-to-end diversion direction traffic density identification method according to claim 1 is characterized in that: The step of cropping the image to be identified according to the lane detection result to generate a single lane image comprises: Crop each single lane into a single lane image according to the lane detection results; In each single lane image, the portion that does not belong to the single lane is filled with pure white.
4. The end-to-end diversion direction traffic density identification method according to claim 2 is characterized in that: Calculating the lane traffic density according to the top-view single-lane image includes: Obtain the length of each vehicle detection frame and the total length of the single lane in the top-down single lane image, and calculate the lane traffic density : ; In the formula, For the The traffic density of a single lane is The top-down view of the single lane image The length of the vehicle detection frame and the total length of a single lane.
5. The end-to-end diversion direction traffic density identification method according to claim 2 is characterized in that: The training of the lane traffic density recognition model by using the labeled single lane image includes: Initializing model parameters of the lane traffic density recognition model, wherein the model parameters include parameters of a feature extraction network and a multi-layer perceptron; Repeat the following steps until the loss function value is lower than the preset threshold: The annotated single lane image is input into the feature extraction network for feature extraction, and the feature extraction result is input into the multi-layer perceptron to predict the lane traffic density; The loss function is calculated by predicting the traffic density of the lane and the traffic density of the marked lane; Based on the loss function value, gradient backpropagation is used to optimize the parameters of the feature extraction network and the multi-layer perceptron through the optimizer.
6. An end-to-end diversion direction traffic density identification device, characterized in that: include: An image acquisition module is configured to acquire a road monitoring image of a traffic intersection as an image to be recognized; A lane arrow recognition module is configured to perform lane detection and arrow recognition on the image to be recognized, and obtain a lane detection result and an arrow on the lane; An image cropping module is configured to crop the image to be identified according to the lane detection result to generate a single lane image; A traffic density recognition module is configured to recognize the single lane image by using a trained lane traffic density recognition model to obtain lane traffic density; The diversion direction average module is configured to calculate the average of the traffic density in each direction according to the arrows on the lane and the traffic density of the lane.
7. The end-to-end diversion direction traffic density identification device according to claim 6, characterized in that: The training process of the lane traffic density recognition model includes: Obtain road monitoring images of traffic intersections as training sample images; Performing lane detection and vehicle detection on the training sample image to obtain a lane detection result and a vehicle detection frame on the lane; The training sample image is cropped according to the lane detection result and the vehicle detection frame on the lane to generate a single lane image retaining the vehicle detection frame; Performing an inverse perspective transformation on the single-lane image retaining the vehicle detection frame to generate a single-lane image from a bird's-eye view; Calculating lane traffic density according to the top-view single-lane image, and marking the lane traffic density to the corresponding single-lane image; A lane traffic density recognition model is constructed, and the lane traffic density recognition model is trained using the labeled single lane image.
8. An electronic device, characterized in that: including processor and storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Methods and Systems for Dynamic Background Extraction and Foreground Object Detection in Surveillance Video
CN110852228B
Pavement arrow identification method and device based on template self-supervision
CN113947763A