A Statistical Investigation Method for Traffic Safety Facilities Based on Deep Learning

By combining GPS/IMU data to construct a corrected position transformation matrix, perform monocular depth estimation and image registration, use YOLO algorithm to identify traffic safety facilities, and combine it with fuzzy inference system to evaluate the status, the problems of low efficiency and insufficient accuracy of traffic safety facilities in the existing technology are solved, and efficient and accurate three-dimensional reconstruction and state judgment are achieved.

CN120107906BActive Publication Date: 2025-07-11JIANGXI HIGHWAY RES & DESIGN INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510579040.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-11
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

In the prior art, the survey and statistical work efficiency of road traffic safety facilities is low, the data dimension is limited, the state judgment is highly subjective, and the image recognition results are difficult to effectively correlate with spatial structure information, which affects the comprehensive judgment of the three-dimensional reconstruction accuracy and the facility status.

Method used

By collecting video images and GPS/IMU data, a corrected pose transformation matrix is constructed, and three-dimensional point cloud reconstruction is realized by combining monocular depth estimation. The YOLO algorithm is used to identify targets, generate structured data, and a fuzzy inference system is introduced to judge the status of the facility, and the output is segmented and summarized according to the station number.

Benefits of technology

The spatial consistency and depth information alignment accuracy between multi-frame images is improved, the three-dimensional geometric parameters of the facility are accurately extracted, flexible and intelligent facility status evaluation is realized, and segmented statistics and report output of the facility are supported, which improves the degree of quantifiability of data and the value of engineering application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107906B_ABST
    Figure CN120107906B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of traffic safety facility recognition, and specifically provides a statistical survey method for traffic safety facilities based on deep learning. The present invention includes: collecting video frame images and synchronously obtaining GPS and IMU data of the camera device; constructing a corrected pose transformation matrix through a fusion pose estimation algorithm, generating a three-dimensional point cloud by combining monocular depth estimation and multi-frame depth map fusion, and extracting geometric attribute information of traffic safety facilities; using the YOLO algorithm to identify traffic safety facilities in the image, extracting the facility category and location, and associating them with three-dimensional attributes to generate structured data; judging the facility status level based on a fuzzy inference method; performing segmented statistics and classified summaries according to road mileage, and outputting a statistical report. The present invention can realize three-dimensional perception of facilities, intelligent evaluation of status, and classified statistics, and is applicable to the operation and maintenance of intelligent roads and the management of digital traffic facilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of traffic safety facility identification, and specifically to a method for statistical investigation of traffic safety facilities based on deep learning. Background Art

[0002] With the continuous development of the transportation network and the advancement of intelligent infrastructure construction, road traffic safety facilities (such as signboards, guardrails, markings, contour markers, etc.), as an important part of ensuring driving safety and conveying road information, the rationality of their layout and the intact state directly affect road traffic efficiency and driving safety. Therefore, the automatic identification, state perception, and structured management of traffic safety facilities have become key links in highway maintenance and intelligent transportation systems.

[0003] Currently, the investigation and statistical work of road traffic safety facilities mainly rely on manual on-site inspection or semi-automatic equipment collection, which has problems such as low work efficiency, insufficient coverage, limited data dimensions, and strong subjectivity in state judgment. Even in scenarios where image recognition or laser point cloud equipment is introduced, there are also technical bottlenecks such as difficulty in accurately obtaining the three-dimensional dimensions of facilities, serious interference from image ghosting, and difficulty in structuring and fusing data, which are difficult to meet the actual needs of digital management of the new generation of intelligent transportation facilities. In addition, existing technologies often lack an accurate modeling and correction mechanism for the spatial pose changes between the collected data, resulting in difficulty in aligning multiple frames of images or depth information in space, affecting the accuracy of three-dimensional reconstruction. The image recognition results are also difficult to be effectively associated with the spatial structure information, limiting the comprehensive judgment ability of the actual state of the facilities.

[0004] Therefore, a method for statistical investigation of traffic safety facilities based on deep learning is proposed. Summary of the Invention

[0005] The present invention provides a method for statistical investigation of traffic safety facilities based on deep learning. By collecting video image data and synchronously obtaining GPS / IMU data, a corrected pose transformation matrix is constructed, and three-dimensional point cloud reconstruction is realized in combination with monocular depth estimation to extract the geometric attributes of the facilities; the YOLO algorithm is used for target recognition to generate structured data; and a fuzzy inference system is introduced to judge the state of the facilities, and the statistics and summary are output by segmenting according to the stake number.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A method for statistical investigation of traffic safety facilities based on deep learning, comprising:

[0008] Collect video frame image data, synchronously obtain the GPS and IMU data of the camera device, calculate and correct the relative pose transformation matrix between video frame images based on the GPS and IMU data, and obtain the corrected pose transformation matrix;

[0009] Perform monocular depth estimation on each video frame image to generate a dense depth map; combine the dense depth maps of multiple consecutive frames with the corresponding corrected pose transformation matrix for spatial alignment and fusion, construct a three-dimensional scene point cloud, and extract the geometric attribute information of traffic safety facilities;

[0010] Use the corrected pose transformation matrix as the motion constraint for optical flow estimation, combine the inter-frame image features, perform image registration and motion compensation, and generate a de-blurred image;

[0011] Use the YOLO object detection algorithm to identify traffic safety facilities in the de-blurred image, obtain the categories of the facilities, and associate them with the geometric attribute information of the traffic safety facilities to generate structured data of traffic safety facilities;

[0012] Perform fuzzy inference on the structured data of traffic safety facilities and the standard attributes in the traffic safety facility database to judge the status information of the facilities;

[0013] Based on the structured data and status information, perform segmented statistics and classification summary on traffic safety facilities according to the preset road mileage range, and output a statistical report.

[0014] Further, the steps of obtaining the corrected pose change include:

[0015] Based on the three-axis angular velocity and acceleration data collected by the IMU sensor, use the integration method to estimate the rotation angle and displacement vector of the camera device between consecutive frame images, and construct a relative pose transformation matrix;

[0016] Combine the GPS data to correct the relative pose transformation matrix to obtain the corrected pose transformation matrix.

[0017] Further, the steps of constructing a three-dimensional scene include:

[0018] Input each frame of image into a convolutional neural network for monocular depth estimation to generate a dense depth map;

[0019] Combine the corrected pose transformation matrix between adjacent frame images to perform spatial coordinate transformation on the corresponding dense depth map to achieve three-dimensional alignment of dense depth maps of different frames;

[0020] Use a point cloud fusion algorithm to perform cumulative fusion on multiple frames of dense depth maps, including voxel filtering for noise reduction, point density equalization, and edge preservation enhancement, to generate a dense three-dimensional point cloud.

[0021] Further, the steps of image registration and motion compensation include:

[0022] Based on the corrected pose transformation matrix, predict the overall motion trend between consecutive image frames, generate an initial displacement estimate of the image region, and use it as the initial input for optical flow calculation;

[0023] Combining the feature point matching relationship between image frames, the LK optical flow algorithm based on the image pyramid structure is used to perform multi-scale optical flow estimation to obtain an optical flow vector field;

[0024] During the calculation of the optical flow field, the direction predicted by the corrected pose transformation matrix is used as guiding information to constrain and weighted optimize the direction of the optical flow vector field;

[0025] Based on the optimized optical flow vector field, pixel-level registration and motion compensation are performed on the image frames to generate an image sequence free of motion blur.

[0026] Further, the steps of generating structured data of traffic safety facilities include:

[0027] Input the de-blurred image into the YOLO object detection network to obtain the bounding box position, class label and recognition confidence of traffic safety facilities;

[0028] Perform spatial mapping of the recognized class label to the corresponding area in the 3D scene point cloud, and extract the geometric attribute information of the facility according to the image coordinate and point cloud projection relationship;

[0029] Based on the class label and geometric attribute information, construct structured data entries, and associate the station number range where each facility is located.

[0030] Further, the steps of fuzzy inference include:

[0031] Set the facility category, size deviation, position deviation and recognition confidence as input fuzzy variables, and define their linguistic values and membership functions;

[0032] Construct a fuzzy rule base containing the mapping relationship between input variables and output variables, where the output variable is the facility status level;

[0033] Adopt the Mamdani type fuzzy inference system, and combine the maximum membership degree method for defuzzification processing, and the output result is a level label representing the facility status.

[0034] Further, the steps of segmented statistics and classification summary include:

[0035] According to the spatial position information and station number range in the structured data, divide all recognized traffic safety facilities into pile segments at a set interval;

[0036] In each pile segment, count the quantity, size range and status level of various facilities, and generate a corresponding statistical table.

[0037] The beneficial effects of the present invention are:

[0038] 1. The present invention fuses IMU inertial data and GPS position information to construct a corrected pose transformation matrix between continuous image frames. Compared with the traditional registration method based only on image features, it can effectively reduce the errors caused by integral drift and image jitter, and significantly improve the spatial consistency and depth information alignment accuracy between multiple frame images, providing a stable spatial foundation for subsequent 3D reconstruction and geometric extraction.

[0039] 2. This invention introduces a monocular depth estimation network and combines it with posture transformation to achieve multi-frame dense depth map fusion, generate a dense point cloud model, and accurately extract the length, width, height, clearance and other three-dimensional geometric parameters of the facility. Compared with traditional image recognition that can only output two-dimensional position and category, it improves the spatial dimension and quantification degree of facility information, providing key support for facility modeling, deviation judgment and digital operation and maintenance.

[0040] 3. Based on the information of traffic facility type, size deviation, position deviation and recognition confidence, the present invention constructs a fuzzy rule base and uses the Mamdani reasoning system to grade the facility status, overcoming the problems of threshold rigidity and error sensitivity in traditional judgment methods and realizing a flexible and intelligent status evaluation mechanism. At the same time, in conjunction with the pile number division logic, it supports segmented statistics and report output of facility status, and has strong engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0042] Figure 1 This is a flow chart of a traffic safety facility statistical survey method based on deep learning provided by the present invention. DETAILED DESCRIPTION

[0043] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0044] Embodiment 1

[0045] A statistical survey method for traffic safety facilities based on deep learning, such as Figure 1 As shown, including:

[0046] S100: Collecting video frame image data, synchronously acquiring GPS and IMU data of the camera device, calculating and correcting the relative posture transformation matrix between the video frame images based on the GPS and IMU data, and obtaining the corrected posture transformation matrix;

[0047] Furthermore, the step of obtaining the corrected posture change includes:

[0048] Based on the three-axis angular velocity and acceleration data collected by the IMU sensor, the integral method is used to estimate the rotation angle and displacement vector of the imaging device between consecutive frame images, and a relative pose transformation matrix is constructed;

[0049] Combined with the GPS data, the relative pose transformation matrix is corrected to obtain a corrected pose transformation matrix.

[0050] Specifically, continuous image frames are collected by a camera installed on a mobile acquisition platform (such as a vehicle). At the same time, the three-axis angular velocity and three-axis acceleration collected by the IMU, as well as the three-dimensional position data and absolute timestamp provided by the GPS, are obtained. First, according to the IMU angular velocity data, the direction cosine matrix integration method is used to calculate the relative rotation change of the camera between two frames of images , where represents two adjacent frames. Then, the acceleration data is double-integrated to obtain the translation vector , and the rotation change and the translation vector constitute the relative pose transformation matrix between image frames :

[0051] ;

[0052] Since there is a time accumulation error in the IMU integration, in order to improve the pose estimation accuracy, GPS data is introduced to correct the pose calculated by the above IMU. The correction process includes: synchronizing and interpolating the GPS and IMU data for alignment; using the extended Kalman filter to compensate for the position deviation of the translation amount. The corrected result constitutes the corrected pose transformation matrix finally used for image alignment and point cloud fusion.

[0053] By fusing GPS and IMU sensor data, constructing and correcting the relative pose transformation matrix between video frame images can effectively improve the accuracy and stability of pose estimation between image frames. Through the high-frequency inertial data of the IMU, the real-time calculation of the continuous motion trajectory of the camera is realized, and then combined with the absolute position and attitude reference information provided by the GPS for correction, which can effectively suppress the cumulative drift error caused by the IMU integration, improve the robustness of pose estimation in complex dynamic environments, provide an accurate spatial reference for subsequent depth map alignment, three-dimensional scene reconstruction and image registration, and enhance the environmental perception ability and engineering application reliability of the entire system.

[0054] S200: Perform monocular depth estimation on each frame of video frame image to generate a dense depth map; combine the dense depth maps of multiple consecutive frames with the corresponding corrected pose transformation matrix for spatial alignment and fusion, construct a three-dimensional scene point cloud, and extract the geometric attribute information of traffic safety facilities;

[0055] Furthermore, the steps of constructing a three-dimensional scene include:

[0056] Each frame of image is input into a convolutional neural network for monocular depth estimation to generate a dense depth map;

[0057] Combined with the corrected pose transformation matrix between adjacent frame images, spatial coordinate transformation is performed on the corresponding dense depth maps to achieve three-dimensional alignment of the dense depth maps of different frames;

[0058] A point cloud fusion algorithm is used to cumulatively fuse multiple frames of dense depth maps, including voxel filtering for noise reduction, point density equalization, and edge-preserving enhancement, to generate a dense three-dimensional point cloud.

[0059] Specifically, each frame of video image collected is input into a pre-trained convolutional neural network model for monocular depth estimation. The depth estimation network uses a deep learning model based on an encoder-decoder structure, such as Monodepth, DPT, or their improved versions, and outputs a dense depth map with the same size as the original image. In the dense depth map, each pixel contains the relative depth value of the scene point corresponding to the pixel. Then, combined with the corrected pose transformation matrix corresponding to each frame of image, three-dimensional spatial coordinate transformation is performed on each frame of depth map. Specifically, each depth pixel (u, v, d) is back-projected into a three-dimensional point (X, Y, Z) in the camera coordinate system through the camera intrinsic matrix, and is uniformly transformed to the common world coordinate system through the corrected pose transformation matrix to complete the three-dimensional alignment of multiple frames of depth maps. The back-projection transformation process is:

[0060] ;

[0061] where X, Y, and Z represent the coordinates of the X-axis, Y-axis, and Z-axis in the three-dimensional space coordinate system, u, v, and d represent the row index, column index, and depth value of the image pixel, and K represents the camera intrinsic matrix. After that, a point cloud fusion algorithm is used to fuse the three-dimensional aligned multi-frame point cloud data, including voxel grid filtering to reduce noise and redundant data, point density equalization to improve spatial consistency, and an edge-preserving algorithm to enhance the details of the facility contour, and finally generate a dense and continuous three-dimensional point cloud model.

[0062] Through this step, high-precision and multi-frame fusion three-dimensional scene reconstruction can be achieved, which not only effectively overcomes the instability of single-frame depth estimation, but also improves the extraction accuracy of the geometric attributes of traffic safety facilities and the spatial positioning accuracy, providing a reliable three-dimensional structure basis for subsequent facility recognition and status evaluation.

[0063] S300: Use the corrected pose transformation matrix as the motion constraint for optical flow estimation, combine the inter-frame image features, perform image registration and motion compensation, and generate a de-blurred image;

[0064] Furthermore, the steps of image registration and motion compensation include:

[0065] Based on the corrected pose transformation matrix, predict the overall motion trend between consecutive image frames, generate an initial displacement estimate of the image region, and use it as the initial input for optical flow calculation;

[0066] Combined with the feature point matching relationship between image frames, use the LK optical flow algorithm based on the image pyramid structure to perform multi-scale optical flow estimation and obtain the optical flow vector field;

[0067] During the calculation of the optical flow field, use the direction predicted by the corrected pose transformation matrix as the guiding information to constrain and weight-optimize the direction of the optical flow vector field;

[0068] Based on the optimized optical flow vector field, perform pixel-level registration and motion compensation on the image frames to generate an image sequence without motion blur.

[0069] Specifically, based on the corrected pose transformation matrix obtained in the previous step, preliminarily model the motion trend between two consecutive frames of images. Through the rotation information and displacement information contained in this matrix, the overall motion direction and speed of different regions in the image frame can be predicted, which are used to estimate the initial displacement vector of pixel blocks or key points and serve as the initial value for optical flow estimation. Then, extract and match the feature points between the inter-frame images to obtain the corresponding set of feature points, which are used to construct the matching window in optical flow estimation. Use the LK optical flow algorithm based on the image pyramid structure to perform multi-scale optical flow estimation on the consecutive frame images, and solve layer by layer from low resolution to high resolution. Finally, obtain the optical flow vector field F(x, y) of the entire image. This vector field represents the displacement direction and amplitude of each pixel point in the front and back frame images. During the calculation of the optical flow vector field, use the predicted displacement direction in the corrected pose transformation matrix as the guiding information to perform direction consistency verification with the locally obtained optical flow direction in the image; in weak texture regions or fast motion scenes, perform directional regularization and noise suppression on the optical flow field through weighted optimization, thereby improving the global stability and local robustness of optical flow estimation. Finally, based on the optimized optical flow vector field, perform pixel-level registration and image resampling on the image frames, and perform position compensation and brightness adjustment on the blurred and motion-blurred regions caused by camera motion to generate a clear image sequence without motion blur.

[0070] By introducing the fused corrected pose transformation matrix into the optical flow estimation process, not only the stability and accuracy of the traditional Lucas-Kanade optical flow algorithm in complex scenarios are improved, but also the direction consistency of image registration is significantly enhanced. Combining the multi-scale pyramid structure and the direction regularization optimization mechanism can still maintain a good pixel-level registration effect in cases of fast motion, occlusion, weak texture, etc. The finally output de-blurred image provides a high-quality input guarantee for subsequent YOLO recognition and geometric attribute extraction, enhancing the reliability and engineering practicability of the entire system.

[0071] S400: Use the YOLO object detection algorithm to identify traffic safety facilities in the de-blurred image, obtain the categories of the facilities, and associate them with the geometric attribute information of the traffic safety facilities to generate structured data of traffic safety facilities;

[0072] Furthermore, the steps of generating structured data of traffic safety facilities include:

[0073] Input the de-blurred image into the YOLO object detection network to obtain the bounding box position, category label, and recognition confidence of traffic safety facilities;

[0074] Perform spatial mapping of the recognized category label to the corresponding area in the 3D scene point cloud, and extract the geometric attribute information of the facility according to the image coordinate and point cloud projection relationship;

[0075] Construct structured data entries based on the category label and geometric attribute information, and associate the station number range where each facility is located.

[0076] Specifically, the image frame that has completed motion compensation and ghosting removal is input into a pre-trained YOLO object detection model. The YOLO model can be YOLO v5, YOLO v7, YOLO v8, or other object detection architectures with lightweight deployment capabilities. The model outputs include the bounding box position, class label, and recognition confidence. Next, the bounding box region of each object in the image is spatially corresponded with the 3D point cloud data. Specifically, based on the internal parameter consistency between the image and the point cloud, the depth map or point cloud corresponding to the pixel points in the bounding box is filtered and back-projected to establish a mapping relationship from the image detection result to the 3D point cloud. Through this mapping, the image region corresponding to the YOLO detection result is converted into a 3D point cloud segment, so that the actual geometric attribute information of the facility can be extracted, including dimension parameters (length, width, height), spatial position (X, Y, Z), installation height, pitch angle, etc. Finally, combining the recognized class label, recognition confidence, and geometric attribute information, a structured data entry for the traffic safety facility is constructed. The structured entry can include the following fields: facility category (such as speed limit sign, guardrail, contour marker, etc.), spatial position (such as center coordinates, station number where it is located), dimension parameters (such as width, height, clearance), recognition confidence, and the range of the road station number to which it belongs (calculated based on GPS coordinates and road matching).

[0077] Through this step, not only the efficient recognition of traffic safety facilities from images is realized, but also the spatial fusion of the recognition results and 3D geometric information is completed, and then a structured data format with practical usability is constructed. This method takes into account visual intelligence and spatial accuracy, providing a solid data foundation for the digital management and status assessment of facility assets.

[0078] S500: Perform fuzzy inference on the structured data of traffic safety facilities and the standard attributes in the traffic safety facility database to judge the status information of the facilities;

[0079] Furthermore, the steps of fuzzy inference include:

[0080] Set the facility category, dimension deviation, position deviation, and recognition confidence as input fuzzy variables, and define their linguistic values and membership functions;

[0081] Construct a fuzzy rule base that contains the mapping relationship between the input variables and the output variable, where the output variable is the facility status level;

[0082] Adopt a Mamdani-type fuzzy inference system and perform defuzzification processing in combination with the maximum membership degree method. The output result is a level label representing the facility status.

[0083] Specifically, multiple input fuzzy variables for judging the facility status are set, including but not limited to: facility category (such as speed limit signs, corrugated guardrails, crossing marker posts, etc.), size deviation (relative error between the detected size and the standard size), position deviation (distance between the actual installation position of the facility and the designed position), and recognition confidence (confidence score output by the YOLO model). Corresponding linguistic variable values and membership functions are defined for each input variable respectively. For example, the linguistic values of size deviation include "small", "normal", and "large", and its membership function can adopt triangular or trapezoidal functions; the linguistic values of position deviation include "slightly smaller", "moderate", and "slightly larger"; the linguistic values of recognition confidence include "low", "medium", and "high". Taking size deviation as an example, its fuzzy membership function can be set as:

[0084] ;

[0085] Secondly, a fuzzy rule base is constructed. The rules map the combinations between input variables to the output variable "facility status level", and its linguistic values can be: "qualified", "acceptable", "abnormal", "severely abnormal". Example rules include:

[0086] IF the size deviation is normal AND the position deviation is moderate AND the confidence is high THEN the status is qualified;

[0087] IF the size deviation is large OR the position deviation is large THEN the status is acceptable;

[0088] IF the size deviation is large AND the position deviation is large AND the confidence is low THEN the status is abnormal. Then, a Mamdani-type fuzzy inference system is used to perform fuzzy logic calculations on the above rules, and the inference engine uses the min-max composition method to perform fuzzy set superposition and aggregation on multiple rules. Finally, the defuzzification process is carried out using the maximum membership degree method, and the final facility status level label is extracted from the output fuzzy set as the evaluation result of the facility. The label can be selected as status identifiers such as "qualified", "slight deviation", "unqualified", etc.

[0089] Through this fuzzy inference system, under the conditions of high uncertainty in facility recognition and complex environmental interference, the facility status can be flexibly judged by comprehensively considering multi-dimensional attribute indicators, avoiding misjudgment or missed judgment caused by threshold judgment, improving the intelligence and fault tolerance of facility evaluation, and providing more stable and reasonable data support for subsequent statistics, maintenance, and management.

[0090] S600: Based on the structured data and status information, the traffic safety facilities are segmented and statistically analyzed and classified and summarized according to the preset road mileage range, and a statistical report is output.

[0091] Furthermore, the steps of segmented statistics and classified summary include:

[0092] According to the spatial position information and mileage range in the structured data, all identified traffic safety facilities are divided into mileage segments at a set interval.

[0093] Within each mileage segment, count the quantity, size range, and status level of various facilities, and generate corresponding statistical tables.

[0094] Specifically, according to the spatial position information (GPS coordinates or the center point coordinates of the three-dimensional point cloud) of each traffic safety facility in the structured data and its identified road name or number, perform mileage mapping on the position of the facility in the actual road. The mileage range or mileage point value to which each facility belongs can be determined by referring to the standard road mileage database or projecting the GPS coordinates into the existing road linear reference model. Then, according to the set statistical interval (such as every 100 meters or 500 meters as a mileage segment), the road is divided into equal-distance segments by mileage. Each segment is used as a statistical unit, corresponding to a set of facility data. For each mileage segment, based on the content of the structured data, count the following information: the quantity of various traffic safety facilities (such as signs, guardrails, markings, etc.), the size range (maximum value, minimum value, average value) of each type of facility, and the distribution of status levels (such as the proportion of the quantity of qualified, slightly abnormal, and severely abnormal). The statistical process can be realized by traversing the structured data and performing clustering operations according to the belonging mileage, supporting online processing or offline batch processing. Optionally, the system can also sort, filter, and mark anomalies for the statistical results.

[0095] Through the segmented statistical mechanism based on mileage, accurate classified summary of traffic safety facilities in different road sections can be realized, facilitating the unified management and horizontal comparative analysis of various facility data. The output results not only meet the technical evaluation requirements but also have good engineering interface compatibility, and are suitable for the direct integration and invocation of road traffic management systems, infrastructure maintenance systems, and intelligent transportation platforms.

[0096] Embodiment 2

[0097] In this embodiment, a 12-kilometer expressway test section is selected, and various traffic safety facilities are equipped along the line, including speed limit signs, corrugated guardrails, markings, contour markers, etc. A vehicle platform equipped with a front monocular camera, a GPS positioning module, and an IMU inertial sensor is used to collect data at a normal passing speed (about 80 km / h).

[0098] During the collection process, the camera continuously records video images at a frequency of 20 frames per second, the IMU synchronously collects three-axis angular velocity and acceleration data, and the GPS module provides geographical coordinates and timestamp information.

[0099] In the post - processing stage, first, based on the integration of IMU data, the relative rotation and displacement estimation results between every two frames of images are obtained, and the cumulative drift is corrected by GPS coordinates to form a corrected pose transformation matrix between consecutive image frames.

[0100] Then, a deep neural network (such as Monodepth2) is used to perform monocular depth estimation on each frame of image to generate a dense depth map; by cooperating with the corrected pose transformation matrix, multiple depth maps are spatially aligned in the same coordinate system and fused into a three - dimensional scene point cloud.

[0101] Meanwhile, the YOLOv5 model is used to perform object detection on the image sequence after removing motion blur, identify the facility categories such as speed limit signs and guardrails and their bounding box information, and perform back - projection in combination with the point cloud to extract geometric attributes such as their actual spatial dimensions, positions, and installation heights.

[0102] Subsequently, the identified structured data is used as input variables, and a Mamdani - type fuzzy inference system is adopted to input the size deviation, position deviation, and recognition confidence, and output the facility status level. The facility status results include three categories: "qualified", "slight deviation", and "unqualified".

[0103] Finally, according to the rule of taking every 500 meters as a pile segment, the identified facilities are statistically segmented. The output includes the number of facilities, type distribution, proportion of status levels, size range, etc. within each segment.

[0104] As shown in Table 1, through the image acquisition, three - dimensional reconstruction, facility recognition, and fuzzy inference analysis process described in the present invention, various types of traffic safety facilities have been successfully identified on a 12 - kilometer expressway, and their three - dimensional size extraction and status level determination have been completed. The table shows partial segmented statistical results in units of 1 kilometer, including the number of facilities, average size, status level distribution (qualified, slight deviation, abnormal), and special remarks information.

[0105] Table 1 Statistical Table of Traffic Safety Facilities Segmentation

[0106]

[0107] Finally, it should be noted that the above - mentioned are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent substitution on some of the technical features. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A traffic safety facility statistical survey method based on deep learning, characterized in that, Including: Collect video frame image data, synchronously obtain GPS and IMU data of the camera device, calculate the relative pose transformation matrix between video frame images based on the GPS and IMU data and correct it, and obtain the corrected pose transformation matrix; Perform monocular depth estimation on each frame of video frame image to generate a dense depth map, combine the dense depth maps of multiple consecutive frames with the corresponding corrected pose transformation matrix for spatial alignment and fusion, construct a three-dimensional scene point cloud, and extract the geometric attribute information of traffic safety facilities; Use the corrected pose transformation matrix as the motion constraint for optical flow estimation, combine the inter-frame image features, perform image registration and motion compensation, and generate a de-blurred image; Use the YOLO object detection algorithm to identify traffic safety facilities in the de-blurred image, obtain the category of the facilities, and associate it with the geometric attribute information of the traffic safety facilities to generate structured data of traffic safety facilities; Perform fuzzy inference on the structured data of traffic safety facilities and the standard attributes in the traffic safety facility database to judge the status information of the facilities; Based on the structured data and status information, perform segmented statistics and classification summary of traffic safety facilities according to the preset road stake number range, and output a statistical report; The steps of image registration and motion compensation include: Based on the corrected pose transformation matrix, predict the overall motion trend between consecutive image frames, generate an initial displacement estimate of the image area, and use it as the initial input for optical flow calculation; Combine the feature point matching relationship between image frames, and use the LK optical flow algorithm based on the image pyramid structure to perform multi-scale optical flow estimation to obtain an optical flow vector field; During the calculation of the optical flow field, use the direction predicted by the corrected pose transformation matrix as the guiding information to constrain and weight-optimize the direction of the optical flow vector field; Based on the optimized optical flow vector field, perform pixel-level registration and motion compensation on the image frames to generate an image sequence without motion blur; The steps of generating structured data of traffic safety facilities include: Input the de-blurred image into the YOLO object detection network to obtain the bounding box position, class label and recognition confidence of traffic safety facilities; Perform spatial mapping of the recognized class label to the corresponding area in the three-dimensional scene point cloud, and extract the geometric attribute information of the facility according to the image coordinate and point cloud projection relationship; Construct structured data entries based on the class label and geometric attribute information, and associate the stake number range where each facility is located.

2. The traffic safety facility statistical survey method based on deep learning according to claim 1, characterized in that The steps of obtaining the corrected pose change include: Based on the three-axis angular velocity and acceleration data collected by the IMU sensor, use the integration method to estimate the rotation angle and displacement vector of the camera device between consecutive frame images, and construct a relative pose transformation matrix; Combine the GPS data to correct the relative pose transformation matrix to obtain the corrected pose transformation matrix.

3. A statistical survey method for traffic safety facilities based on deep learning according to claim 1, characterized in that, The steps of constructing a three-dimensional scene include: Input each frame of image into a convolutional neural network for monocular depth estimation to generate a dense depth map; Combine the corrected pose transformation matrix between adjacent frame images to perform spatial coordinate transformation on the corresponding dense depth map to achieve three-dimensional alignment of dense depth maps of different frames; Use a point cloud fusion algorithm to perform cumulative fusion on multiple frames of dense depth maps, including voxel filtering for noise reduction, point density equalization and edge-preserving enhancement, to generate a dense three-dimensional point cloud.

4. A traffic safety facility statistical survey method based on deep learning according to claim 1, characterized in that, The steps of fuzzy inference include: Set the facility category, dimension deviation, position deviation, and recognition confidence as input fuzzy variables, and define their linguistic values and membership functions; Construct a fuzzy rule base that contains the mapping relationship between input variables and output variables, where the output variable is the facility status level; Adopt a Mamdani-type fuzzy inference system and combine it with the maximum membership degree method for defuzzification processing. The output result is a level label that characterizes the facility status.

5. A traffic safety facility statistical survey method based on deep learning according to claim 1, characterized in that, The steps of sectional statistics and classification summary include: According to the spatial position information and mileage range in the structured data, divide all identified traffic safety facilities into pile segments at a set interval; Statistically count the quantity, dimension range, and status level of various facilities within each pile segment, and generate corresponding statistical tables.

Citation Information

Patent Citations

  • Depth map and IMU-based high-dynamic scene three-dimensional reconstruction method and system

    CN110310362A

  • Multi-sensor fusion dense VSLAM method

    CN118298262A