Deep learning-based traffic safety facility statistical survey method

Through a deep learning-based method, combined with GPS/IMU data and monocular depth estimation, high-precision three-dimensional reconstruction and state evaluation of road traffic safety facilities are achieved, and the problems of low efficiency, insufficient data and strong subjectivity of state judgment in the existing technology are solved, and the accurate identification and intelligent state judgment of the facilities are realized.

CN120107906AActive Publication Date: 2025-06-06JIANGXI HIGHWAY RES & DESIGN INST CO LTD

Patent Information

Application Number
CN202510579040.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

In the survey and statistics of road traffic safety facilities, the existing technology has problems such as low work efficiency, insufficient coverage, limited data dimensions, and strong subjectivity of state judgments. It is difficult to accurately obtain the three-dimensional size and image schema of the facilities and severe interference, and it is difficult to structure and fusion of the data.

Method used

Using a deep learning method, a corrected position transformation matrix is ​​constructed by collecting video images and synchronizing the acquisition of GPS/IMU data, and a modified position transformation matrix is ​​constructed, combined with monocular depth estimation to realize three-dimensional point cloud reconstruction, extract the geometric properties of the facility; using the YOLO algorithm for target recognition, and generating structured data; and a fuzzy inference system is introduced to judge the facility status, and statistics and summary outputs are performed according to the station number.

Benefits of technology

It significantly improves the spatial consistency and depth information alignment accuracy between multi-frame images, accurately extracts the three-dimensional geometric parameters of the facility, realizes a flexible and intelligent facility status evaluation mechanism, and supports segmented statistics and report output of facility status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107906A_ABST
    Figure CN120107906A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of traffic safety facility identification, in particular to a traffic safety facility statistical survey method based on deep learning. The method comprises the following steps: acquiring a video frame image and synchronously acquiring GPS and IMU data of camera equipment; constructing a corrected pose transformation matrix through a fusion pose estimation algorithm, generating a three-dimensional point cloud by combining monocular depth estimation and multi-frame depth map fusion, and extracting geometric attribute information of traffic safety facilities; identifying the traffic safety facilities in the image by using a YOLO algorithm, extracting the types and positions of the facilities, and associating the types and positions with three-dimensional attributes to generate structured data; performing grade judgment on the facility state based on a fuzzy reasoning method; and performing segmented statistics and classified summarization according to the road stake numbers, and outputting a statistical report. The method can realize facility three-dimensional perception, state intelligent evaluation and classified statistics, and is suitable for intelligent highway operation and maintenance and digital traffic facility management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of traffic safety facility identification, and in particular to a traffic safety facility statistical investigation method based on deep learning. Background Art

[0002] With the continuous development of transportation networks and the advancement of intelligent infrastructure construction, road traffic safety facilities (such as signs, guardrails, markings, delineators, etc.) are important components for ensuring driving safety and conveying road information. The rationality of their layout and good condition directly affect road traffic efficiency and driving safety. Therefore, the automatic identification, state perception and structured management of traffic safety facilities have become key links in highway maintenance and intelligent transportation systems.

[0003] At present, the investigation and statistics of road traffic safety facilities mainly rely on manual surveys or semi-automatic equipment collection, which has problems such as low work efficiency, insufficient coverage, limited data dimensions, and strong subjectivity in status judgment. Even in the scenario where image recognition or laser point cloud equipment is introduced, there are technical bottlenecks such as the difficulty in accurately obtaining the three-dimensional dimensions of the facilities, serious image smear interference, and difficulty in structured data fusion, which makes it difficult to meet the actual needs of the digital management of the new generation of smart transportation facilities. In addition, the existing technology often lacks accurate modeling and correction mechanisms for the spatial posture changes between the collected data, which makes it difficult to align multiple frames of images or depth information in space, affecting the accuracy of three-dimensional reconstruction. Image recognition results are also difficult to effectively associate with spatial structural information, limiting the ability to comprehensively judge the actual status of the facilities.

[0004] To this end, a traffic safety facility statistical survey method based on deep learning is proposed. Summary of the invention

[0005] The present invention provides a traffic safety facility statistical survey method based on deep learning. The method collects video images and synchronously obtains GPS / IMU data to construct a corrected pose transformation matrix, combines monocular depth estimation to achieve three-dimensional point cloud reconstruction, and extracts facility geometric attributes. The YOLO algorithm is used for target recognition to generate structured data. A fuzzy reasoning system is introduced to judge the status of the facility, and statistics and summary output are performed in sections according to the pile number.

[0006] To achieve the above object, the present invention provides the following technical solutions: A traffic safety facility statistical survey method based on deep learning, comprising: Collect video frame image data, synchronously obtain GPS and IMU data of the camera device, calculate and correct the relative posture transformation matrix between the video frame images based on the GPS and IMU data, and obtain the corrected posture transformation matrix; Perform monocular depth estimation on each video frame to generate a dense depth map; combine the dense depth maps of multiple consecutive frames with the corresponding corrected pose transformation matrix for spatial alignment and fusion, construct a three-dimensional scene point cloud, and extract the geometric attribute information of traffic safety facilities; The corrected pose transformation matrix is ​​used as the motion constraint for optical flow estimation, and combined with the inter-frame image features, image registration and motion compensation are performed to generate a de-smeared image. The YOLO target detection algorithm is used to identify the traffic safety facilities in the de-smeared image, obtain the category of the facilities, and associate them with the geometric attribute information of the traffic safety facilities to generate structured data of the traffic safety facilities; Perform fuzzy reasoning on the structured data of traffic safety facilities and the standard attributes in the traffic safety facility database to determine the status information of the facilities; Based on the structured data and status information, the traffic safety facilities are segmented and classified according to the preset road pile number range, and a statistical report is output.

[0007] Furthermore, the step of obtaining the corrected posture change includes: Based on the three-axis angular velocity and acceleration data collected by the IMU sensor, the rotation angle and displacement vector of the camera device between consecutive frame images are estimated using the integration method, and the relative posture transformation matrix is ​​constructed; Combined with GPS data, the relative posture transformation matrix is ​​corrected to obtain the corrected posture transformation matrix.

[0008] Furthermore, the steps of constructing a three-dimensional scene include: Input each frame of image into the convolutional neural network for monocular depth estimation to generate a dense depth map; Combined with the corrected pose transformation matrix between adjacent frame images, the corresponding dense depth map is transformed into spatial coordinates to achieve three-dimensional alignment of dense depth maps of different frames. A point cloud fusion algorithm is used to accumulate and fuse multiple frames of dense depth maps, including voxel filtering denoising, point density balancing and edge preservation enhancement, to generate a dense 3D point cloud.

[0009] Furthermore, the steps of image registration and motion compensation include: Based on the corrected pose transformation matrix, the overall motion trend between consecutive image frames is predicted to generate an initial displacement estimate of the image region as the initial input for optical flow calculation; Combined with the feature point matching relationship between image frames, the LK optical flow algorithm based on the image pyramid structure is used to perform multi-scale optical flow estimation to obtain the optical flow vector field; In the process of calculating the optical flow field, the direction predicted by the corrected pose transformation matrix is ​​used as the guiding information to constrain and weight the direction of the optical flow vector field. Based on the optimized optical flow vector field, the image frames are registered at the pixel level and motion compensated to generate an image sequence without motion artifacts.

[0010] Furthermore, the step of generating traffic safety facility structured data includes: The de-smeared image is input into the YOLO target detection network to obtain the bounding box position, category label and recognition confidence of the traffic safety facility; Spatial mapping is performed between the identified category labels and the corresponding areas in the 3D scene point cloud, and the geometric attribute information of the facility is extracted based on the relationship between the image coordinates and the point cloud projection; Structured data entries are constructed based on category labels and geometric attribute information, and the station range in which each facility is located is associated.

[0011] Furthermore, the steps of fuzzy reasoning include: Set facility category, size deviation, location deviation and recognition confidence as input fuzzy variables, and define their linguistic values ​​and membership functions; Construct a fuzzy rule base containing the mapping relationship between input variables and output variables, where the output variable is the facility status level; The Mamdani type fuzzy inference system is used and combined with the maximum membership method for defuzzification, and the output result is a grade label that characterizes the status of the facility.

[0012] Furthermore, the steps of segmented statistics and classified aggregation include: Based on the spatial location information and pile number range in the structured data, all identified traffic safety facilities are divided into pile sections according to the set spacing; The number, size range and status level of various facilities in each pile segment are counted, and corresponding statistical tables are generated.

[0013] The beneficial effects of the present invention are: 1. The present invention fuses IMU inertial data and GPS position information to construct a corrected pose transformation matrix between continuous image frames. Compared with the traditional registration method based only on image features, it can effectively reduce the errors caused by integral drift and image jitter, and significantly improve the spatial consistency and depth information alignment accuracy between multiple frame images, providing a stable spatial foundation for subsequent 3D reconstruction and geometric extraction.

[0014] 2. This invention introduces a monocular depth estimation network and combines it with posture transformation to achieve multi-frame dense depth map fusion, generate a dense point cloud model, and accurately extract the length, width, height, clearance and other three-dimensional geometric parameters of the facility. Compared with traditional image recognition that can only output two-dimensional position and category, it improves the spatial dimension and quantification degree of facility information, providing key support for facility modeling, deviation judgment and digital operation and maintenance.

[0015] 3. Based on the information of traffic facility type, size deviation, position deviation and recognition confidence, the present invention constructs a fuzzy rule base and uses the Mamdani reasoning system to grade the facility status, overcoming the problems of threshold rigidity and error sensitivity in traditional judgment methods and realizing a flexible and intelligent status evaluation mechanism. At the same time, in conjunction with the pile number division logic, it supports segmented statistics and report output of facility status, and has strong engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a flow chart of a traffic safety facility statistical survey method based on deep learning provided by the present invention. DETAILED DESCRIPTION

[0017] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0018] Embodiment 1 A traffic safety facility statistical survey method based on deep learning, such as Figure 1 As shown, including: S100: Collecting video frame image data, synchronously acquiring GPS and IMU data of the camera device, calculating and correcting the relative posture transformation matrix between the video frame images based on the GPS and IMU data, and obtaining the corrected posture transformation matrix; Furthermore, the step of obtaining the corrected posture change includes: Based on the three-axis angular velocity and acceleration data collected by the IMU sensor, the rotation angle and displacement vector of the camera device between consecutive frame images are estimated using the integration method, and the relative posture transformation matrix is ​​constructed; In combination with the GPS data, the relative posture transformation matrix is ​​corrected to obtain a corrected posture transformation matrix.

[0019] Specifically, continuous image frames are collected by a camera installed on a mobile acquisition platform (such as a vehicle), and at the same time, the three-axis angular velocity and three-axis acceleration collected by the IMU, as well as the three-dimensional position data and absolute timestamp provided by the GPS, are obtained. First, based on the IMU angular velocity data, the direction cosine matrix integration method is used to calculate the relative rotation change of the camera between two frames of images. ,in Represents two adjacent frames, then the acceleration data is double integrated to obtain the translation vector The rotation change and translation vector constitute the relative posture transformation matrix between image frames : ; Since IMU integration has time accumulation error, in order to improve the accuracy of pose estimation, GPS data is introduced to correct the pose estimated by the above IMU. The correction process includes: time synchronization and interpolation alignment of GPS and IMU data; using extended Kalman filter to compensate for the position deviation of translation. The corrected result constitutes the final corrected pose transformation matrix used for image alignment and point cloud fusion.

[0020] By fusing GPS and IMU sensor data, constructing and correcting the relative pose transformation matrix between video frame images, the accuracy and stability of pose estimation between image frames can be effectively improved. The real-time inference of the camera's continuous motion trajectory is achieved through the high-frequency inertial data of the IMU, and then combined with the absolute position and attitude reference information provided by the GPS for correction, which can effectively suppress the cumulative drift error caused by IMU integration, improve the robustness of pose estimation in complex dynamic environments, provide accurate spatial references for subsequent depth map alignment, three-dimensional scene reconstruction and image registration, and enhance the environmental perception ability and engineering application reliability of the entire system.

[0021] S200: performing monocular depth estimation on each video frame image to generate a dense depth map; combining the dense depth maps of multiple consecutive frames with the corresponding corrected pose transformation matrix to perform spatial alignment and fusion, constructing a three-dimensional scene point cloud, and extracting geometric attribute information of traffic safety facilities; Furthermore, the steps of constructing a three-dimensional scene include: Input each frame of image into the convolutional neural network for monocular depth estimation to generate a dense depth map; Combined with the corrected pose transformation matrix between adjacent frame images, the corresponding dense depth map is transformed into spatial coordinates to achieve three-dimensional alignment of dense depth maps of different frames. A point cloud fusion algorithm is used to accumulate and fuse multiple frames of dense depth maps, including voxel filtering denoising, point density balancing and edge preservation enhancement, to generate a dense 3D point cloud.

[0022] Specifically, each frame of the collected video image is input into a pre-trained convolutional neural network model for monocular depth estimation. The depth estimation network adopts a deep learning model based on an encoder-decoder structure, such as Monodepth, DPT or its improved version, and outputs a dense depth map with the same size as the original image. In the dense depth map, each pixel contains the relative depth value of the scene point corresponding to the pixel. Then, combined with the corrected pose transformation matrix corresponding to each frame of the image, the three-dimensional spatial coordinate transformation of each frame of the depth map is performed. Specifically, each depth pixel (u, v, d) is back-projected into a three-dimensional point (X, Y, Z) in the camera coordinate system through the camera intrinsic parameter matrix, and it is uniformly transformed to the common world coordinate system through the corrected pose transformation matrix to complete the three-dimensional alignment of multiple frames of depth maps. The back-projection transformation process is: ; Among them, X, Y and Z represent the coordinates of the X-axis, Y-axis and Z-axis in the three-dimensional space coordinates, u, v and d represent the row index, column index and depth value of the image pixel, and K represents the camera intrinsic parameter matrix. After that, the point cloud fusion algorithm is used to fuse the multi-frame point cloud data after three-dimensional alignment, including voxel grid filtering to reduce noise and redundant data, point density balancing to improve spatial consistency, and edge preservation algorithm to enhance the contour details of the facility, and finally generate a dense and continuous three-dimensional point cloud model.

[0023] Through this step, high-precision, multi-frame fusion three-dimensional scene reconstruction can be achieved, which not only effectively overcomes the instability of single-frame depth estimation, but also improves the accuracy of geometric attribute extraction and spatial positioning of traffic safety facilities, providing a reliable three-dimensional structure basis for subsequent facility identification and status assessment.

[0024] S300: Using the corrected posture transformation matrix as a motion constraint for optical flow estimation, combining inter-frame image features, performing image registration and motion compensation, and generating a de-smeared image; Furthermore, the steps of image registration and motion compensation include: Based on the corrected pose transformation matrix, the overall motion trend between consecutive image frames is predicted to generate an initial displacement estimate of the image region as the initial input for optical flow calculation; Combined with the feature point matching relationship between image frames, the LK optical flow algorithm based on the image pyramid structure is used to perform multi-scale optical flow estimation to obtain the optical flow vector field; In the process of calculating the optical flow field, the direction predicted by the corrected pose transformation matrix is ​​used as the guiding information to constrain and weight the direction of the optical flow vector field. Based on the optimized optical flow vector field, the image frames are registered at the pixel level and motion compensated to generate an image sequence without motion artifacts.

[0025] Specifically, based on the corrected posture transformation matrix obtained in the previous step, the motion trend between two consecutive frames of images is preliminarily modeled. Through the rotation information and displacement information contained in the matrix, the overall motion direction and speed of different regions in the image frame can be predicted, which is used to estimate the initial displacement vector of the pixel block or key point as the initial value of the optical flow estimation. Then, the inter-frame image is subjected to feature point extraction and matching to obtain the corresponding feature point set, which is used to construct the matching window in the optical flow estimation. The LK optical flow algorithm based on the image pyramid structure is used to estimate the multi-scale optical flow of the continuous frame images, and the solution is progressively solved from low resolution to high resolution layer by layer, and finally the optical flow vector field F(x,y) of the entire image is obtained. The vector field represents the displacement direction and amplitude of each pixel in the previous and next frame images. In the calculation process of the optical flow vector field, the predicted displacement direction in the corrected posture transformation matrix is ​​used as the guiding information, and the direction consistency is checked with the optical flow direction obtained locally in the image; in the weak texture area or fast motion scene, the optical flow field is directional regularized and noise suppressed by weighted optimization, thereby improving the global stability and local robustness of the optical flow estimation. Finally, based on the optimized optical flow vector field, the image frames are pixel-level registered and resampled, and the image blur and smear areas caused by camera motion are compensated for and the brightness is adjusted to generate a clear image sequence without motion smear.

[0026] By introducing the fused corrected pose transformation matrix into the optical flow estimation process, not only the stability and accuracy of the traditional Lucas-Kanade optical flow algorithm in complex scenes are improved, but also the directional consistency of image registration is significantly enhanced. Combining the multi-scale pyramid structure with the directional regularization optimization mechanism, it can maintain a good pixel-level registration effect under conditions of fast motion, occlusion, weak texture, etc. The final output de-smeared image provides high-quality input guarantee for subsequent YOLO recognition and geometric attribute extraction, enhancing the reliability and engineering practicality of the entire system.

[0027] S400: using the YOLO target detection algorithm to identify the traffic safety facilities in the de-smeared image, obtain the category of the facilities, and associate them with the geometric attribute information of the traffic safety facilities to generate structured data of the traffic safety facilities; Furthermore, the step of generating traffic safety facility structured data includes: The de-smeared image is input into the YOLO target detection network to obtain the bounding box position, category label and recognition confidence of the traffic safety facility; Spatial mapping is performed between the identified category labels and the corresponding areas in the 3D scene point cloud, and the geometric attribute information of the facility is extracted based on the relationship between the image coordinates and the point cloud projection; Structured data entries are constructed based on category labels and geometric attribute information, and the station range in which each facility is located is associated.

[0028] Specifically, the image frame that has completed motion compensation and smear removal is input into the pre-trained YOLO target detection model. The YOLO model can be YOLO v5, YOLO v7, YOLO v8 or other target detection architectures with lightweight deployment capabilities. The model output includes the bounding box position, category label and recognition confidence. Then, the bounding box area of ​​each target in the image is spatially corresponded to the three-dimensional point cloud data. Specifically, based on the internal parameter consistency of the image and the point cloud, the depth map or point cloud corresponding to the pixel point in the bounding box is screened and back-projected to establish a mapping relationship from the image detection result to the three-dimensional point cloud. Through this mapping, the image area corresponding to the YOLO detection result is converted into a three-dimensional point cloud fragment, so that the actual geometric attribute information of the facility can be extracted, including size parameters (length, width, height), spatial position (X, Y, Z), installation height, pitch angle, etc. Finally, the structured data entry of the traffic safety facility is constructed by combining the recognized category label, recognition confidence and geometric attribute information. Structured entries may include the following fields: facility category (such as speed limit signs, guardrails, delineators, etc.), spatial location (such as center coordinates, pile number), size parameters (such as width, height, clearance), identification confidence, and road pile number range (calculated based on GPS coordinates and road matching).

[0029] Through this step, not only can the efficient recognition of traffic safety facilities from images be achieved, but also the spatial fusion of recognition results and three-dimensional geometric information is completed, thereby constructing a structured data format with practical usability. This method takes into account both visual intelligence and spatial accuracy, and provides a solid data foundation for the digital management and status assessment of facility assets.

[0030] S500: Perform fuzzy reasoning on the traffic safety facility structured data and the standard attributes in the traffic safety facility database to determine the facility status information; Furthermore, the steps of fuzzy reasoning include: Set facility category, size deviation, location deviation and recognition confidence as input fuzzy variables, and define their linguistic values ​​and membership functions; Construct a fuzzy rule base containing the mapping relationship between input variables and output variables, where the output variable is the facility status level; The Mamdani type fuzzy inference system is used and combined with the maximum membership method for defuzzification, and the output result is a grade label that characterizes the status of the facility.

[0031] Specifically, multiple input fuzzy variables are set for judging the status of facilities, including but not limited to: facility category (such as speed limit signs, corrugated guardrails, crossing posts, etc.), size deviation (relative error between the detected size and the standard size), position deviation (the distance from the actual installation position of the facility to the designed position) and recognition confidence (confidence score output by the YOLO model). For each input variable, the corresponding linguistic variable value and membership function are defined. For example, the linguistic values ​​of size deviation include "small", "normal" and "large", and its membership function can adopt a triangular or trapezoidal function; the linguistic values ​​of position deviation include "small", "moderate" and "large"; the linguistic values ​​of recognition confidence include "low", "medium" and "high". Taking size deviation as an example, its fuzzy membership function can be set as: ; Secondly, a fuzzy rule base is constructed. The rules are mapped from the combination of input variables to the output variable "facility status level", and its language values ​​can be: "qualified", "acceptable", "abnormal", "severe abnormality". Example rules include: IF size deviation is NORMAL AND position deviation is MODERATE AND confidence is HIGH THEN status is ACCEPTED; IF the size deviation is too large OR the position deviation is too large THEN the status is acceptable; IF size deviation is too large AND position deviation is too large AND confidence is low THEN status is abnormal. Then, the Mamdani type fuzzy inference system is used to perform fuzzy logic calculation on the above rules, and the inference engine uses the minimum-maximum synthesis method to superimpose and aggregate multiple rules. Finally, the maximum membership method is used for defuzzification, and the final facility status level label is extracted from the output fuzzy set as the evaluation result of the facility. The label can be selected as: "qualified", "slight deviation", "unqualified" and other status identifiers.

[0032] Through this fuzzy reasoning system, under conditions of high uncertainty in facility identification and complex environmental interference, the facility status can be flexibly judged by integrating multi-dimensional attribute indicators, avoiding misjudgment or missed judgment due to threshold judgment, improving the intelligence and fault tolerance of facility evaluation, and providing more stable and reasonable data support for subsequent statistics, maintenance and management.

[0033] S600: Based on the structured data and status information, traffic safety facilities are segmented and classified according to a preset road pile number range, and a statistical report is output.

[0034] Furthermore, the steps of segmented statistics and classified aggregation include: Based on the spatial location information and pile number range in the structured data, all identified traffic safety facilities are divided into pile sections according to the set spacing; The number, size range and status level of various facilities in each pile segment are counted, and corresponding statistical tables are generated.

[0035] Specifically, according to the spatial location information (GPS coordinates or coordinates of the center point of the three-dimensional point cloud) of each traffic safety facility in the structured data and the road name or number identified by it, the location of the facility in the actual road is mapped to the pile number. The pile number range or pile number point value of each facility can be determined by consulting the standard road pile number database or projecting the GPS coordinates into the existing road linear reference model. Then, according to the set statistical interval (such as one pile segment every 100 meters or 500 meters), the road is divided into equidistant segments according to the pile number. Each segment is a statistical unit corresponding to a set of facility data. For each pile segment, based on the structured data content, the following information is counted: the number of various types of traffic safety facilities (such as signs, guardrails, markings, etc.), the size range of each type of facility (maximum value, minimum value, average value), and the distribution of status levels (such as the proportion of qualified, slightly abnormal, and severely abnormal). The statistical process can be achieved by traversing the structured data and performing clustering operations according to the pile numbers, supporting online processing or offline batch processing. Optionally, the system can also sort, filter, and mark abnormalities for the statistical results.

[0036] Through the segmented statistical mechanism based on pile numbers, accurate classification and summary of traffic safety facilities in different road sections can be achieved, which is convenient for unified management and horizontal comparative analysis of data of various facilities. The output results not only meet the technical evaluation requirements, but also have good engineering interface compatibility, and are suitable for direct integration and calling of road traffic management systems, infrastructure maintenance systems and smart transportation platforms.

[0037] Embodiment 2 In this embodiment, a 12-kilometer highway test section is selected, along which a variety of traffic safety facilities are installed, including speed limit signs, corrugated guardrails, markings, contour markers, etc. A vehicle platform equipped with a front monocular camera, a GPS positioning module, and an IMU inertial sensor is used to collect data at a normal speed (about 80 km / h).

[0038] During the acquisition process, the camera continuously records video images at a frequency of 20 frames / second, the IMU synchronously collects three-axis angular velocity and acceleration data, and the GPS module provides geographic coordinates and timestamp information.

[0039] In the post-processing stage, the relative rotation and displacement estimation results between every two frames of images are first obtained based on the IMU data integration, and the accumulated drift is corrected through the GPS coordinates to form a corrected posture transformation matrix between consecutive image frames.

[0040] Then, a deep neural network (such as Monodepth2) is used to perform monocular depth estimation on each frame of the image to generate a dense depth map; by coordinating with the corrected pose transformation matrix, multiple depth maps are spatially aligned in the same coordinate system and fused into a three-dimensional scene point cloud.

[0041] At the same time, the YOLOv5 model is used to perform target detection on the image sequence after removing the smear, identify the categories of facilities such as speed limit signs and guardrails and their bounding box information, and perform back-projection in combination with the point cloud to extract their actual spatial size, position, installation height and other geometric attributes.

[0042] Then, the identified structured data is used as input variables, and the Mamdani type fuzzy inference system is used to input the size deviation, position deviation and recognition confidence, and output the facility status level. The facility status results include three categories: "qualified", "slight deviation" and "unqualified".

[0043] Finally, the identified facilities are segmented and counted according to the rule that every 500 meters is a pile segment. The output includes the number of facilities in each segment, type distribution, status level ratio, size range, etc.

[0044] As shown in Table 1, through the image acquisition, 3D reconstruction, facility identification and fuzzy reasoning analysis process described in the present invention, various types of traffic safety facilities were successfully identified on a 12-kilometer highway, and their 3D size extraction and status level determination were completed. The table shows some segmented statistical results per kilometer, including the number of facilities, average size, status level distribution (qualified, slightly deviated, abnormal) and special remarks.

[0045] Table 1 Traffic safety facilities segment statistics

[0046] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A traffic safety facility statistical survey method based on deep learning, characterized in that: include: Collect video frame image data, synchronously obtain GPS and IMU data of the camera device, calculate and correct the relative posture transformation matrix between the video frame images based on the GPS and IMU data, and obtain the corrected posture transformation matrix; Perform monocular depth estimation on each video frame to generate a dense depth map. Combine the dense depth maps of multiple consecutive frames with the corresponding corrected pose transformation matrix for spatial alignment and fusion to construct a 3D scene point cloud and extract the geometric attribute information of traffic safety facilities. The corrected pose transformation matrix is ​​used as the motion constraint for optical flow estimation, and combined with the inter-frame image features, image registration and motion compensation are performed to generate a de-smeared image. The YOLO target detection algorithm is used to identify the traffic safety facilities in the de-smeared image, obtain the category of the facilities, and associate them with the geometric attribute information of the traffic safety facilities to generate structured data of the traffic safety facilities; Perform fuzzy reasoning on the structured data of traffic safety facilities and the standard attributes in the traffic safety facility database to determine the status information of the facilities; Based on the structured data and status information, the traffic safety facilities are segmented and classified according to the preset road pile number range, and a statistical report is output.

2. The method for statistical investigation of traffic safety facilities based on deep learning according to claim 1, characterized in that: The steps to obtain the corrected pose change include: Based on the three-axis angular velocity and acceleration data collected by the IMU sensor, the rotation angle and displacement vector of the camera device between consecutive frame images are estimated using the integration method, and the relative posture transformation matrix is ​​constructed; Combined with GPS data, the relative posture transformation matrix is ​​corrected to obtain the corrected posture transformation matrix.

3. The method for statistical investigation of traffic safety facilities based on deep learning according to claim 1, characterized in that: The steps to build a 3D scene include: Input each frame of image into the convolutional neural network for monocular depth estimation to generate a dense depth map; Combined with the corrected pose transformation matrix between adjacent frame images, the corresponding dense depth map is transformed into spatial coordinates to achieve three-dimensional alignment of dense depth maps of different frames. A point cloud fusion algorithm is used to accumulate and fuse multiple frames of dense depth maps, including voxel filtering denoising, point density balancing and edge preservation enhancement, to generate a dense 3D point cloud.

4. The method for statistical investigation of traffic safety facilities based on deep learning according to claim 1, characterized in that: The steps of image registration and motion compensation include: Based on the corrected pose transformation matrix, the overall motion trend between consecutive image frames is predicted to generate an initial displacement estimate of the image region as the initial input for optical flow calculation; Combined with the feature point matching relationship between image frames, the LK optical flow algorithm based on the image pyramid structure is used to perform multi-scale optical flow estimation to obtain the optical flow vector field; In the process of calculating the optical flow field, the direction predicted by the corrected pose transformation matrix is ​​used as the guiding information to constrain and weight the direction of the optical flow vector field. Based on the optimized optical flow vector field, the image frames are registered at the pixel level and motion compensated to generate an image sequence without motion artifacts.

5. The method for statistical investigation of traffic safety facilities based on deep learning according to claim 1, characterized in that: The steps to generate structured data of traffic safety facilities include: The de-smeared image is input into the YOLO target detection network to obtain the bounding box position, category label and recognition confidence of the traffic safety facility; Spatial mapping is performed between the identified category labels and the corresponding areas in the 3D scene point cloud, and the geometric attribute information of the facility is extracted based on the relationship between the image coordinates and the point cloud projection; Structured data entries are constructed based on category labels and geometric attribute information, and the station range in which each facility is located is associated.

6. The method for statistical investigation of traffic safety facilities based on deep learning according to claim 1, characterized in that: The steps of fuzzy reasoning include: Set facility category, size deviation, location deviation and recognition confidence as input fuzzy variables, and define their linguistic values ​​and membership functions; Construct a fuzzy rule base containing the mapping relationship between input variables and output variables, where the output variable is the facility status level; The Mamdani type fuzzy inference system is used and combined with the maximum membership method for defuzzification, and the output result is a grade label that characterizes the status of the facility.

7. The method for statistical investigation of traffic safety facilities based on deep learning according to claim 1, characterized in that: The steps of segmented statistics and classification summary include: Based on the spatial location information and pile number range in the structured data, all identified traffic safety facilities are divided into pile sections according to the set spacing; The number, size range and status level of various facilities in each pile segment are counted, and corresponding statistical tables are generated.

Citation Information

Patent Citations

  • Depth map and IMU-based high-dynamic scene three-dimensional reconstruction method and system

    CN110310362A

  • Multi-sensor fusion dense VSLAM method

    CN118298262A

  • IMU (Inertial Measurement Unit)-assisted moving target detection method under dynamic condition

    CN119273715A

  • IMU (Inertial Measurement Unit) and RGB-D (Red Green Blue-Depth) fusion driven real-time three-

    CN119399363A

  • Joint learning of geometry and motion with three-dimensional holistic understanding

    US20200211206A1

Cited By

  • Video fusion method and device based on three-dimensional scene and storage medium

    CN121052994A