Method and system for sensing subsidence and jolt of road surface

An AI-driven method processes vehicle camera data to automate road surface distress detection, addressing inefficiencies and inaccuracies in existing systems by using existing vehicle cameras and neural networks for accurate and cost-effective road condition analysis.

CN120318793AActive Publication Date: 2025-07-15SHANGHAI TONGLU CLOUD TRANSPORTATION TECH CO LTD

Patent Information

Application Number
CN202510804456.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The existing operating vehicle dash recorder cannot effectively collect acceleration data, resulting in low automation of road subsidence and bump disease analysis, high hardware cost, complex data fusion, and video object detection is susceptible to occlusion, resulting in insufficient misjudgment and insufficient accuracy.

Method used

Using the existing operating vehicle dash recorder video, single-frame recognition is performed through frame extraction and pre-trained neural network models, point coordinates, target objects and scene categories are extracted, combined with timing data analysis, abnormal bumps are judged and stored in a graded manner, so as to realize road surface disease perception without additional sensors.

Benefits of technology

It improves the degree of automation and accuracy of road subsidence and bump analysis, reduces hardware costs and manual intervention, and realizes efficient data collection and storage, which facilitates subsequent data query and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318793A_ABST
    Figure CN120318793A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of road engineering, highway engineering and machine vision, in particular to a method and system for sensing subsidence and jolt of a road surface, and the method comprises the steps: obtaining a video collected by an operating vehicle driving recorder, carrying out the frame extraction of the video, and obtaining an image sequence; performing single-frame identification on the image sequence based on a pre-trained neural network model, extracting driving information of an operating vehicle in each frame of image, vanishing point coordinates, a target object and a scene category, and performing summarization and sorting to obtain time sequence data; slicing the time sequence data, analyzing the position change condition of the vanishing point in the running process of the operating vehicle based on the sliced data, and judging abnormal bumping according to the position change condition of the vanishing point; and grading the abnormal jolts, establishing association between the related identification data and the abnormal jolts based on the time sequence data, and storing the associated identification data and the abnormal jolts in a database. According to the invention, the image analysis technology is utilized to realize bumping analysis of the patrol vehicle, so that pavement subsidence and abnormal bumping points can be sensed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of road engineering, highway engineering, and machine vision technology, and particularly relates to a method and system for sensing pavement subsidence and bumps. Background Art

[0002] Pavement subsidence and bumps are common disease problems in urban roads, which seriously affect driving safety and comfort. Timely sensing and discovery of pavement subsidence and bumps, and early maintenance and treatment are important tasks in current road management. The traditional sensing of pavement subsidence and bumps mainly relies on manual measurement methods and intelligent inspection equipment measurement methods. Although the accuracy can be guaranteed, the automation degree is not high, and it consumes manpower and material resources. With the development of information technology, driving recorders are generally installed on operating vehicles. These devices can collect video data in real time during vehicle driving, and record information such as time and geographical location. Compared with arranging special inspection vehicles and personnel to perform operations, operating vehicles have a wide coverage, high monitoring frequency, and low cost. However, there are the following problems if operating vehicles are used to collect pavement diseases such as bumps and subsidence:

[0003] (1) Limitations of the driving recorder function: Currently, most driving recorders installed on operating vehicles generally only collect video data and do not collect acceleration data. In addition, positioning, speed, and time information are usually embedded in the video in text form, rather than stored in independent data forms. Therefore, pavement bump data processing algorithms such as laser and acceleration are difficult to use in this case.

[0004] (2) Challenges of hardware cost and data fusion: If new sensors (such as laser and acceleration sensors) are specially installed on operating vehicles, on the one hand, it will increase the hardware cost, and on the other hand, it will also involve the problems of fusion storage and parsing of existing video data from different manufacturers.

[0005] (3) Challenges of data processing and analysis: If only acceleration sensors are installed on operating vehicles, and based on acceleration data to analyze the subsidence and abnormal bumps passed during vehicle driving, there may be a lack of in-depth analysis of the disease causes, resulting in misjudgment. For example, simply relying on acceleration signals may not be able to distinguish whether the bump is caused by pavement diseases or vehicle self-vibration.

[0006] If video analysis is added on the basis of acceleration analysis, it is necessary to synchronize the timing of acceleration data, video, longitude and latitude, etc., adjust the sampling interval, and perform data slicing adjustment. For example, when abnormal bumps occur in acceleration data, the video image collected is already an image after the abnormal bumps, and the time difference between the two is related to the installation positions of the image sensor and the acceleration sensor, which will significantly increase the complexity of the processing program.

[0007] (4) Limitations of existing video target detection and analysis: In most current analyses of road surface subsidence and bumps, image detectors are mainly used to detect objects (such as manhole covers, speed bumps, potholes, etc.) in the video that may cause vehicle bumps. However, in detector recognition, missed detections may occur due to occlusion (such as vehicle or pedestrian occlusion), insufficient detector performance, etc., thus affecting the accuracy of cause analysis.

[0008] From the above analysis, it can be seen that although using operating vehicles to collect road surface disease data has broad application prospects, it still faces many technical challenges and limitations in practical applications. Therefore, the present invention proposes a method and system for sensing road surface subsidence and bumps. Summary of the Invention

[0009] The purpose of the present invention is to provide a method and system for sensing road surface subsidence and bumps. Without adding additional hardware sensors, the videos collected by the current driving recorders of operating vehicles are analyzed, that is, AI image analysis technology is used to realize the bump analysis of the patrol vehicles, and then the road surface subsidence and abnormal bump points are sensed, and the relevant diseases are classified and structurally stored according to the bumps.

[0010] To achieve the above purpose, on the one hand, the present invention provides a method for sensing road surface subsidence and bumps, including:

[0011] Obtain the video collected by the driving recorder of the operating vehicle, extract frames from the video, and obtain an image sequence;

[0012] Based on a pre-trained neural network model, perform single-frame recognition on the image sequence, extract the driving information of the operating vehicle, vanishing point coordinates, objects, and scene categories in each frame of the image, and summarize and sort them to obtain time-series data;

[0013] Slice the time-series data, analyze the change of the position of the vanishing point during the driving process of the operating vehicle based on the sliced data, and judge abnormal bumps according to the change of the position of the vanishing point;

[0014] Classify the abnormal bumps, and establish an association between the relevant recognition data and the abnormal bumps based on the time-series data, and store them in the database.

[0015] Optionally, the conditions for obtaining the video collected by the driving recorder of the operating vehicle are:

[0016] The video clarity meets the preset range, the proportion of the road surface area in the video meets the preset proportion, the installation method, installation position, and angle of the driving recorder meet the preset requirements, and there are clear and fixed time information, speed information, longitude and latitude information in the video.

[0017] Optionally, perform single-frame recognition on the image sequence based on a pre-trained neural network model, and extract the driving information of the operating vehicle in each frame of the image, as well as the vanishing point coordinates, target objects, and scene categories, including:

[0018] Perform single-frame recognition on the image sequence based on a pre-trained first neural network model, and extract the vanishing point coordinates, target objects, and scene categories;

[0019] Perform single-frame recognition on the image sequence based on a pre-trained second neural network model, and extract the time information, speed information, longitude and latitude information.

[0020] Optionally, the first neural network model includes: a recognition backbone network, a first branch head network, a second branch head network, and a third branch head network. The recognition backbone network is used to perform multi-scale feature extraction on the input image and output several feature maps; the first branch head network is used to recognize the height of the vanishing point in the feature map, the second branch head network is used to recognize the target objects and target object types that cause bumps in the feature map, and the third branch head network is used to recognize the scene type of the image acquisition in the feature map.

[0021] Optionally, the second neural network model uses a neural network for text OCR recognition.

[0022] Optionally, slice the time series data, analyze the position change of the vanishing point during the driving process of the operating vehicle based on the sliced data, and judge abnormal bumps according to the position change of the vanishing point, including:

[0023] Slice the time series data at a preset first time interval to obtain several first sliced data;

[0024] Extract the vanishing point height time series data and speed time series data from the first sliced data, and calculate the average speed;

[0025] Use the least squares method to fit the vanishing point height time series data, obtain the fitting line and calculate the vanishing point deviation, obtain the vanishing point deviation time series data and determine the average bump condition by calculating the root mean square;

[0026] Slice the vanishing point deviation time series data at a preset second time interval to obtain several second sliced data, and select the second sliced data with the largest root mean square as the short-term maximum bump;

[0027] Introduce average speed correction, calculate the bump ratio based on the average bump condition and the short-term maximum bump, and compare the bump ratio with a preset bump threshold to determine the abnormal bump.

[0028] Optionally, an average speed correction is introduced, and a bump ratio is calculated based on the average bump condition and the short-term maximum bump, including:

[0029] ;

[0030] where R1 is the bump ratio, v is the average speed of the first slice of data, alpha is a preset value set according to the vehicle type, rms1_sub_max is the short-term maximum bump, and rms1 is the average bump condition.

[0031] On the other hand, the present invention also provides a perception system for road surface subsidence bumps, including:

[0032] A video acquisition module for acquiring the video collected by the driving recorder of the operating vehicle, extracting frames from the video, and obtaining an image sequence;

[0033] A video analysis module for performing single-frame recognition on the image sequence based on a pre-trained neural network model, extracting the driving information of the operating vehicle, as well as the vanishing point coordinates, target objects, and scene categories in each frame of the image, and summarizing and sorting them to obtain time-series data;

[0034] A bump analysis module for slicing the time-series data, analyzing the position change of the vanishing point during the driving process of the operating vehicle based on the sliced data, and judging abnormal bumps according to the position change of the vanishing point;

[0035] A hierarchical storage module for classifying the abnormal bumps, and establishing an association between the relevant recognition data and the abnormal bumps based on the time-series data, and storing them in a database.

[0036] The beneficial effects of the present invention are as follows:

[0037] The present invention utilizes the existing driving recorder of the operating vehicle, without the need for additional installation of sensors, greatly reducing the complexity and cost of data acquisition. Through video decoding, frame extraction, and analysis by the neural network model, it can efficiently identify and analyze road surface subsidence and bumps, timely detect road diseases, and improve the real-time performance and accuracy of data acquisition. The whole process of the present invention from video decoding, frame extraction, image recognition to data parsing and storage in the database is highly automated, reducing manual intervention and improving work efficiency; through the multi-branch head design of the neural network model, it can simultaneously perform vanishing point height detection, target detection, and scene classification, improving the accuracy and efficiency of recognition; and the recognition results are stored in a structured manner, facilitating subsequent data query and analysis. By using a vector database, specific types of bump subsidence data can be quickly retrieved, improving the availability of data. Description of the Drawings

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 It is a flowchart of a method for perceiving road surface subsidence and bumps in an embodiment of the present invention;

[0040] Figure 2 It is a flowchart of video frame extraction analysis in an embodiment of the present invention;

[0041] Figure 3 It is a schematic diagram of vanishing point height inference in an embodiment of the present invention;

[0042] Figure 4 It is a schematic diagram for determining whether the recognized target intersects with the driving area in an embodiment of the present invention;

[0043] Figure 5 It is a flowchart of time series data analysis in an embodiment of the present invention;

[0044] Figure 6 It is a schematic diagram of the transformation relationship between road surface bumps and image vanishing points in an embodiment of the present invention. Specific Embodiments

[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0046] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0047] Embodiment 1:

[0048] This embodiment provides a method for perceiving road surface subsidence and bumps, including:

[0049] Obtain the video collected by the driving recorder of the operating vehicle, extract frames from the video, and obtain an image sequence;

[0050] Based on a pre-trained neural network model, perform single-frame recognition on the image sequence, extract the driving information, vanishing point coordinates, target objects, and scene categories of the operating vehicle in each frame of the image, and summarize and sort them to obtain time series data;

[0051] Slice the time series data, analyze the position change of the vanishing point during the driving of the operating vehicle based on the sliced data analysis, and judge abnormal bumps according to the position change of the vanishing point;

[0052] Classify the abnormal bumps, and establish an association between the relevant recognition data and the abnormal bumps based on the time series data, and store them in the database.

[0053] Specifically, this embodiment uses the existing on-vehicle driving recorder of the operating vehicle, without the need to install additional sensors, which greatly reduces the complexity and cost of data collection. Through video decoding, frame extraction, and analysis by a neural network model, it can efficiently identify and analyze road surface subsidence and bumps, timely detect road surface diseases, and improve the real-time performance and accuracy of data collection. In this embodiment, from video decoding, frame extraction, image recognition to data parsing and storage, the whole process is highly automated, reducing manual intervention and improving work efficiency; through the multi-branch head design of the neural network model, it can simultaneously perform vanishing point height, target detection, and scene classification, improving the accuracy and efficiency of recognition; and the recognition results are stored in a structured manner, facilitating subsequent data query and analysis. By using a vector database, specific types of bump and subsidence data can be quickly retrieved, improving the availability of data.

[0054] Further, the conditions for the video collected by the on-vehicle driving recorder of the operating vehicle to be obtained are:

[0055] The video clarity conforms to the preset range, the proportion of the road surface area in the video meets the preset proportion, the installation method, installation position, and angle of the driving recorder meet the preset requirements, and there are clear and fixed time information, speed information, and longitude and latitude information in the video.

[0056] Further, perform single-frame recognition on the image sequence based on a pre-trained neural network model, and extract the driving information of the operating vehicle in each frame of the image, as well as the vanishing point coordinates, target objects, and scene categories, including:

[0057] Perform single-frame recognition on the image sequence based on a pre-trained first neural network model, and extract the vanishing point coordinates, target objects, and scene categories;

[0058] Perform single-frame recognition on the image sequence based on a pre-trained second neural network model, and extract the time information, speed information, and longitude and latitude information.

[0059] Further, the first neural network model includes: an identification backbone network, a first branch head network, a second branch head network, and a third branch head network. The identification backbone network is used to perform multi-scale feature extraction on the input image and output a plurality of feature maps. The first branch head network is used to identify the vanishing point height in the feature maps. The second branch head network is used to identify the objects that cause jolts and the types of the objects in the feature maps. The third branch head network is used to identify the scene type of the image acquisition in the feature maps.

[0060] Further, the second neural network model is a neural network for performing text OCR recognition.

[0061] Further, slicing the time series data, analyzing the position change of the vanishing point during the driving of the operating vehicle based on the sliced data, and judging abnormal jolts according to the position change of the vanishing point, including:

[0062] Slicing the time series data at a preset first time interval to obtain a plurality of first sliced data;

[0063] Extracting the vanishing point height time series data and the speed time series data from the first sliced data, and calculating the average speed;

[0064] Fitting the vanishing point height time series data by using the least square method, obtaining a fitting straight line and calculating the vanishing point deviation, obtaining the vanishing point deviation time series data and determining the average jolt condition by calculating the root mean square;

[0065] Slicing the vanishing point deviation time series data at a preset second time interval to obtain a plurality of second sliced data, and selecting the second sliced data with the largest root mean square as the short-term maximum jolt;

[0066] Introducing average speed correction, calculating a jolt ratio based on the average jolt condition and the short-term maximum jolt, and comparing the jolt ratio with a preset jolt threshold to determine the abnormal jolt.

[0067] Further, introducing average speed correction, calculating a jolt ratio based on the average jolt condition and the short-term maximum jolt, including:

[0068] ;

[0069] wherein, R1 is the jolt ratio, v is the average speed of the first sliced data, alpha is a preset value set according to the vehicle type, rms1_sub_max is the short-term maximum jolt, and rms1 is the average jolt condition.

[0070] The following combines Figures 1-6 to elaborate in detail on the method provided in this embodiment, as Figure 1As shown in the figure, first, decode and extract frames from the video, use a neural network model to parse and recognize the images after frame extraction, summarize the recognition results, sort them according to the frame numbers of the image frames, and summarize the recognition results to obtain time-series data Q; then slice and parse the time-series data Q, analyze the position change of the vanishing point during the vehicle driving process from the sliced data, determine whether the vehicle has abnormal bumps, classify the abnormal bumps, and establish an association between the relevant recognition data and the abnormal bumps based on the time-series data Q to complete data storage.

[0071] 1. Video frame extraction and parsing, as Figure 2 shown, specifically including:

[0072] (1) Decode and extract frames from the video Video collected by the on-vehicle driving recorder to obtain an image sequence image_list. The data format of image_list is [[f_1, image_1], [f_2, image_2], [f_3,image_3], ……[f_n, image_n], ……], where image_1, image_2, …image_n, … are single-frame images in the video, and f_1, f_2, …f_n, … are the frame numbers corresponding to the frame images.

[0073] The requirements for the on-vehicle driving recorder video Video are as follows:

[0074] 1) The resolution ≥ 1 million pixels, the video frame rate ≥ 24fps, the collected video is clear, the proportion of the road surface area captured in the image ≥ 20%, there is no occlusion, and there will be no frame dropping or screen distortion.

[0075] 2) The used on-vehicle driving recorder is installed and fixed, and will not be interfered or occluded by other objects in the vehicle during driving. The shooting angle is towards the vehicle driving direction, the image is horizontally without obvious rolling or deflection, and the position of the vanishing point in the image is within the visible range of the image.

[0076] 3) There are text such as time information, speed information, longitude and latitude information, etc. in the video, and the text is clear and the text position is fixed.

[0077] (2) Use neural network model one to perform recognition processing on single-frame images. Taking the nth frame in the image sequence as an example, the frame number is f_n and the frame image is image_n. The processing steps are as follows:

[0078] 1) Preprocess the image image_n and input it into neural network model one for inference to calculate the corresponding inference result.

[0079] The model structure design of Neural Network Model 1 refers to the current mainstream object detection network (such as yolov11), including 1 recognition backbone network backbone, and 3 branch head networks, namely head_1, head_2, and head_3. The multiple branch heads share the inference feature map of the backbone network backbone to reduce the computational load and improve the recognition efficiency.

[0080] A) The backbone network backbone adopts a deep convolutional neural network structure, combines technologies such as residual connection and multi-scale feature extraction to achieve efficient and powerful feature extraction capabilities. The backbone network can use pre-trained model weights or be trained together with the head_2 detection class data.

[0081] B) head_1 is a classification head network. Taking one of the feature maps output by the backbone network backbone as input, after operations such as convolution, pooling, and self-attention, it is mapped to m categories through softmax. Among them, m refers to dividing m equal parts along the height H direction of the video image, m≥512. The category with the highest probability calculated by softmax is used as the interval where the vanishing point is located, that is, the j-th equal division interval counted from top to bottom. The height of the midpoint of this interval on the image is used as the height of the vanishing point of this image on the image, recorded as vp_h_n, as Figure 3 shown. Use the supervised learning method to train head_1, and freeze the backbone network backbone during training.

[0082] C) head_2 is a detection head network. Taking several feature maps output by the backbone network backbone as input, after operations such as convolution and pooling, the recognition detection results [[x1, y1, x2, y2, cls, conf], …] are output, and the intersection over union (IoU) analysis is performed between the recognition result boxes and the driving area of the inspection vehicle in the image. The recognition results outside the driving area of the inspection vehicle are excluded, and the recognition results within the driving area are retained, denoted as detects_n. Among them, cls is the type of the recognition target, including speed bumps, road markings, manholes, road repairs, potholes, road debris, expansion joints, and road subsidence. x1, y1, x2, y2 refer to the coordinate positions of the recognition target on the image, and conf refers to the confidence level of the recognition target. The driving area of the inspection vehicle in the image refers to the range that the wheels of the inspection vehicle can roll over in the image. When the camera is installed and fixed, this area generally does not change. Therefore, a trapezoidal area T can be preset in the image according to the actual image perspective as the driving area of the inspection vehicle, and the IoU intersection over union between each recognition target and the trapezoidal area T is calculated to filter out the recognition target objects with a small IoU, that is, to exclude the recognition target objects outside the driving area of the inspection vehicle, as Figure 4As shown in the figure, head_2 is trained using supervised learning, and the backbone network backbone may not be frozen during training.

[0083] D) head_3 is a classification head network. It takes one of the feature maps output by the backbone network as input. After convolution and pooling operations, it is mapped to k categories through softmax. The collected images are classified into k scenes, including expressway, ground, ramp, tunnel, under the elevated road, on the bridge, etc. The category with the highest probability calculated by softmax is used as the scene of the image collection, recorded as scene_n. Head_3 is trained using supervised learning, and the backbone network is frozen during training.

[0084] (3) Neural network model 2 uses a neural network specifically used for text OCR recognition. Through manual pre-marking or text recognition, the position of text information such as time, longitude and latitude, and speed in image image_n is pre-marked, and the corresponding time information sub-image image_sub1_n, longitude and latitude information sub-image image_sub2_n, and speed information sub-image image_sub3_n are cut out and input into neural network model 2 respectively, thereby identifying the corresponding time text information time_n, longitude and latitude text information lnglat_n, and speed text information v_n.

[0085] (4) The recognition results of the single-frame recognition analysis, such as the image vanishing point height, the detected target in the driving area, the image scene classification, and the time, are summarized and recorded as [f_n, vp_h_n, detects_n, scene_n, time_n, lnglat_n, v_n]; the analysis results of all frames of a video will be summarized into a queue, which is recorded as the time series queue Q. The programs and models involved in video frame extraction and image analysis can be deployed on the edge or in the cloud.

[0086] 2. Analyze the time series data Q and analyze the vehicle bumps;

[0087] The time series data Q records the changes of the vanishing point of the image over time, and the changes in the height of the vanishing point can reflect the bumpy condition of the road to a certain extent, such as Figure 6 Based on the analyzed time series data Q, the road bump condition is further analyzed. The specific steps are as follows: Figure 5 As shown, specifically including:

[0088] (1) Slice the time-series data Q at fixed intervals of time TIME to obtain the sliced time-series data Q1, separate the time-series data of the vanishing point height from it and perform median filtering, denoted as h1 (including the frame number and the vanishing point height), separate the speed time-series data extracted by OCR, and calculate the average speed v of this slice. The interval time TIME is generally a preset fixed time, and the value range is 3 - 5 seconds.

[0089] (2) Using the frame number of the time-series data h1 as the independent variable and the height value of the vanishing point in the time-series data h1 as the dependent variable, perform linear fitting using the least squares method to obtain the fitted straight line L1, and then calculate the deviations of each point on the time-series data h1 from L1 and sum them up to obtain the time-series data h1' of the vanishing point deviation. The deviation refers to the difference between the actual height value of each point and the height value of the fitted straight line L1. For example, for the frame number f_i, the corresponding vanishing point height is h1_i, substitute x = f_1 into the straight line L1, calculate L1_h_i, and calculate the deviation as h1'_i = h1_i - L1_h_i.

[0090] (3) For the time-series data h1' of the vanishing point deviation, calculate the mean value h1'_mean, and then according to the calculation formula , calculate the root mean square of the time-series data h1' of the vanishing point deviation, which is used as the average bump condition of the inspection vehicle within the slice TIME. Among them, x i is the i-th frame of the time-series data h1' of the vanishing point deviation, and N is the number of valid frames of the time-series data h1' of the vanishing point deviation.

[0091] (4) Further slice the time-series data h1' of the vanishing point deviation according to the fixed time TIME' to obtain s sub-slices of the same length, calculate the mean value h1'_sub_mean of the vanishing point deviation of each sub-slice respectively, and then according to the calculation formula , calculate the root mean square of each slice data, and take the largest value from them as the short-term maximum bump of the inspection vehicle within the slice TIME, denoted as rms1_sub_max, and record the frame number corresponding to this maximum short-term bump, denoted as f_sub_max. Among them, N' is the number of valid frames in the sub-slice data, and the fixed time TIME' is a preset value, generally a relatively short time, and the value range is 0.2 - 1.0 s according to experience.

[0092] (5) According to the calculation formula , calculate the ratio of the maximum short-term bump after speed correction to the average bump within this slice data, and use this ratio to characterize the bump condition of the inspection vehicle within the slice TIME. Among them, v is the average speed of the slice Q1, and alpha is a preset value, generally set according to the vehicle type and can be obtained through experimental statistics.

[0093] (6) Determine whether R1 exceeds a certain threshold THREHOLD. If R1 is smaller than THREHOLD, it indicates that the short-term jitter ratio intensity of this slice is relatively low, and there is no obvious abnormal jitter in the vehicle; if R1 is larger than THREHOLD, it indicates that the short-term jitter ratio intensity of this slice of data is relatively high, and the vehicle has experienced obvious jitter.

[0094] (7) Rate the abnormal jitter of the vehicle according to the magnitude of R1. According to the frame numbers (f_sub_max - beta1) to (f_sub_max - beta2), find the corresponding recognition information, speed information, time information, longitude and latitude information from the time-series data Q, and organize the relevant data and store it in the vector database.

[0095] beta1 and beta2 refer to the time difference from when the road surface target is captured in the image to when the road surface target causes the vehicle to jitter. It can be a fixed value or calculated by multiplying the average speed v by a fixed time.

[0096] Considering that some road surface diseases may be difficult to identify, or missed identification due to vehicle occlusion or insufficient detector performance, this part of the data is considered to be classified as "severe jitter with unknown cause". When organizing the recognition information, record the cause of the jitter as res, with the type being a string, and classify the cause according to the following situations:

[0097] 1) If there is only one type of recognized target A within the corresponding frame number range, record res as "A";

[0098] 2) If there are multiple types of recognized targets A, B, C, D within the corresponding frame number range, record res as "A\B\C\D";

[0099] 3) If there is no recognized target within the corresponding frame number range and R1 exceeds a specific threshold THREHOLD2, record res as "severe jitter with unknown cause".

[0100] After completing the organization of the recognition information, form a data record [f_sub_max, res, lnglat, time, v, scene, R1, rms1, rms1_sub_max, vec, image] and record it in the vector database. Among them, vec is a high-dimensional vector obtained by processing the string res using neural network model three. In this embodiment, neural network model three uses NetEase's BECmbedding model.

[0101] (8) When performing a search in the vector database later, it is possible to query the bump and subsidence data within a specified time, space, and scene range based on time, lnglat, and scene. It is also possible to search for specific types of bump and subsidence data using the vector similarity distance. For example, when searching in the query system for road surface bumps and subsidence caused by two diseases, A and B, the neural network model three can be used to process the string "A\B" to obtain the vector vec1, and then perform a similarity search or a hybrid search in the vector database to obtain the corresponding results.

[0102] (9) The slicing, parsing, and storage of time-series data can be deployed at the edge or in the cloud.

[0103] Embodiment 2:

[0104] This embodiment also provides a perception system for road surface subsidence and bumps, including:

[0105] A video acquisition module for acquiring the video collected by the driving recorder of an operating vehicle, extracting frames from the video, and obtaining an image sequence;

[0106] A video analysis module for performing single-frame recognition on the image sequence based on a pre-trained neural network model, extracting the driving information of the operating vehicle, as well as the vanishing point coordinates, target objects, and scene categories in each frame image, summarizing and sorting them to obtain time-series data;

[0107] A bump analysis module for slicing the time-series data, analyzing the position change of the vanishing point during the driving process of the operating vehicle based on the sliced data, and judging abnormal bumps according to the position change of the vanishing point;

[0108] A hierarchical storage module for classifying the abnormal bumps, establishing an association between the relevant recognition data and the abnormal bumps based on the time-series data, and storing them in the database.

[0109] The processing flow of the system proposed in this embodiment is as follows: After exporting the video collected by the bus, upload it to the cloud, and let the cloud system perform frame extraction analysis, time-series data analysis, and complete data storage on the collected video.

[0110] (1) Collecting videos, with the following requirements:

[0111] 1) Collect video data in the driving recorder.

[0112] 2) Resolution ≥ 1 million pixels, video frame rate ≥ 24fps.

[0113] 3) The proportion of the road surface area captured in the image ≥ 20%.

[0114] 4) The video is not blocked, and there will be no frame drops or screen corruption phenomena.

[0115] 5) The driving recorder is installed and fixed, with the shooting angle facing the forward direction of the vehicle's travel. The image is horizontally level without obvious rolling or deflection, and the position of the vanishing point in the image is within the visible range of the image.

[0116] 6) There are texts such as time information, speed information, longitude and latitude information, etc. in the video, and the texts are clear and the text positions are fixed.

[0117] (2) Video frame extraction analysis:

[0118] (1) Decode the collected video to obtain an image sequence image_list in the format [[f_1, image_1], [f_2, image_2], [f_3, image_3], ……[f_n, image_n], ……], where image_1, image_2, …image_n, … are single-frame images in the video, and f_1, f_2, …f_n, … are the frame numbers corresponding to the frame images.

[0119] (2) Frame extraction image parsing:

[0120] 1) Use neural network model one to perform recognition processing on single-frame images. Neural network model one includes a backbone network backbone and multiple branch head networks head_1, head_2, head_3:

[0121] backbone network backbone: Adopt a deep convolutional neural network structure, combine techniques such as residual connection and multi-scale feature extraction to achieve efficient and powerful feature extraction capabilities.

[0122] head_1: Classification head network, used to identify the height of the vanishing point.

[0123] head_2: Detection head network, used to detect objects that may cause vehicle bumps.

[0124] head_3: Classification head network, used to identify the image scene category.

[0125] 2) Preset the text information area in the image, crop out the sub-image according to the predefined coordinates, and use neural network model two to perform recognition on the sub-image to extract information such as time, longitude and latitude, and speed.

[0126] 3) Summarize the recognition results of the single-frame image image_n by the neural network model 1 and the neural network model 2, record them as [f_n, vp_h_n, detects_n, scene_n, time_n, lnglat_n, v_n], and summarize them into the time series queue Q, where:

[0127] f_n is the frame number corresponding to image image_n;

[0128] vp_h_n is the height of the image vanishing point, which is calculated by the neural network model 1 through the head_1 branch network;

[0129] detects_n is the identified targets in image_n that may cause road bumps, including speed bumps, road markings, manhole covers, road repairs, potholes, foreign objects on the road, expansion joints, road subsidence, and other targets that may cause vehicle bumps. They are calculated by neural network model 1 through the head_2 branch network;

[0130] scene_n is the scene category of image_n recognized by neural network model 1, such as expressway, ground, ramp, tunnel, under viaduct, on bridge, etc., which is calculated by neural network model 1 through head_3 branch network;

[0131] time_n is the time information of image_n, which is recognized by neural network model 2;

[0132] lnglat_n is the longitude and latitude information of image_n, which is recognized by neural network model 2;

[0133] v_n is the velocity information of image_n, which is recognized by neural network model 2.

[0134] Neural network model 2: used for text OCR recognition and extraction of time, longitude, latitude, speed and other information.

[0135] Neural network model three: used to convert recognition results into high-dimensional vectors for easy storage and query in vector databases.

[0136] (III) Q-slice analysis of time series data:

[0137] (1) Time series data slicing:

[0138] 1) Slice the time series data Q according to the fixed interval time TIME to obtain the sliced time series data Q1, and take the interval time TIME=5s.

[0139] 2) Separate the time series data of the vanishing point height from Q1 and perform median filtering to obtain h1.

[0140] 3) Isolate the speed time series data and calculate the average speed v of this slice.

[0141] (2) Linear fitting, deviation calculation, and root mean square calculation:

[0142] 1) Using the frame number of the time series data h1 as the independent variable and the height value of the vanishing point as the dependent variable, perform linear fitting using the least squares method to obtain the fitted line L1.

[0143] 2) Calculate the deviation of each point on the time series data h1 from L1 to obtain the vanishing point deviation time series data h1'.

[0144] 3) Calculate the mean h1'_mean of the vanishing point deviation time series data h1', and according to the calculation formula , calculate the root mean square of the vanishing point deviation time series data h1', which is used as the average bumpiness of the inspection vehicle within the TIME of this slice.

[0145] 4) Further slice the vanishing point deviation time series data h1' at a fixed time TIME' = 0.2s to obtain 5÷0.2 = 25 sub-slices of the same length. Calculate the mean h1'_sub_mean of the vanishing point deviation for each sub-slice, and then according to the calculation formula calculate the root mean square of each slice of data, and take the maximum value from them as the short-term maximum bumpiness of the inspection vehicle within the TIME of this slice, denoted as rms1_sub_max. At the same time, record the frame number corresponding to this maximum short-term bumpiness, denoted as f_sub_max.

[0146] 5) Take out the maximum value rms1_sub_max and record the frame number f_sub_max corresponding to this maximum short-term bumpiness.

[0147] (3) Bumpiness assessment:

[0148] 1) Calculate , to evaluate the bumpiness of the inspection vehicle within the TIME of this slice.

[0149] 2) Determine whether R1 exceeds the threshold THREHOLD to evaluate whether the vehicle has abnormal bumpiness: If R1 is greater than the threshold THREHOLD, rate the abnormal bumpiness of the vehicle according to the size of R1 and perform subsequent calculations; if R1 is less than THREHOLD, skip this slice of data and then take the next slice of data from the time series data Q for calculation.

[0150] 3) According to the frame numbers (f_sub_max - beta1) to (f_sub_max - beta2), find the corresponding recognition information, speed information, time information, longitude and latitude information from the timing data Q, and organize the relevant data and store it in the vector database.

[0151] Among them, beta1 and beta2 refer to the time difference from when the road surface target is captured in the image to when the road surface target causes the vehicle to jolt, which is calculated by multiplying the average speed v by the fixed time t'.

[0152] Considering that some road surface diseases may be difficult to identify, or may be missed due to vehicle occlusion or insufficient detector performance, so consider classifying this part of the data as "severe jolts of unknown cause". When organizing the recognition information, record the cause of the jolt as res, with the type being a string, and classify the cause according to the following situations:

[0153] If there is exactly one type of recognized target A within the corresponding frame number range, record res as "A";

[0154] If there are multiple types of recognized targets A, B, C, D within the corresponding frame number range, record res as "A\B\C\D";

[0155] If there is no recognized target within the corresponding frame number range and R1 exceeds a specific threshold THREHOLD2, record res as "severe jolts of unknown cause".

[0156] After completing the organization of the recognition information, form a data record [f_sub_max, res, lnglat, time, v, scene, R1, rms1, rms1_sub_max, vec, image], and record it in the vector database MILVUS. Among them, vec is a high-dimensional vector obtained by processing the string res using the neural network model three. In this embodiment, the neural network model three uses NetEase's BECmbedding model.

[0157] (4) Subsequent data query usage:

[0158] Subsequently, when searching in the vector database, it is possible to query the jolt and subsidence data within a specified time, space, and scene range according to time, lnglat, and scene, or search for specific types of jolt and subsidence data using the vector similarity distance. For example, when searching in the query system for road surface jolts and subsidence caused by diseases A and B, the vector vec1 can be obtained by processing the string "A\B" using the neural network model three, and then perform similarity search or hybrid search in the vector database to obtain the corresponding results.

[0159] This embodiment utilizes an existing driving recorder, eliminating the need for additional installation of sensors and significantly reducing the hardware cost. Compared with traditional intelligent inspection devices, it reduces the expenses for equipment procurement and maintenance. By leveraging the extensive coverage and high-frequency driving of operating vehicles, it can achieve comprehensive monitoring of urban roads, reducing the investment in dedicated inspection vehicles and personnel. The structured storage and rapid retrieval of data enable maintenance personnel to quickly locate problem sections, improving the efficiency and effectiveness of maintenance work. By promptly detecting and recording road surface diseases, maintenance and repair can be carried out in advance to avoid further deterioration of the diseases, reducing the repair cost. It can also reduce traffic accidents caused by road surface problems, lowering the economic losses and social costs brought by traffic accidents.

[0160] The embodiments described above are merely descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for perceiving pavement settlement and bumps, characterized in that Including: Obtain the video collected by the driving recorder of the operating vehicle, extract frames from the video, and obtain an image sequence; Based on a pre-trained neural network model, perform single-frame recognition on the image sequence, extract the driving information of the operating vehicle in each frame of the image, as well as the vanishing point coordinates, target objects, and scene categories, and summarize and sort them to obtain time-series data; Slice the time-series data, analyze the change of the position of the vanishing point during the driving process of the operating vehicle based on the sliced data, and judge abnormal bumps according to the change of the position of the vanishing point; Classify the abnormal bumps, establish an association between relevant recognition data and the abnormal bumps based on the time-series data, and store them in the database.

2. The method for perceiving road surface subsidence and bumps according to claim 1, wherein, The condition for obtaining the video collected by the driving recorder of the operating vehicle is: The video clarity conforms to the preset range, the proportion of the road surface area in the video meets the preset proportion, the installation method, installation position, and angle of the driving recorder meet the preset requirements, and there are clear and fixed time information, speed information, longitude and latitude information in the video.

3. The method for perceiving pavement settlement and bumpiness according to claim 1, characterized in that, Based on a pre-trained neural network model, perform single-frame recognition on the image sequence, extract the driving information of the operating vehicle in each frame of the image, as well as the vanishing point coordinates, target objects, and scene categories, including: Based on a pre-trained first neural network model, perform single-frame recognition on the image sequence, and extract the vanishing point coordinates, target objects, and scene categories; Based on a pre-trained second neural network model, perform single-frame recognition on the image sequence, and extract the time information, speed information, longitude and latitude information.

4. The method for perceiving pavement settlement and bumpiness according to claim 3, characterized in that, The first neural network model includes: a recognition backbone network, a first branch head network, a second branch head network, and a third branch head network. The recognition backbone network is used to perform multi-scale feature extraction on the input image and output several feature maps; the first branch head network is used to recognize the height of the vanishing point in the feature map, the second branch head network is used to recognize the target objects that cause bumps and the types of target objects in the feature map, and the third branch head network is used to recognize the scene type of the image acquisition in the feature map.

5. The method for perceiving pavement settlement and bumps according to claim 3, characterized in that, The second neural network model uses a neural network for text OCR recognition.

6. The method for perceiving pavement settlement and bumps according to claim 1, wherein Slice the time-series data, analyze the change of the position of the vanishing point during the driving process of the operating vehicle based on the sliced data, and judge abnormal bumps according to the change of the position of the vanishing point, including: Slice the time-series data at a preset first time interval to obtain several first slice data; Extract the vanishing point height time-series data and speed time-series data from the first slice data, and calculate the average speed; Use the least squares method to fit the vanishing point height time-series data, obtain the fitted straight line and calculate the vanishing point deviation, obtain the vanishing point deviation time-series data and determine the average bump condition by calculating the root mean square; Slice the vanishing point deviation time-series data at a preset second time interval to obtain several second slice data, and select the second slice data with the largest root mean square as the short-term maximum bump; Introduce average speed correction, calculate the bump ratio based on the average bump condition and the short-term maximum bump, compare the bump ratio with a preset bump threshold, and determine the abnormal bump.

7. The method for perceiving pavement settlement and bump according to claim 6, characterized in that, Introduce the average speed correction, and calculate the jolt ratio based on the average jolt condition and the short-term maximum jolt, including: ; Among them, R1 is the jolt ratio, v is the average speed of the first slice of data, alpha is a preset value set by vehicle type, rms1_sub_max is the short-term maximum jolt, and rms1 is the average jolt condition.

8. A perception system for pavement subsidence and bumpiness, characterized in that, Including: A video acquisition module, which is used to acquire the video collected by the driving recorder of the operating vehicle, extract frames from the video, and obtain an image sequence; A video analysis module, which is used to perform single-frame recognition on the image sequence based on a pre-trained neural network model, extract the driving information of the operating vehicle in each frame of the image, as well as the vanishing point coordinates, target objects, and scene categories, and summarize and sort them to obtain time-series data; A jolt analysis module, which is used to slice the time-series data, analyze the position change of the vanishing point during the driving process of the operating vehicle based on the sliced data, and judge abnormal jolts according to the position change of the vanishing point; A hierarchical storage module, which is used to classify the abnormal jolts, establish an association between the relevant recognition data and the abnormal jolts based on the time-series data, and store them in the database.

Citation Information

Patent Citations

  • Highway pavement detection method based on image processing

    CN108171695A

  • Pavement disease positioning and size calculation method and system

    CN119198728A

  • Position detector

    JP2002259995A

  • Apparatus and method for detecting road surface unevenness

    JP2013011454A

Cited By

  • An intelligent detection and analysis method based on image multi-modal analysis

    CN122780741A