Method, system and equipment for dynamically diagnosing road surface state of automobile and medium
By employing synchronous data acquisition and multimodal feature fusion, and utilizing a large-scale model enhanced with road knowledge for road condition diagnosis, this approach addresses the shortcomings of existing systems in terms of the refinement and intelligent reasoning required for road health status detection. It achieves high-precision defect detection and prediction, adapting to the detection needs of different roads.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGSHA YIFENG AUTOMOBILE TECHNOLOGY CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-21
Smart Images

Figure CN121901845A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent connected vehicle technology, specifically relating to a method, system, device, and medium for dynamic diagnosis of road conditions of a vehicle. Background Technology
[0002] With the rapid development of intelligent connected vehicle technology, the perception and data processing capabilities of vehicles are constantly improving, providing new technological pathways for road condition detection. Existing 360° surround-view vehicle systems primarily focus on driver assistance and parking, using multiple wide-angle cameras around the vehicle to collect images, which are then stitched and corrected to form a panoramic overhead view, eliminating blind spots. While some advanced systems integrate AI algorithms to identify obstacles such as pedestrians and vehicles, or attempt to combine inertial measurement unit data for simple road condition analysis, existing technologies have significant limitations: Existing systems primarily serve real-time display and obstacle warning, lacking the ability to perform refined quantitative detection and assessment of the road surface's health condition. Furthermore, while some solutions mention multi-sensor fusion, these are mostly limited to the overlay of location information, failing to perform deep feature correlation and joint analysis of surround-view images, vehicle motion posture, geographical location, and historical data in a spatiotemporally synchronized manner, resulting in insufficient information utilization. Moreover, existing technologies lack intelligent reasoning and prediction capabilities regarding the causes and evolution trends of road surface defects. Traditional detection systems have fixed parameters, making it difficult to adapt to the differences in detection standards across different regions and road grades, and they typically cannot explain the detection results and their basis to management personnel in a natural and interactive way. Summary of the Invention
[0003] The primary objective of this invention is to overcome the shortcomings of the prior art and provide a method, system, device, and medium for dynamic diagnosis of road conditions of automobiles.
[0004] A method for dynamic diagnosis of road conditions of a vehicle includes the following steps: S1: Synchronously acquire vehicle surround view images, inertial measurement unit data, and positioning data, and synchronize them via a clock signal; S2: Preprocess and extract features from the collected data at the edge to obtain multimodal feature vectors; the preprocessing includes panoramic image stitching and data spatiotemporal alignment, and the feature extraction includes extracting feature vectors of the disease area from the vehicle surround view image and extracting vibration feature vectors related to road surface smoothness from the inertial measurement unit data; S3: Input the multimodal feature vector into the road knowledge enhancement big model to perform data fusion, state diagnosis, causal reasoning and trend prediction. The road knowledge enhancement big model is a domain expert model formed by injecting professional road engineering knowledge into a general big model pre-trained with text, code and image data and fine-tuning it with instructions. S4: Based on the output of the large model, generate visualization results and a detection report that can be interacted with in natural language.
[0005] Furthermore, in step S1, the synchronously collected data also includes road spectrum data and vehicle bus data. The vehicle bus data includes vehicle speed and turn signal. Data acquisition is achieved through multiple wide-angle cameras, inertial measurement units, high-precision GPS receivers, and vehicle communication units arranged on the vehicle body. The wide-angle cameras acquire high-resolution video streams and image sequences. The inertial measurement units acquire high-frequency motion attitude data such as three-axis acceleration, three-axis angular velocity, and three-axis magnetic field strength of the vehicle. The GPS receiver acquires the vehicle's latitude and longitude coordinates, altitude, precise timestamp, and real-time speed. The vehicle communication unit, as the data transmission hub, is responsible for data aggregation and transmission.
[0006] Furthermore, in step S2, the panoramic image stitching steps include: performing distortion correction on the original fisheye image captured by the wide-angle camera, and stitching it into a panoramic image through feature point matching and viewpoint transformation; data spatiotemporal alignment to correlate and match data collected by different sensors at the same time stamp; feature extraction including extracting the depth feature vector of the image; and extracting the vibration feature vector after filtering and coordinate transformation of the inertial measurement signal.
[0007] Furthermore, in step S3, data fusion uses the multimodal data understanding interface of the large model to vectorize visual features, vibration signals, geographical locations, etc., and map them to the unified semantic space of the large model; state diagnosis includes the accurate classification and quantification of road surface defects and the calculation of road surface driving quality; causal reasoning is used to infer the potential causes of defects based on defect features and historical data; and trend prediction is used to predict the development trend of defects.
[0008] Furthermore, the precise classification and quantification of pavement defects involves pixel-level semantic segmentation of panoramic images to identify defects including cracks, potholes, and repairs, calculate their geometric features, and cross-validate the visual recognition results using vibration data. The calculation of pavement driving quality includes analyzing the power spectral density of the vertical acceleration signal, combining it with the current vehicle speed, inverting the pavement longitudinal profile through the vehicle-road transfer function, and calculating the International Roughness Index.
[0009] Furthermore, in step S4, the visualization results are displayed in the form of electronic maps and charts. On the electronic map, the location of the disease and the flatness are displayed by overlaying icons of different colors and heat maps. The detection report that can be interacted with in natural language can be queried in natural language. The large model generates a personalized report with explanatory text based on the user's query.
[0010] Furthermore, it also includes a continuous learning step: each detection data and result is used as a new sample to optimize and fine-tune the parameters of the large model, so as to realize the self-evolution of the model performance. The new sample includes the detected disease information, flatness data, and the implementation status of maintenance suggestions.
[0011] A dynamic road condition diagnostic system for automobiles, comprising: Data acquisition module: includes multiple wide-angle cameras, inertial measurement unit, global positioning system receiver and vehicle communication unit arranged on the vehicle body. All sensors are synchronized by clock signal to synchronously acquire vehicle surround view images, inertial measurement data, high-precision positioning data, road spectrum data and vehicle bus data. Edge computing unit: connected to the data acquisition module, responsible for preprocessing and feature extraction of raw data. The preprocessing includes panoramic image stitching and data spatiotemporal alignment. Feature extraction includes extracting feature vectors of the diseased area from the image and extracting vibration feature vectors related to road surface smoothness from the inertial measurement signal, and performing compression encoding. Large Model Intelligent Analysis Center: Includes a multimodal data understanding interface, a large model for road knowledge enhancement, a diagnostic reasoning and prediction engine, and a report and suggestion generator, used to achieve multimodal data fusion, in-depth diagnosis of road conditions, causal reasoning and evolution prediction, generation of structured inspection reports, and provision of maintenance decision suggestions; Results Output and Interaction Module: This module displays the analysis results in a visual manner, allowing users to query detailed information using natural language, thus enabling the output and interaction of the detection results.
[0012] A computer device includes a processor and a memory connected to the processor. The memory stores one or more programs that are executed by the processor to implement the steps in the above-described method for dynamic diagnosis of road conditions of a vehicle.
[0013] A computer-readable storage medium storing one or more programs, which are executed by a processor to implement the steps in the above-described method for dynamic diagnosis of road conditions of a vehicle.
[0014] Beneficial effects: This invention uses a large model as the core analysis engine, enabling the system not only to detect road surface defects, but also to assess their severity, deduce potential causes, and predict their development trends. This provides unprecedented technical support for precision maintenance and solves the problem of superficial analysis in existing technologies.
[0015] This invention uses a large model to correlate and reason about visual, inertial, and positioning information within a unified semantic space, achieving deep fusion that goes beyond simple data overlay. This significantly improves detection accuracy and anti-interference capabilities, overcoming the data isolation limitations of existing technologies. Through the generation and reasoning capabilities of the large model, managers can conduct in-depth queries via natural language dialogue, improving decision-making efficiency and system usability, and resolving the poor interactivity issues of existing systems.
[0016] This invention uses a continuous learning mechanism to treat each detection data and result as a new sample for optimizing and fine-tuning the parameters of the large model, thereby continuously adapting to different road types and degradation modes and achieving continuous performance optimization. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic flowchart illustrating the main method steps of the present invention. Detailed Implementation
[0019] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0020] In the description of this application, it should be understood that "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. In this application, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise explicitly specified. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0021] Reference Figure 1As shown, this embodiment provides a method for dynamic diagnosis of road conditions of a vehicle, including the following steps: S1: Synchronously acquire vehicle surround view images, inertial measurement unit data, and positioning data, and synchronize them via a clock signal; S2: Preprocess and extract features from the collected data at the edge to obtain multimodal feature vectors; the preprocessing includes panoramic image stitching and data spatiotemporal alignment, and the feature extraction includes extracting feature vectors of the disease area from the vehicle surround view image and extracting vibration feature vectors related to road surface smoothness from the inertial measurement unit data; S3: Input the multimodal feature vector into the road knowledge enhancement big model to perform data fusion, state diagnosis, causal reasoning and trend prediction. The road knowledge enhancement big model is a domain expert model formed by injecting professional road engineering knowledge into a general big model pre-trained with text, code and image data and fine-tuning it with instructions. S4: Based on the output of the large model, generate visualization results and a detection report that can be interacted with in natural language.
[0022] In step S1, 360-degree surround view images of the vehicle, inertial measurement unit (IMU) data, and high-precision positioning data are simultaneously acquired. All sensors are synchronized via a high-precision clock signal to ensure data time consistency. The synchronously acquired data also includes road spectrum data and vehicle bus data, including vehicle speed and turn signal signals. Data acquisition is achieved through multiple wide-angle cameras, an inertial measurement unit (IMU), a high-precision global positioning system (GPS) receiver, and an onboard communication unit (T-Box) deployed on the vehicle body.
[0023] Specifically, the wide-angle camera primarily captures high-resolution video streams and image sequences. Through a surround-view stitching algorithm, it generates a 360-degree bird's-eye view sequence of images surrounding the vehicle. This visual data serves as the basis for identifying road surface defects. The inertial measurement unit (IMU) collects high-frequency motion attitude data of the vehicle, including triaxial acceleration, triaxial angular velocity, and triaxial magnetic field strength. This data serves as crucial raw signals for analyzing vehicle vibration and calculating road surface smoothness. The high-precision GPS receiver collects the vehicle's latitude and longitude coordinates, altitude, precise timestamp, and real-time speed, providing accurate spatial location labels for all detection results and enabling visualization of road surface defects. Speed data is a necessary parameter for calculating smoothness indicators. The onboard communication unit itself does not directly collect physical world data but acts as a data transmission hub. It is responsible for aggregating the data collected by the aforementioned sensors through the onboard network and, with the aid of cellular networks or V2X technology, transmitting data packets in real time to the edge computing unit or the cloud.
[0024] The test vehicle was an electric SUV with intelligent connectivity features, equipped with a CAN bus interface to acquire signals such as vehicle speed and turn signals from the vehicle bus. The data acquisition module included four 190° ultra-wide-angle fisheye cameras, a high-precision 9-axis IMU, a high-precision GPS module supporting RTK technology, and an onboard communication unit (T-Box). The four wide-angle cameras were mounted near the front grille, the rear license plate lights, and the bottom of the left and right side mirrors, ensuring 360° coverage around the vehicle. The IMU sampling frequency was set to 100Hz to collect attitude data such as three-axis acceleration and angular velocity. The GPS module achieved centimeter-level positioning accuracy, outputting the vehicle's latitude and longitude coordinates, altitude, timestamp, and real-time speed at a frequency of 1Hz. All sensors were synchronously triggered via a unified hardware clock signal, and the cameras captured video streams at a frame rate of 30fps, with timestamp errors controlled within milliseconds. The onboard communication unit aggregated the data collected by the sensors via the CAN bus and transmitted the data packets to the edge computing unit in real time using a 5G network.
[0025] In step S2, the collected data is preprocessed and features are extracted at the edge to obtain multimodal feature vectors. The preprocessing includes panoramic image stitching and data spatiotemporal alignment. Feature extraction includes extracting feature vectors of the affected areas from the images and extracting vibration feature vectors related to road surface smoothness from the IMU signals, followed by compression encoding.
[0026] Specifically, the panoramic image stitching process is as follows: distortion correction is performed on the original fisheye image captured by the wide-angle camera, and lens distortion is corrected using camera calibration parameters to obtain a distortion-free image; then, feature point matching is performed using a feature point detection algorithm, and multiple images are projected onto a unified bird's-eye view model through perspective transformation to stitch them into a seamless 360-degree panoramic image. Data spatiotemporal alignment involves associating and matching data collected by different sensors at the same time stamp to ensure data consistency in both time and space dimensions.
[0027] In feature extraction, instead of directly transmitting massive panoramic images, a lightweight convolutional neural network or the encoder portion of a visual transformer is used to perform forward propagation on the panoramic image, extracting highly abstract depth feature vectors. These vectors may only have a few hundred or a thousand dimensions, far smaller than the millions of pixels in the original image, significantly compressing the data volume. For IMU signals, filtering and coordinate transformation are performed first, converting them from the vehicle coordinate system to the geodetic coordinate system. Then, the frequency domain features are analyzed using a fast Fourier transform, or the root mean square value of acceleration is calculated as the time domain feature to extract vibration feature vectors related to road surface smoothness.
[0028] The edge computing unit utilizes an onboard industrial control computer equipped with an NVIDIA Jetson AGX Orin module. First, the received raw data packets are parsed and verified, then classified. For the 360° video stream, image distortion correction is performed first, using camera calibration parameters to correct lens distortion and obtain a distortion-free image. Then, the ORB algorithm is used for feature point matching, and perspective transformation is used to project the four images onto a unified bird's-eye view model, stitching them together into a seamless 360° panoramic image. Finally, the encoder part of the MobileNet lightweight convolutional neural network is used to perform forward propagation on the panoramic image to extract depth feature vectors. For the raw IMU data, coordinate transformation is performed, converting acceleration from the vehicle coordinate system to the geodetic coordinate system. A Butterworth low-pass filter is then used to filter out high-frequency vibration noise, retaining low-frequency signals that reflect the macroscopic smoothness of the road surface. Fast Fourier Transform is then used to analyze its frequency domain characteristics, extracting energy from specific frequency bands as vibration feature vectors. For GPS / Bus data, information such as position, speed, and timestamps is parsed to generate standardized position and state vectors. Finally, the extracted image depth feature vectors, vibration feature vectors, and position and state vectors are aligned by time, multimodal feature fusion is performed, a structured feature data package is generated, and sent to the large model analysis center.
[0029] In step S3, the multimodal feature vectors are input into the pavement knowledge-enhanced large model for data fusion, state diagnosis, causal reasoning, and trend prediction. The pavement knowledge-enhanced large model is a domain expert model formed by pre-training a general-purpose model with massive amounts of text, code, and image data, and then injecting professional road engineering knowledge, including the Pavement Disorder Classification Standard (PASER), maintenance specifications, material degradation model literature, and a large amount of pavement image and text pairing data, through instruction fine-tuning.
[0030] Specifically, data fusion uses the multimodal data understanding interface of a large model to vectorize visual features, vibration signals, geographic locations, etc., and map them to a unified semantic space that the large model can understand. The large model is equipped with a visual encoder to map the input panoramic image feature vectors to the same semantic space as the large model's text representation. At this point, the system constructs a prompt word containing multimodal information and inputs it to the large model.
[0031] Condition diagnosis includes the precise classification and quantification of pavement defects and the accurate calculation of pavement ride quality. Precise classification and quantification of pavement defects involves pixel-level semantic segmentation of panoramic images to identify defects such as cracks, potholes, and repairs, and accurately calculates their geometric features. Vibration data is then used to cross-validate the visual recognition results, reducing false positives. Accurate calculation of pavement ride quality involves analyzing the power spectral density of vertical acceleration signals, combining this with the current vehicle speed, and using the known vehicle-road transfer function to invert the longitudinal profile of the pavement, thereby calculating objective indicators such as the International Roughness Index.
[0032] Causal reasoning is based on the characteristics of road defects and historical data to infer the potential causes of defects. For example, when a network of cracks is identified and accompanied by a moderate IRI value, it is inferred that the road section has fatigue damage to the base layer, mainly due to heavy traffic. Trend prediction is to predict the development trend of defects, such as predicting that defects will accelerate in the next rainy season and that the risk of potholes is high.
[0033] The large-scale intelligent analysis center is deployed on a cloud server, using the Qwen series of general-purpose large-scale models as the base model. First, structured road engineering knowledge data is constructed, including maintenance specifications such as my country's "Highway Technical Condition Assessment Standard" (JTGH20), design and construction specifications, historical inspection reports, maintenance records, expert diagnostic opinions, and professional textbooks and papers on pavement distress mechanisms and material degradation models. This knowledge is then converted into structured question-and-answer pairs of standard clauses, specific requirements, and evaluation indicators, as well as samples of case backgrounds, problem descriptions, causal analysis, and remedial solutions. Then, the LoRA parameter efficient fine-tuning method is used. Most parameters of the original model are frozen, a LoRA adapter is injected, and the adapter and specific layers are trained. The base model is then fine-tuned using commands to obtain a pavement knowledge-enhanced domain expert model.
[0034] After receiving feature data packets from the edge computing unit, the large model vectorizes visual features, vibration signals, and geographical locations through a multimodal data understanding interface and maps them to a unified semantic space. Prompt words containing multimodal information are constructed and input into the large model, such as: "Analyze the panoramic image features [feature vector], vertical vibration signal [feature vector], and historical road condition data of the current vehicle location (39.9°N, 116.4°E). Please diagnose the road surface condition and predict the trend over the next three months." The large-scale model performs in-depth analysis based on prompts to achieve accurate classification and quantification of pavement defects, precise calculation of pavement driving quality, causal inference, and evolution prediction. For example, the model identifies obvious linear features in the image feature vector and, combined with the fact that the energy of a specific frequency band in the vibration feature vector is significantly higher than the threshold, infers a longitudinal crack. Through pixel-level semantic segmentation, the crack length is calculated to be approximately 2.5 meters and the width to be 4 mm, with a moderate severity. Based on the power spectral density of the vertical acceleration signal and the current vehicle speed, the International Roughness Index (IRI) is calculated to be 3.5 m / km, with a roughness level of medium. Combining this with historical data showing that this section of road has reported minor network cracks in the past 3 months, the model infers that the longitudinal crack may be caused by base fatigue or reflective cracking. It predicts that if left untreated, the crack will accelerate its expansion after the rainy season, and its width may increase to 6 mm within 3 months, accompanied by the formation of new network cracks. At the same time, maintenance recommendations are generated, suggesting that the crack be repaired within 2 weeks to prevent further development of the defect.
[0035] In step S4, based on the output of the large model, a visualization and a natural language interactive inspection report are generated. The visualization results are presented in the form of electronic maps and charts. On the electronic map, the location of defects and the flatness status are displayed by overlaying icons of different colors and heat maps. Clicking on the icons allows users to view a detailed diagnostic report. The natural language interactive inspection report allows users to query detailed information through natural language. Users can ask questions on the terminal, and the large model will generate human-like answers and automatically generate a structured report containing a defect distribution map, flatness curves, and maintenance suggestions.
[0036] The method also includes a continuous learning step: each detection data and result is used as a new sample to optimize and fine-tune the parameters of the large model, achieving self-evolution of model performance. The new samples include detected defect information, flatness data, and the implementation status of maintenance recommendations. Detection results and raw data can be encrypted and transmitted back to the cloud data center via 4G / 5G network. This data will serve as new samples for periodic fine-tuning of the large model.
[0037] The results output and interaction module displays visualized results on the vehicle's central control screen and connected tablet. The location of longitudinal cracks is marked with red icons on the electronic map, and the smoothness of the area is shown in a heat map. Clicking the red icon allows viewing a detailed diagnostic report, including the crack's geometric characteristics, smoothness data, cause analysis, development trend prediction, and maintenance recommendations. Users input the natural language question "Does this crack need immediate repair?" through the terminal. The system interprets the user's intent, retrieves relevant current and historical analysis data, constructs question-and-answer prompts, and submits them to the large model. The large model generates a natural language answer: "Based on analysis, the longitudinal crack is approximately 2.5 meters long and 4 mm wide, with a moderate severity. It may be caused by base layer fatigue or reflective cracking. If not repaired promptly, the width is expected to increase to 6 mm within 3 months, accompanied by the formation of new network cracks. Repair is recommended within 2 weeks to prevent further damage and ensure road safety." Continuous learning phase: The data obtained from this inspection, including road defect information, smoothness data, maintenance recommendations, and the implementation status of subsequent maintenance recommendations, are used as new samples. These samples are encrypted and transmitted back to the cloud data center via a 5G network. These new samples are used periodically to fine-tune the large model, optimize model parameters, and improve the model's accuracy in identifying road defects, the accuracy of causal reasoning, and the reliability of trend prediction, thus achieving self-evolution of model performance.
[0038] The various modules of the vehicle road condition dynamic diagnostic system of this invention work collaboratively. The data acquisition module continuously collects multi-source data, the edge computing unit performs real-time preprocessing and feature extraction, the large-model intelligent analysis center completes the core diagnosis, reasoning, and prediction, and the result output and interaction module realizes the visualization of information and natural language interaction. A continuous learning mechanism ensures that the system performance is continuously optimized. This system can adapt to the detection needs of roads in different regions and of different grades, providing comprehensive, accurate, and efficient technical support for road maintenance and management.
[0039] This embodiment also provides a dynamic road condition diagnostic system for automobiles, including: Data acquisition module: includes multiple wide-angle cameras, inertial measurement unit, global positioning system receiver and vehicle communication unit arranged on the vehicle body. All sensors are synchronized by clock signal to synchronously acquire vehicle surround view images, inertial measurement data, high-precision positioning data, road spectrum data and vehicle bus data. Edge computing unit: connected to the data acquisition module, responsible for preprocessing and feature extraction of raw data. The preprocessing includes panoramic image stitching and data spatiotemporal alignment. Feature extraction includes extracting feature vectors of the diseased area from the image and extracting vibration feature vectors related to road surface smoothness from the inertial measurement signal, and performing compression encoding. Large Model Intelligent Analysis Center: Includes a multimodal data understanding interface, a large model for road knowledge enhancement, a diagnostic reasoning and prediction engine, and a report and suggestion generator, used to achieve multimodal data fusion, in-depth diagnosis of road conditions, causal reasoning and evolution prediction, generation of structured inspection reports, and provision of maintenance decision suggestions; Results Output and Interaction Module: This module displays the analysis results in a visual manner, allowing users to query detailed information using natural language, thus enabling the output and interaction of the detection results.
[0040] This embodiment also provides a computer device, which includes a processor and a memory. The memory is connected to the processor and stores one or more programs. The one or more programs are executed by the processor to implement the steps in the above-described method for dynamic diagnosis of road conditions of a vehicle.
[0041] This embodiment also provides a computer-readable storage medium storing one or more programs, which are executed by a processor to implement the steps in the above-described method for dynamic diagnosis of road conditions of a vehicle.
[0042] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0043] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for dynamic diagnosis of road conditions of a vehicle, characterized in that, Includes the following steps: The system simultaneously collects vehicle surround view images, inertial measurement unit data, and positioning data, and synchronizes them via a clock signal. The collected data is preprocessed and features extracted at the edge to obtain multimodal feature vectors. The preprocessing includes panoramic image stitching and spatiotemporal data alignment. Feature extraction includes extracting feature vectors of the defect area from vehicle surround view images and extracting vibration feature vectors related to road surface smoothness from inertial measurement unit data. The multimodal feature vectors are input into a large-scale road knowledge enhancement model for data fusion, state diagnosis, causal reasoning, and trend prediction. The large-scale road knowledge enhancement model is a domain expert model formed by injecting professional road engineering knowledge into a general large-scale model pre-trained with text, code, and image data and fine-tuning it. Based on the output of the large-scale model, a visualization result and a detection report that can be interacted with in natural language are generated.
2. The method for dynamic diagnosis of road surface conditions of a vehicle according to claim 1, characterized in that, In step S1, the synchronously collected data also includes road spectrum data and vehicle bus data. The vehicle bus data includes vehicle speed and turn signal. Data acquisition is achieved through multiple wide-angle cameras, inertial measurement units, high-precision GPS receivers, and vehicle communication units arranged on the vehicle body. The wide-angle cameras acquire high-resolution video streams and image sequences. The inertial measurement units acquire high-frequency motion attitude data such as three-axis acceleration, three-axis angular velocity, and three-axis magnetic field strength of the vehicle. The GPS receiver acquires the vehicle's latitude and longitude coordinates, altitude, precise timestamp, and real-time speed. The vehicle communication unit, as the data transmission hub, is responsible for data aggregation and transmission.
3. The method for dynamic diagnosis of road surface conditions of a vehicle according to claim 1, characterized in that, In step S2, the panoramic image stitching steps include: performing distortion correction on the original fisheye image captured by the wide-angle camera, and stitching it into a panoramic image through feature point matching and viewpoint transformation; data spatiotemporal alignment to correlate and match data collected by different sensors at the same time stamp; feature extraction including extracting the depth feature vector of the image; and extracting the vibration feature vector after filtering and coordinate transformation of the inertial measurement signal.
4. The method for dynamic diagnosis of road surface conditions of a vehicle according to claim 1, characterized in that, In step S3, data fusion uses the multimodal data understanding interface of the large model to vectorize visual features, vibration signals, geographical locations, etc., and map them to the unified semantic space of the large model; state diagnosis includes the accurate classification and quantification of road surface defects and the calculation of road surface driving quality; causal reasoning is used to infer the potential causes of defects based on defect features and historical data; and trend prediction is used to predict the development trend of defects.
5. The method for dynamic diagnosis of road surface conditions of a vehicle according to claim 4, characterized in that, Accurate classification and quantification of road surface defects involves pixel-level semantic segmentation of panoramic images to identify defects including cracks, potholes, and repairs, and to calculate their geometric features. The visual recognition results are cross-validated by combining vibration data; the road surface driving quality calculation includes analyzing the power spectral density of the vertical acceleration signal, combining the current vehicle speed, inverting the longitudinal profile of the road surface through the vehicle-road transfer function, and calculating the international roughness index.
6. The method for dynamic diagnosis of road surface conditions of a vehicle according to claim 1, characterized in that, In step S4, the visualization results are displayed in the form of electronic maps and charts. The location of the disease and the flatness are displayed on the electronic map by overlaying icons of different colors and heat maps. The detection report can be interactive with natural language. Through natural language queries, the large model generates personalized reports with explanatory text based on user queries.
7. The method for dynamic diagnosis of road surface conditions of a vehicle according to claim 1, characterized in that, It also includes a continuous learning step: each detection data and result is used as a new sample to optimize and fine-tune the parameters of the large model, so as to realize the self-evolution of the model performance. The new sample includes the detected disease information, flatness data, and the implementation status of maintenance suggestions.
8. A dynamic road condition diagnostic system for automobiles, used to implement the dynamic road condition diagnostic method for automobiles as described in any one of claims 1-7, characterized in that, include: Data acquisition module: includes multiple wide-angle cameras, inertial measurement unit, global positioning system receiver and vehicle communication unit arranged on the vehicle body. All sensors are synchronized by clock signal to synchronously acquire vehicle surround view images, inertial measurement data, high-precision positioning data, road spectrum data and vehicle bus data. Edge computing unit: connected to the data acquisition module, responsible for preprocessing and feature extraction of raw data. The preprocessing includes panoramic image stitching and data spatiotemporal alignment. Feature extraction includes extracting feature vectors of the diseased area from the image and extracting vibration feature vectors related to road surface smoothness from the inertial measurement signal, and performing compression encoding. Large Model Intelligent Analysis Center: Includes a multimodal data understanding interface, a large model for road knowledge enhancement, a diagnostic reasoning and prediction engine, and a report and suggestion generator, used to achieve multimodal data fusion, in-depth diagnosis of road conditions, causal reasoning and evolution prediction, generation of structured inspection reports, and provision of maintenance decision suggestions; Results Output and Interaction Module: This module displays the analysis results in a visual manner, allowing users to query detailed information using natural language, thus enabling the output and interaction of the test results.
9. A computer device, characterized in that, The computer device includes a processor and a memory connected to the processor. The memory stores one or more programs, which are executed by the processor to implement the steps in the dynamic road condition diagnosis method for a vehicle as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which are executed by a processor to implement the steps in the method for dynamic diagnosis of road conditions of a vehicle as described in any one of claims 1-7.