Road unmanned inspection system and inspection method based on multi-modal perception and large model

By constructing a collaborative framework of multimodal perception and large models, a collaborative system is built for the vehicle, cloud, and scheduling terminals. This solves the problem of autonomous inspection closed loop in unmanned road patrol systems, achieves efficient defect identification and task scheduling, and improves automation and robustness.

CN121884294APending Publication Date: 2026-04-17WUHAN ZHONGXIANG TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN ZHONGXIANG TECH CO LTD
Filing Date
2025-12-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies lack a closed-loop autonomous inspection system, have insufficient cross-modal self-supervised calibration, and struggle to achieve online self-evolution of models, making it difficult to support unmanned inspections. This results in insufficient automation and robustness of unmanned road patrol systems.

Method used

A road unmanned inspection system based on multimodal perception and large models is adopted. Through the collaborative framework of vehicle, cloud and dispatch terminal, it realizes multimodal data acquisition, autonomous driving control, cross-modal deep recognition, disease knowledge base management and task scheduling, and builds a complete closed loop of multimodal self-supervised perception-temporal prediction-policy generation-vehicle control-online self-learning.

Benefits of technology

It significantly improves the automation and robustness of unmanned road inspection, reduces manual operation, improves the accuracy of disease location and identification, supports predictive maintenance, generates structured inspection reports, reduces system latency and improves reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884294A_ABST
    Figure CN121884294A_ABST
Patent Text Reader

Abstract

The invention provides a road unmanned inspection system and inspection method based on multi-modal perception and a large model. The system comprises a vehicle end system used for multi-modal data acquisition, automatic driving control, real-time perception fusion and edge reasoning of a lightweight large model; the cloud platform is used for cross-modal depth identification, disease knowledge base management, GIS visual display, inspection task management and large model training iteration; the scheduling end is used for task issuing, result auditing, and visual interaction of human-aided decision making and road health assessment. On the basis of the three-layer architecture, a large model collaboration framework is further constructed, and the framework runs across a vehicle end, a cloud end and a scheduling end, is responsible for reasoning collaboration, intelligent supplementary shooting, task priority generation, structured report generation and interpretable reasoning chain construction of a multi-mode large model, and is a key mechanism for realizing closed-loop routing inspection of an intelligent collaboration system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an unmanned road inspection system and method based on multimodal perception and large models. Background Technology

[0002] Currently, mainstream research and applications in the fields of road inspection, road defect identification, and traffic facility operation and maintenance have made some progress in multimodal defect detection, knowledge extraction, or crack segmentation. However, they still suffer from problems such as a lack of closed-loop autonomous inspection systems, insufficient cross-modal self-supervised calibration, difficulty in online self-evolution of models, and difficulty in supporting unmanned inspections. Therefore, they are unable to support truly autonomous driving road inspection systems.

[0003] For example, patent CN202510706030 proposes a road inspection method based on a cloud-based large model. This method fuses multimodal data such as camera data, point cloud data, vehicle vibration data, and traffic flow data, and combines this with prior physical data such as material elastic modulus to achieve defect detection, risk prediction, and maintenance recommendations. It also employs a precise time protocol and spline interpolation to achieve cross-sensor spatiotemporal synchronization. While this technology has certain advantages in multimodal fusion and risk analysis, its processing flow remains at the "analysis layer" and does not involve autonomous control, path planning, or closed-loop execution of the inspection vehicle.

[0004] Patent CN202511030189 focuses on knowledge extraction from large-scale intelligent transportation models. It employs methods such as hyperdimensional vectors, information field theory, and risk potential functions to construct a multimodal entity space for driving behavior understanding and risk analysis. Although it provides new methods for semantic fusion and high-dimensional representation, it is mainly geared towards traffic management scenarios and does not cover the actual control execution of road inspection vehicles, cross-modal self-supervised calibration, or disease evolution prediction.

[0005] Patent CN202511050512 focuses on crack detection and large-scene stitching. This solution utilizes ultra-high-resolution images acquired by drones, employing a two-stage training process: supervised fine-tuning and pseudo-label self-supervised training using MobileSAM. Crack distribution maps are then generated through image cropping, segmentation, and stitching, and a pseudo-label filtering mechanism improves model accuracy. While this technology can handle crack extraction in large-image scenes, its primary application is offline image segmentation and visualization, and it does not involve cross-modal fusion, vehicle control, or inspection task strategy generation. Summary of the Invention

[0006] This invention provides a road unmanned inspection system and method based on multimodal perception and large model, which solves the defects in the existing technology and realizes a complete closed loop of "multimodal self-supervised perception - temporal prediction - strategy generation - vehicle control - online self-learning" for unmanned road inspection system, ensuring a significant improvement in the degree of inspection automation, robustness and long-term reliability.

[0007] In a first aspect, the present invention provides a road unmanned inspection system based on multimodal perception and a large model, comprising a vehicle-side terminal, a cloud-side terminal, and a scheduling terminal, which are coordinated and scheduled by a large model collaborative framework, wherein: The vehicle-side components are used for multimodal data acquisition, autonomous driving control, real-time perception fusion, and edge inference of lightweight large models. The cloud is used for cross-modal deep recognition, disease knowledge base management, GIS visualization, inspection task management, and large model training and iteration; The dispatch terminal is used for task assignment, result review, human-assisted decision-making, and visual interaction for road health assessment.

[0008] According to the present invention, an unmanned road inspection system based on multimodal perception and large model is provided. The system is deployed on an unmanned inspection vehicle. The hardware includes a high-performance integrated display, multiple GMSL2 cameras, LiDAR, GPS / RTK high-precision positioning unit, IMU inertial sensor, and a lightweight large model inference engine. The software includes a data acquisition and autonomous driving module, a multimodal fusion module, a large-scale edge inference module, a fault location module, and a vehicle-side autonomous decision-making module. The data acquisition and autonomous driving module is used for mapping, localization, path tracking, and stable cruise. The multimodal fusion module is used to generate a stable BEV characterization by fusing data from cameras, LiDAR, GPS / RTK, and IMU. The large-scale edge inference module is used to realize disease identification, scene understanding and structured output; The disease location module is used to achieve centimeter-level geolocation based on RTK, IMU, and odometry; The vehicle-side autonomous decision-making module uses a lightweight LLM to determine whether deceleration, reshooting, or perspective adjustment is needed, thus achieving closed-loop control on the acquisition side.

[0009] According to the present invention, an unmanned road inspection system based on multimodal perception and a large model is provided, wherein the multimodal fusion module is specifically used for: By acquiring camera images at any given moment using multiple GMSL2 cameras, and transforming the pixels of any camera image into camera coordinates based on the camera intrinsic parameter matrix, the coordinates of the disease point in the camera coordinate system are obtained. By using the rotation and offset matrices from the camera coordinate system to the vehicle coordinate system, the coordinates of the defect point in the camera coordinate system are transformed to obtain the position of the defect point in the vehicle coordinate system. Obtain the laser point cloud at any given time, and transform the laser point cloud using the rotation and translation matrices of the laser point cloud to the vehicle body coordinate system to obtain the position of the laser point cloud in the vehicle body coordinate system. Obtain the horizontal and vertical coordinate grid resolution of the BEV, project the positions of the defect points in the vehicle coordinate system and the positions of the laser point cloud in the vehicle coordinate system onto the BEV grid index, and obtain the BEV features of the laser radar point cloud and the BEV features of the camera image. The weighted summation of the BEV features from the lidar point cloud and the camera image is used to obtain the fused BEV features.

[0010] According to the present invention, an unmanned road inspection system based on multimodal perception and large model is provided, wherein the large-scale edge inference module is specifically used for: The input vector is composed of the fused BEV features, the position of the laser point cloud in the vehicle coordinate system, the absolute position of RTK / GNSS, and the rotation matrix of the IMU attitude from the vehicle to the world coordinate system at any given time. The normalization function based on the self-attention mechanism is used to abstract the model from the input vector to obtain attention weights, and the weighted sum is used to obtain the model output. The model output is normalized and solved by using a parameter set to obtain the disease category prediction results; Geometric regression dimensions are obtained by regressing the length, width, and depth of the disease image based on the model output; The confidence score is obtained based on the model output. By combining the disease category prediction results, geometric regression dimensions, confidence output, and bounding boxes, a structured candidate set output is obtained.

[0011] According to the present invention, an unmanned road inspection system based on multimodal perception and large model is provided, wherein the defect location module is specifically used for: Obtain the location of the defect points in the vehicle coordinate system; Using the RTK / GNSS absolute position at any given time, the rotation matrix of the IMU attitude from the vehicle body to the world coordinate system at any given time, and the position of the defect point in the vehicle body coordinate system, the position of the defect in the world coordinate system can be obtained. Based on vehicle speed, the position of the defects in the world coordinate system is compensated for in a temporal manner, and a structured defect record is output.

[0012] According to the present invention, an unmanned road inspection system based on multimodal perception and large model is deployed in the cloud on the system's central computing and management platform, including a road defect review and correction module, a road defect knowledge base, a GIS visualization module, an inspection task management module, a large model training and iteration module, and a road health index assessment module, wherein: The defect review and correction module is used to perform high-precision identification and correction of the results from the vehicle end. The road disease knowledge base is used to record the characteristics, evolution patterns, and threshold systems of various diseases; The GIS visualization module is used to display the trajectory, location of defects, density of defects, and current road conditions; The inspection task management module is used for task assignment, result collection, and progress monitoring; The large model training and iteration module is responsible for model updates, retraining, and version management; The road health index assessment module is used to generate road inspection reports and health scores; The cloud also includes receiving structured defect records from the vehicle, quantifying and scoring each defect, aggregating them across road segments, and calculating defect scores, road segment health indices, and task priorities.

[0013] According to the present invention, an unmanned road inspection system based on multimodal perception and large model is provided, which calculates road defect scores and road segment health indices, including: The basic score of the disease is obtained by weighted summation based on the area, length, width and depth of the disease. Determine the type factor, and obtain type correction based on the type factor and the basic disease score; The type correction is adjusted using confidence level to obtain the final single disease severity. The total severity of a road defect is obtained by aggregating the severity of multiple individual defects within a preset road section area. Damage density is obtained by dividing the total severity of damage by the preset road section area; The road health index is derived from the damage density and the maximum allowable density threshold.

[0014] According to the present invention, an unmanned road inspection system based on multimodal perception and large model is provided, which calculates task priority, including: Extract the severity of a single disease at any two different times to obtain the growth rate of a single disease; Linear prediction of a single disease is obtained by combining the severity and growth rate of a single disease. Task priority scores are obtained based on linear prediction of single disease, single disease growth rate, and road health index.

[0015] According to the present invention, an unmanned road inspection system based on multimodal perception and large model is provided, wherein the scheduling terminal is specifically used for: It provides functions such as configuring and issuing inspection tasks, adjusting task priorities, reviewing and manually verifying defects, displaying daily / weekly / monthly road network inspection reports, and providing anomaly warnings and maintenance strategies.

[0016] Secondly, the present invention also provides a method for unmanned road inspection based on multimodal perception and large model, comprising: Lightweight inference is performed on the vehicle side, and the decision to reshoot is made based on the confidence level of the defect, the shooting angle, and the vehicle posture. The large language model automatically generates task ranking based on the defect level, road level, and historical trends, and outputs daily, weekly, road segment-level reports and health assessment reports. High-precision inference is performed in the cloud, predicting future risks based on time series, disease types and multimodal characteristics, outputting the reasoning basis for each inspection decision, and synchronizing the results to the scheduling terminal; Daily reports, weekly reports, road segment reports, and health assessment reports are automatically generated by the dispatching terminal.

[0017] The unmanned road inspection system and method based on multimodal perception and large model provided by this invention have the following beneficial effects: (1) Achieve automated closed-loop inspection tasks and significantly reduce manual intervention. Through the linkage of autonomous decision-making of the vehicle-side big model and task scheduling in the cloud, a fully automated closed loop of "collection → judgment → re-shooting → uploading → priority scheduling" is achieved, eliminating the need for manual judgment of re-shooting timing and reducing manual operation by more than 80%.

[0018] (2) Centimeter-level lesion localization based on multimodal fusion significantly improves localization accuracy. The vehicle-side camera, lidar, RTK and IMU are fused to generate a unified BEV representation, and centimeter-level lesion localization is achieved through motion compensation and attitude calculation. The localization error is reduced by more than 60% compared with traditional vision methods.

[0019] (3) Lightweight large model edge inference enables real-time recognition and reduces dependence on the cloud. Through model distillation, MOE and Q-LoRA quantization, large models can run on edge devices such as Orin to achieve real-time recognition capabilities, with reduced latency compared to pure cloud inference solutions.

[0020] (4) Introduce a large-scale model collaboration framework to improve the accuracy of disease identification and the efficiency of task scheduling. Cross-end inference collaboration allows the vehicle-side and cloud-side models to complement each other: the vehicle-side is responsible for real-time identification and decision-making, while the cloud-side is responsible for high-precision supplementary judgment. Ultimately, this improves the overall identification accuracy and task scheduling efficiency.

[0021] (5) To realize automated and structured inspection reports and improve the efficiency of road asset management. The system automatically outputs inspection reports containing the type, size, location, grade and road health index of defects, replacing manual compilation and improving the efficiency of report generation by more than 10 times.

[0022] (6) Explainable reasoning chains improve the auditability and credibility of the system. The system generates natural language reasoning chains in key steps such as identification, positioning, reshooting, and scheduling, which facilitates the review by regulatory authorities and improves the transparency and credibility of the system.

[0023] (7) Support road health trend prediction and realize predictive maintenance. Based on the time series model of disease evolution, the health index can be predicted, which can identify potential high-risk road sections in advance and help optimize road maintenance budget and operation and maintenance plan. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0025] Figure 1 This is a schematic diagram of the structure of the unmanned road inspection system based on multimodal perception and large model provided by the present invention; Figure 2 This is a flowchart illustrating the unmanned road inspection method based on multimodal perception and large model provided by the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0027] To address the shortcomings of existing technologies, this invention constructs an unmanned road inspection system that can realize a complete closed loop of "multimodal self-supervised perception - temporal prediction - strategy generation - vehicle control - online self-learning", ensuring a significant improvement in the degree of automation, robustness and long-term reliability of inspection.

[0028] This invention proposes an intelligent collaborative system for road inspection scenarios. The system adopts a three-layer architecture of "vehicle-cloud-dispatch terminal", and on this basis, it builds a cross-terminal large-scale model collaborative framework to realize automatic collection, intelligent identification, structured expression, precise positioning, health assessment and task closed-loop scheduling of road defects.

[0029] The intelligent collaborative system consists of three parts: the vehicle-side system for multimodal data acquisition, autonomous driving control, real-time perception fusion, and edge inference of lightweight large models; the cloud platform for cross-modal deep recognition, disease knowledge base management, GIS visualization, inspection task management, and large model training iteration; and the dispatch terminal for task issuance, result review, human-assisted decision-making, and visualized interaction of road health assessment.

[0030] Based on the above three-layer architecture, this invention further constructs a large-scale model collaboration framework. This framework runs across the vehicle, cloud, and scheduling terminals and is responsible for multimodal large-scale model inference collaboration, intelligent reshooting, task priority generation, structured report generation, and interpretable inference chain construction. It is a key mechanism for intelligent collaborative systems to achieve closed-loop inspection.

[0031] Figure 1 This is a schematic diagram of the structure of the unmanned road inspection system based on multimodal perception and large model provided in an embodiment of the present invention, as shown below. Figure 1 As shown, it includes: (a) Vehicle end The vehicle-mounted system is deployed on an unmanned inspection vehicle and consists of the following components: a high-performance integrated display, four GMSL2 cameras, a 64-line LiDAR, a GPS / RTK high-precision positioning unit, an IMU inertial sensor, and a lightweight large-model inference engine.

[0032] On the software side, the vehicle-side functional modules integrate six parts: vehicle-side data collection and autonomous driving module, multimodal fusion module, large-scale edge inference module, defect localization module, and vehicle-side autonomous decision-making module. Details are as follows: The vehicle-side data collection and autonomous driving module is responsible for mapping, localization, path tracking, and stable cruise control.

[0033] The multimodal fusion module generates a stable BEV characterization by fusing data from cameras, LiDAR, GPS / RTK, and IMU.

[0034] The large model edge reasoning module is used to realize disease identification, scene understanding and structured output.

[0035] The disease location module achieves centimeter-level geolocation based on RTK, IMU, and odometer.

[0036] The vehicle-side autonomous decision-making module uses a lightweight LLM to determine whether deceleration, reshooting, or perspective adjustment is needed, thus achieving closed-loop control on the acquisition side.

[0037] In terms of algorithms, the vehicle-side involves multimodal fusion models, disease localization models, and lightweight large-scale model inference, as detailed below: 1. Vehicle-side multimodal fusion module: The core objective of the vehicle-side multimodal fusion module is to integrate multiple GMSL2 cameras, LiDAR point clouds, RTK high-precision positioning data, and IMU attitude data. The entire model involves five core stages: data preprocessing, spatiotemporal registration, feature extraction, feature fusion, and output application.

[0038] (1) Pixel to camera coordinate transformation For a pixel point p=(u,v,1) in the image T : p Where P C Let K be the coordinates of the defect point in the camera coordinate system, and K be the camera intrinsic parameter matrix. (2) Camera coordinates to vehicle coordinates conversion

[0039] Where P veh It is the location of the defect point in the vehicle body coordinate system, R cv t cv This is the rotation and offset matrix from the camera coordinate system to the vehicle coordinate system.

[0040] (3) Lidar point cloud projection onto vehicle coordinates Lidar points Transform to vehicle body unified coordinate system for registration

[0041] Where R lv , t lv The rotation and translation of the external parameters of the LiDAR → vehicle body can be obtained through calibration.

[0042] (4) BEV mesh mapping (unified planar representation) Projecting vehicle coordinates (x, y, z) onto the BEV mesh index

[0043] Where r x ,r y BEV grid resolution (m / grid) (5) Fusion characteristics

[0044] Where F lidar For the BEV features of the lidar point cloud, F cam For the BEV features of the camera image, F fuse For the merged BEV characteristics, Vehicle-side fusion module output: F fuse (For subsequent disease monitoring and location) In summary, the input to the vehicle-side multimodal fusion module is (I t P t G t R t v t ), I t The camera image at time t, P t G represents the Lidar point cloud at time t. t R represents the absolute RTK / GNSS position at time t. t The rotation matrix v represents the IMU attitude from the vehicle body to the world coordinate system at time t. t Represents vehicle speed; output is (F fuse )F fuse This represents the fusion output characteristics of multimodal data from the vehicle.

[0045] 2. Lightweight large model inference: (1) Input vector construction

[0046] (2) Model abstraction

[0047] (3) Disease category prediction

[0048] (4) Geometric Dimension Regression

[0049] (5) Confidence output

[0050] (6) Output (structured)

[0051] Where C represents the semantic category of the disease (discrete label); y i For category prediction; L, W, and D represent the length, width, and depth (meters) of the lesion, respectively; C i ∈[0,1]: Confidence of the large model; b i This represents the detection bounding box or pixel-level mask.

[0052] In summary, the input to lightweight large model inference is the output of the vehicle-side multimodal fusion module (F). fuse ), F here fuse Represents the BEV / ROI fusion features, outputting a structured candidate set. y i Represents disease category, Represents the geometric dimensions obtained from the regression. Represents confidence level. Represents a bounding box or segmentation mask.

[0053] 3. Vehicle-side defect location module: The vehicle-side positioning uses RTK as the absolute coordinate reference and IMU as the attitude reference. Based on the monitoring result C, combined with the defect point representation in the vehicle coordinate system, a unified positioning is achieved from the image coordinate system to the vehicle coordinate system to the world coordinate system.

[0054] (1) The defective pixels are projected onto the vehicle coordinate system P. veh Assume the lesion corresponds to camera coordinate point P on the image plane. c Transformed to the vehicle coordinate system using extrinsic parameters:

[0055] (2) Projecting the vehicle coordinate system onto the world coordinate system Using the vehicle position information G provided by RTK t With IMU pose R t :

[0056] Where P w The location of the disease in the world coordinate system (3) Timing compensation If there is a data collection delay

[0057] (4) Output structured disease records

[0058] In summary, the input to the vehicle-side defect location module is (C, P) t G t ,R t ,v t ,T c→v ), P t Let G be the point cloud data at time t. t For RTK world coordinates, R t For IMU attitude (rotation matrix), v t T represents vehicle speed. c→v Represents the total transformation matrix from camera to vehicle body; the output is yi represents the category prediction; L, W, and D represent the length, width, and depth (meters) of the lesion, respectively; Ci∈[0,1] represents the confidence level of the large model. This represents the coordinates of the defect in the vehicle body coordinate system. This represents the coordinates of the defect in the world coordinate system.

[0059] (ii) Cloud The cloud serves as the system's central computing and management platform, comprising a disease review and correction module, a road disease knowledge base, a GIS visualization module, an inspection task management module, a large model training and iteration module, and a road health index assessment module.

[0060] Disease review and correction module: performs high-precision identification and correction of the results from the vehicle end.

[0061] Road disease knowledge base: records the characteristics, evolution patterns and threshold systems of various diseases.

[0062] GIS visualization module: Displays the trajectory, location of defects, density of defects, and current road conditions.

[0063] Inspection task management module: task assignment, result collection, and progress monitoring.

[0064] Large Model Training and Iteration Module: Responsible for model updates, retraining, and version management.

[0065] Road Health Index Assessment Module: Generates road inspection reports and health scores.

[0066] In terms of algorithms, the cloud receives structured defect data (R) from the vehicle, quantifies and scores each defect, aggregates the data for each road segment, and calculates the defect score (DSS), road segment health index (RQI), and task priority (PS).

[0067] 1. Road Defect Score (DSS) and Road Section Health Index (RQI) (1) Basic score of single disease

[0068] Where A represents the diseased area, and S... b Basic score for disease (2) Type correction (weighting by disease category)

[0069] Where T(y) is the type factor, S b This is the basic score for the disease.

[0070] (3) Confidence level correction yields the final single disease severity DSS.

[0071] DSS: Single Disease Severity Score (0–100) (4) Road segment aggregation and ROI Road section area A road It contains n diseases:

[0072] Damage density:

[0073] Road health index:

[0074] Among them, RQI is the road health index of the road segment (0–100, 100 is the best), D max This is the maximum allowable density threshold.

[0075] 2. Trend Forecasting and Task Prioritization (1) Single disease growth rate: the severity of disease i at times t1 and t2 (DSS)

[0076] (2) Linear prediction of single disease:

[0077] (3) Task priority scoring:

[0078] Where γ1, γ2, and γ3 are weights (configurable), and v is the disease growth rate.

[0079] In summary, the input to the cloud is (R, A) road R represents disease, A road The output is (DSS, RQI, PS), where DSS represents the disease score, RQI represents the road segment health index, and PS represents the task priority score.

[0080] (III) Scheduling Terminal The dispatch terminal is designed for road management personnel or platform operators, providing functions such as configuring and issuing inspection tasks, adjusting task priorities, reviewing and manually verifying defects, displaying daily / weekly / monthly road network inspection reports, and providing anomaly warnings and maintenance strategies.

[0081] Therefore, the input to the dispatcher is (DSS, RQI, PS), and the output is (refresh shot / dispatch / maintenance command).

[0082] Figure 2 This is a flowchart illustrating the unmanned road inspection method based on multimodal perception and large model provided in an embodiment of the present invention, as shown below. Figure 2 As shown, it includes: Step 100: Lightweight inference is performed on the vehicle side. It determines whether to reshoot based on the confidence level of the defect, the shooting angle, and the vehicle posture. The large language model automatically generates task sorting based on the defect level, road level, and historical trends, and outputs daily, weekly, road segment reports and health assessment reports. Step 200: High-precision inference is performed in the cloud to predict future risks based on time series, disease type and multimodal characteristics, output the reasoning basis for each inspection decision, and synchronize the results to the scheduling terminal; Step 300: The dispatch terminal automatically generates daily reports, weekly reports, road segment reports, and health assessment reports.

[0083] Specifically, the large-scale model collaborative framework proposed in this embodiment of the invention is not a single independent endpoint, but rather a set of intelligent decision-making and reasoning mechanisms spanning the vehicle end, cloud end, and scheduling end, specifically including: Cross-device inference collaboration: First, lightweight inference is performed on the vehicle side, then high-precision inference is performed on the cloud side, and finally the results are synchronized to the scheduling end from the cloud side.

[0084] Intelligent reshooting and closed-loop control: The vehicle-side LLM determines whether to reshoot based on factors such as defect confidence, shooting angle, and vehicle posture.

[0085] Automatic task priority generation: The large model automatically generates task sorting based on disease level, road level, and historical trends.

[0086] Structured inspection report generation: Automatically generates daily reports, weekly reports, road section-level reports, and health assessment reports.

[0087] Road health prediction model: predicting future risks based on time series, disease type and multimodal characteristics.

[0088] Explainable reasoning chain generation: Outputs the reasoning basis for each inspection decision, improving security and auditing capabilities.

[0089] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0090] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-modal perception and large model based road unmanned inspection system, characterized in that, This includes the vehicle-side, cloud-side, and scheduling-side components, all coordinated by a large-scale model coordination framework. The vehicle-side components are used for multimodal data acquisition, autonomous driving control, real-time perception fusion, and edge inference of lightweight large models. The cloud is used for cross-modal deep recognition, disease knowledge base management, GIS visualization, inspection task management, and large model training and iteration; The dispatch terminal is used for task assignment, result review, human-assisted decision-making, and visual interaction for road health assessment.

2. The multi-modal perception and large model based road unmanned inspection system according to claim 1, wherein, The vehicle-mounted system is deployed in unmanned inspection vehicles. The hardware includes a high-performance integrated display, multiple GMSL2 cameras, LiDAR, GPS / RTK high-precision positioning unit, IMU inertial sensor, and a lightweight large-model inference engine. The software includes a data acquisition and autonomous driving module, a multimodal fusion module, a large-scale edge inference module, a defect localization module, and a vehicle-mounted autonomous decision-making module. The data acquisition and autonomous driving module is used for mapping, localization, path tracking, and stable cruise. The multimodal fusion module is used to generate a stable BEV characterization by fusing data from cameras, LiDAR, GPS / RTK, and IMU. The large-scale edge inference module is used to realize disease identification, scene understanding and structured output; The disease location module is used to achieve centimeter-level geolocation based on RTK, IMU, and odometry; The vehicle-side autonomous decision-making module uses a lightweight LLM to determine whether deceleration, reshooting, or perspective adjustment is needed, thus achieving closed-loop control on the acquisition side.

3. The multi-modal perception and large model based road unmanned inspection system according to claim 2, wherein, The multimodal fusion module is specifically used for: By acquiring camera images at any given moment using multiple GMSL2 cameras, and transforming the pixels of any camera image into camera coordinates based on the camera intrinsic parameter matrix, the coordinates of the disease point in the camera coordinate system are obtained. By using the rotation and offset matrices from the camera coordinate system to the vehicle coordinate system, the coordinates of the defect point in the camera coordinate system are transformed to obtain the position of the defect point in the vehicle coordinate system. Obtain the laser point cloud at any given time, and transform the laser point cloud using the rotation and translation matrices of the laser point cloud to the vehicle body coordinate system to obtain the position of the laser point cloud in the vehicle body coordinate system. Obtain the horizontal and vertical coordinate grid resolution of the BEV, project the positions of the defect points in the vehicle coordinate system and the positions of the laser point cloud in the vehicle coordinate system onto the BEV grid index, and obtain the BEV features of the laser radar point cloud and the BEV features of the camera image. The weighted summation of the BEV features from the lidar point cloud and the camera image is used to obtain the fused BEV features.

4. The multi-modal perception and large model based road unmanned inspection system according to claim 3, wherein, The large-scale edge inference module is specifically used for: The input vector is composed of the fused BEV features, the position of the laser point cloud in the vehicle coordinate system, the absolute position of RTK / GNSS, and the rotation matrix of the IMU attitude from the vehicle to the world coordinate system at any given time. The normalization function based on the self-attention mechanism is used to abstract the model from the input vector to obtain attention weights, and the model output is obtained by weighted summation. The model output is normalized and solved by using a parameter set to obtain the disease category prediction results; Geometric regression dimensions are obtained by regressing the length, width, and depth of the disease image based on the model output; The confidence score is obtained based on the model output. By combining the disease category prediction results, geometric regression dimensions, confidence output, and bounding boxes, a structured candidate set output is obtained.

5. The multi-modal perception and large model based road unmanned inspection system according to claim 4, wherein, The disease location module is specifically used for: Obtain the location of the defect points in the vehicle coordinate system; Using the RTK / GNSS absolute position at any given time, the rotation matrix of the IMU attitude from the vehicle body to the world coordinate system at any given time, and the position of the defect point in the vehicle body coordinate system, the position of the defect in the world coordinate system can be obtained. Based on vehicle speed, the position of the defects in the world coordinate system is compensated for in a temporal manner, and a structured defect record is output.

6. The multi-modal perception and large model based road unmanned inspection system according to claim 1, wherein, The cloud-based central computing and management platform of the system includes a road defect review and correction module, a road defect knowledge base, a GIS visualization module, an inspection task management module, a large model training and iteration module, and a road health index assessment module, among which: The defect review and correction module is used to perform high-precision identification and correction of the results from the vehicle end. The road disease knowledge base is used to record the characteristics, evolution patterns, and threshold systems of various diseases; The GIS visualization module is used to display the trajectory, location of defects, density of defects, and current road conditions; The inspection task management module is used for task assignment, result collection, and progress monitoring; The large model training and iteration module is responsible for model updates, retraining, and version management; The road health index assessment module is used to generate road inspection reports and health scores; The cloud also includes receiving structured defect records from the vehicle, quantifying and scoring each defect, aggregating them across road segments, and calculating defect scores, road segment health indices, and task priorities.

7. The multi-modal perception and large model based road unmanned inspection system according to claim 6, wherein, The calculation of road disease scores and road section health index includes: The basic score of the disease is obtained by weighted summation based on the area, length, width and depth of the disease. Determine the type factor, and obtain type correction based on the type factor and the basic disease score; The type correction is adjusted using confidence level to obtain the final single disease severity. The total severity of a road defect is obtained by aggregating the severity of multiple individual defects within a preset road section area. Damage density is obtained by dividing the total severity of road damage by the preset road section area. The road health index is derived from the damage density and the maximum allowable density threshold.

8. The multi-modal perception and large model based road unmanned inspection system according to claim 7, wherein, Calculate task priorities, including: Extract the severity of a single disease at any two different times to obtain the growth rate of a single disease; Linear prediction of a single disease is obtained by combining the severity and growth rate of a single disease. Task priority scores are obtained based on linear prediction of single disease, single disease growth rate, and road health index.

9. The multi-modal perception and large model based road unmanned inspection system according to claim 1, wherein, The scheduling terminal is specifically used for: It provides functions such as configuring and issuing inspection tasks, adjusting task priorities, reviewing and manually verifying defects, displaying daily / weekly / monthly road network inspection reports, and providing anomaly warnings and maintenance strategies.

10. A method for unmanned road inspection based on multimodal perception and large model, based on the unmanned road inspection system based on multimodal perception and large model as described in any one of claims 1 to 9, characterized in that, include: Lightweight inference is performed on the vehicle side, and the decision to reshoot is made based on the confidence level of the defect, the shooting angle, and the vehicle posture. The large language model automatically generates task ranking based on the defect level, road level, and historical trends, and outputs daily, weekly, road segment-level reports and health assessment reports. High-precision inference is performed in the cloud, predicting future risks based on time series, disease types and multimodal characteristics, outputting the reasoning basis for each inspection decision, and synchronizing the results to the scheduling terminal; Daily reports, weekly reports, road segment reports, and health assessment reports are automatically generated by the dispatching terminal.

Citation Information

Patent Citations

  • Large-scene road crack distribution map construction method, device and equipment

    CN120543562B

  • Large model-based road inspection method, apparatus and device, and storage medium

    CN120564168A

  • Multi-modal data knowledge extraction method based on intelligent traffic large model

    CN120805063A