A low-altitude intelligent inspection method based on multi-modal data fusion

The low-altitude intelligent inspection method based on multimodal data fusion solves the problems of low efficiency and low accuracy in existing road inspection technologies, and realizes efficient and intelligent road event identification and response, thereby improving inspection efficiency and safety.

CN122511089APending Publication Date: 2026-08-04BEIJING BADALING LOW ALTITUDE TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BADALING LOW ALTITUDE TECHNOLOGY DEVELOPMENT CO LTD
Filing Date
2026-05-08
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing road inspection technologies rely on manual patrols, fixed monitoring, and single sensors, which suffer from low efficiency, high cost, limited coverage, significant weather impact, low recognition accuracy, and insufficient fusion of multi-source data. This results in high rates of missed and false judgments, failing to meet the demands for intelligent and efficient road inspections.

Method used

A multimodal data fusion method is adopted, which simultaneously collects visible light images, infrared thermal images, lidar point clouds, GPS trajectories and audio data through a low-altitude inspection platform equipped with multimodal sensors. Data preprocessing and spatiotemporal alignment are performed, and data integration and feature extraction are carried out by combining a Transformer+CNN hybrid architecture model to generate road event recognition nodes. Event recognition and early warning are performed through a weighted fusion strategy.

Benefits of technology

It achieves high-precision identification and real-time response to road incidents in complex environments, with an identification latency of less than 100ms and a positioning accuracy of less than 5m, significantly improving inspection efficiency and accuracy, reducing labor costs, and minimizing the traffic risk of delayed incident handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122511089A_ABST
    Figure CN122511089A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of low-altitude intelligent inspection, and relates to a low-altitude intelligent inspection method based on multi-modal data fusion.The method comprises the following steps: S1, acquiring multi-modal data related to road inspection, pre-processing the data and performing space-time alignment, and establishing a multi-modal data integration library and a data correlation index according to the processing result; the present application integrates multi-modal data such as visible light, infrared, laser radar, GPS trajectory and audio, uses RTK / IMU tight coupling positioning to achieve high-precision space-time alignment, breaks the limitation of a single sensor, and fully utilizes the complementary advantages of various modal data - visible light images capture visual details, infrared thermal imaging is suitable for night / bad weather, laser radar provides three-dimensional size and distance information, and GPS trajectory reflects traffic conditions, thereby significantly improving the recognition ability of road events in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of low-altitude intelligent inspection technology, and more specifically, to a low-altitude intelligent inspection method based on multimodal data fusion. Background Technology

[0002] Road inspection is a core component of ensuring traffic safety and improving road maintenance efficiency. Its core objective is to promptly detect and handle various road incidents to prevent secondary accidents or traffic paralysis caused by delayed incident handling.

[0003] Current road inspection technologies mainly rely on manual patrols, fixed surveillance cameras, or low-altitude detection using single sensors, which have significant limitations. Manual patrols are inefficient and costly, making it difficult to achieve large-scale, all-weather coverage. Fixed surveillance has blind spots and is insufficient for capturing remote road sections or sudden moving events. Single sensors (such as visible light cameras) are easily affected by weather (rain, fog, strong light at night), resulting in low event recognition accuracy and an inability to simultaneously monitor both macroscopic scenes and capture microscopic event details. Furthermore, existing technologies lack a deep fusion mechanism for multi-source data, making it difficult to integrate complementary information from heterogeneous data such as images, point clouds, and time-series trajectories. This leads to a high rate of missed and false judgments of road events, failing to meet the needs of intelligent and efficient road inspection. Therefore, this paper proposes a low-altitude intelligent inspection method based on multimodal data fusion. Summary of the Invention

[0004] The purpose of this invention is to provide a low-altitude intelligent inspection method based on multimodal data fusion to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, a low-altitude intelligent inspection method based on multimodal data fusion is provided, comprising the following steps: S1. Acquire multimodal data related to road inspection, preprocess and align the data spatiotemporally, and establish a multimodal data integration library and data association index based on the processing results; S2. The integrated multimodal data is associated and matched according to the road event type, the effective collection time of each modality is counted, the inspection area is divided in combination with the road topology information, and the reference data node of each inspection area is selected. S3. Integrate all related data to perform accuracy verification on the modal data corresponding to the baseline data node. Based on the verification results, retain the data that meets the accuracy standard and substitute them into the corresponding inspection area and baseline data node. S4. In each inspection area, road event identification nodes are generated based on the benchmark data nodes' compliance data and multimodal data features. At the same time, event feature data of each identification node is predicted based on the benchmark data nodes' compliance data. S5. Based on the compliance data retained in S3, perform same-node verification on the identification node data predicted in S4 in each inspection area, and adjust the prediction method according to the verification results until all identification node data meet the compliance data standards. S6. The newly collected multimodal road inspection data is matched with the corresponding inspection area and identification node to carry out event identification. The road event type and level are determined by combining the multimodal feature fusion score and event confidence. The event location, handling suggestions and early warning information are output simultaneously.

[0006] As a further improvement to this technical solution, in S1, a low-altitude inspection platform equipped with multimodal sensors is used to simultaneously collect multimodal data related to road inspection. The multimodal data includes visible light image data, infrared thermal imaging data, lidar point cloud data, GPS trajectory data, and audio data. Perform denoising, image enhancement, and size normalization operations on visible light image data and infrared thermal imaging data; perform distortion correction, ground point filtering, and feature extraction operations on lidar point cloud data. Perform coordinate calibration and time synchronization on GPS trajectory data; perform noise reduction and feature extraction on audio data; Based on the timestamp of GPS trajectory data, a cross-correlation algorithm is used to achieve time synchronization of visible light images, infrared thermal imaging and lidar point clouds. Based on the three-dimensional coordinates of lidar point clouds, spatial registration of image data and point cloud data is achieved through perspective projection transformation. The multimodal data integration library stores preprocessed multimodal data according to data type, collection time, and road segment. The data association index is built according to time, location, and modality type to realize fast association query of multimodal data.

[0007] As a further improvement to this technical solution, in S2, visible light image data and infrared thermal imaging data are associated with the dimensions of road congestion, traffic accidents, and obstacle-occupying events, while lidar point cloud data are associated with the dimensions of road surface damage and road obstacle size measurement. The GPS trajectory data is correlated with the vehicle speed and traffic density calculation dimensions; the audio data is correlated with the abnormal sound recognition dimensions such as collision sounds and warning sounds, and a mapping rule between data and road event types is established. The effective acquisition time is the time during which data for each modality is continuously and completely acquired and the data quality meets the recognition requirements. Statistics are completed by checking the data integrity and clarity frame by frame. Based on road topology information and multimodal data coverage, a regional division threshold is set to divide areas with continuous road segments and continuous data coverage into the same inspection area, ensuring that there are no blind spots in multimodal data coverage within a single inspection area. Selecting a baseline data node means choosing the multimodal data set with the longest effective acquisition time, the highest data integrity, and the best spatiotemporal alignment accuracy as the baseline data node for each inspection area.

[0008] As a further improvement to this technical solution, in step S3, all multimodal data matched to each road event type in step S2 are summarized and grouped and classified according to the inspection area and the reference data node; Acquire industry standards, traffic management regulations, and historical valid inspection data for road event recognition. Set accuracy standard thresholds for each modality of data: event feature recognition error ≤5% for visible light image data, temperature recognition error ≤2℃ for infrared thermal imaging data, and size measurement error ≤3% for lidar point cloud data.

[0009] As a further improvement to this technical solution, in step S3, the modal data of each reference data node are compared with the corresponding accuracy standard threshold, and the data deviation rate is calculated. If the deviation rate is lower than the accuracy standard threshold, it is determined that the accuracy meets the standard; otherwise, it is determined that the accuracy does not meet the standard. If the multimodal data of the reference data nodes in the inspection area does not meet the standard, then a new multimodal data group with the highest accuracy and longest effective acquisition time is selected in the inspection area and substituted into the corresponding inspection area as a new reference data node.

[0010] As a further improvement to this technical solution, in S4, within each inspection area, taking the compliant data of the reference data node as the reference origin, and based on the differences in road event type characteristics, uniformly distributed road event identification nodes are generated according to the principle of equal intervals, so that the generated identification nodes cover all road sections and key locations within the inspection area. A Transformer+CNN hybrid architecture model is constructed. Local detail features are extracted from the feature maps of visible light images and infrared thermal imaging through CNN, and global correlation features are extracted from the 3D features of LiDAR point clouds and the temporal features of GPS trajectories through Transformer. After fusion, multimodal joint features are obtained. The model inputs are the benchmark data nodes' compliance data and road event type features, and the historically labeled road event feature data are used as training labels. After the model training is completed, the multimodal basic data of the identification nodes are input, and the predicted event feature data corresponding to each identification node is output.

[0011] As a further improvement to this technical solution, in step S5, the compliance data retained in step S3 is matched to the identification nodes of the corresponding inspection area according to the road location and event type, and the actual event feature data corresponding to each identification node is obtained. The predicted event feature data of the identification node is compared with the actual event feature data to determine the deviation. A verification threshold is set. If the deviation value is lower than the verification threshold, the identification node is determined to be verified successfully. If the deviation value is higher than the verification threshold, the identification node is determined to be verified unsuccessfully. For identification nodes that fail verification, the deviation features between the predicted data and the actual event feature data are extracted, input into the multimodal fusion prediction model, the feature weights and training parameters of the model are re-optimized, and the predicted data for the identification node is generated again and verified. Repeat the prediction, verification, and optimization process until the prediction data of all identification nodes in the inspection area meets the compliant data standards. Then, synchronize the optimized model and qualified prediction data to the corresponding database.

[0012] As a further improvement to this technical solution, in step S6, the acquisition location, time features and modality type of the newly acquired multimodal data are extracted, the similarity between the new data and the reference data nodes of each inspection area is calculated, the inspection area with the highest similarity is selected, and then the identification node with the closest position in that area is matched. A weighted fusion strategy is used to calculate the multimodal feature fusion score: weights are set according to the contribution of each modality data to the recognition of different road events, and the event recognition scores of each modality are weighted and summed to obtain the multimodal feature fusion score; Based on the consistency of multimodal features and the model prediction probability, event confidence is generated. Event grading standards are set by combining multimodal feature fusion scores and event confidence, resulting in Level 1, Level 2, and Level 3 events. The system pinpoints the exact location of the incident, generates handling suggestions based on the incident type and severity, and outputs early warning information to the traffic management platform via a wireless communication module.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This low-altitude intelligent inspection method based on multimodal data fusion integrates multimodal data such as visible light, infrared, lidar, GPS trajectory, and audio, and uses RTK / IMU tightly coupled positioning to achieve high-precision spatiotemporal alignment. This breaks the limitations of a single sensor and fully leverages the complementary advantages of each modal data: visible light images capture visual details of events, infrared thermal imaging adapts to nighttime / severe weather, lidar provides three-dimensional size and distance information, and GPS trajectory reflects traffic flow status, significantly improving the ability to identify road events in complex environments.

[0014] 2. In this low-altitude intelligent inspection method based on multimodal data fusion, the stability and accuracy of the event recognition model are ensured by dividing the inspection area and setting benchmark data nodes, combined with the closed-loop logic of accuracy verification and prediction optimization. The multimodal fusion prediction model adopts a Transformer+CNN hybrid architecture, which takes into account both local details and global correlation features. With the help of dynamic weight fusion strategy, it effectively solves the problems of missed judgment and high false judgment rate.

[0015] 3. This low-altitude intelligent inspection method based on multimodal data fusion realizes intelligent full-link operation of real-time road event identification, level determination, precise positioning, handling suggestions, and early warning output. The event identification response delay is <100ms and the positioning accuracy is ≤5m. It can quickly respond to major road events, provide accurate decision-making basis for traffic management and operation and maintenance departments, significantly improve road inspection efficiency, reduce labor costs, and reduce traffic risks caused by delayed event handling. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating a low-altitude intelligent inspection method based on multimodal data fusion according to the present invention. Figure 2 This is a flowchart of S1 of the present invention; Figure 3 This is a flowchart of S2 of the present invention; Figure 4 This is a flowchart of S3 of the present invention; Figure 5 This is a flowchart of S4 of the present invention; Figure 6 This is a flowchart of S5 of the present invention; Figure 7 This is a flowchart of S6 of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see Figures 1-7 As shown, the purpose of this embodiment is to provide a low-altitude intelligent inspection method based on multimodal data fusion, including the following steps: S1. Acquire multimodal data related to road inspection, preprocess and align the data spatiotemporally, and establish a multimodal data integration library and data association index based on the processing results; In S1, a low-altitude inspection platform equipped with multimodal sensors synchronously collects multimodal data related to road inspection. The multimodal data includes visible light image data, infrared thermal imaging data, lidar point cloud data, GPS trajectory data, and audio data. Deploy a low-altitude inspection platform (drone) equipped with multimodal sensors to the target road inspection area, and simultaneously start the visible light camera, infrared thermal imager, lidar, GPS module and audio collector to ensure initial time synchronization at the hardware level (synchronization accuracy ≤10ms). Then collect road inspection related data according to the preset inspection route, covering five types of data: visible light image data, infrared thermal imaging data, lidar point cloud data, GPS trajectory data and audio data. Perform denoising, image enhancement, and size normalization operations on visible light image data and infrared thermal imaging data; perform distortion correction, ground point filtering, and feature extraction operations on lidar point cloud data. Perform coordinate calibration and time synchronization on GPS trajectory data; perform noise reduction and feature extraction on audio data; Based on the timestamp of GPS trajectory data, a cross-correlation algorithm is used to achieve time synchronization of visible light images, infrared thermal imaging and lidar point clouds. Based on the three-dimensional coordinates of lidar point clouds, spatial registration of image data and point cloud data is achieved through perspective projection transformation. The UTC timestamp sequence of GPS trajectory data is extracted as a time synchronization reference, and the accuracy target of time synchronization is determined (≤50ms). At the same time, the original timestamp sequences of visible light image, infrared thermal imaging and lidar data are extracted respectively. The cross-correlation algorithm is used to calculate the cross-correlation coefficient between each modal time sequence and the GPS reference time sequence. The time delay corresponding to the maximum value of the cross-correlation coefficient is found. The modal data to be aligned is shifted by this delay, and the time deviation of each modal data after synchronization is verified to ensure ≤50ms. Extract the UTM 3D coordinates (X,Y,Z) of the LiDAR point cloud as the spatial registration reference. Read the camera intrinsic and extrinsic parameter matrices (obtained through pre-calibration). Then, based on the perspective projection transformation formula, map the 3D coordinates of the LiDAR point cloud to the image pixel coordinates. Verify the matching degree between the mapped pixel coordinates and the target in the actual image. Adjust the extrinsic parameter matrix to optimize the registration accuracy. Output the spatially registered multimodal data to ensure that the registration accuracy is ≤10cm. The multimodal data integration library stores preprocessed multimodal data according to data type, collection time, and road segment. The data association index is built according to time, location, and modality type to realize fast association query of multimodal data.

[0019] S2. The integrated multimodal data is associated and matched according to the road event type, the effective collection time of each modality is counted, the inspection area is divided in combination with the road topology information, and the reference data node of each inspection area is selected. In S2, visible light image data and infrared thermal imaging data are associated with the dimensions of road congestion, traffic accidents, and obstacle-occupying events, while lidar point cloud data are associated with the dimensions of road surface damage and road obstacle size measurement. Break down the recognition requirements for each event: road congestion / traffic accidents / obstacles blocking the road require visual feature support; The dimensions of road surface damage / obstacles need to be supported by three-dimensional spatial dimensional features; Vehicle driving status (assisted congestion identification) requires the support of time-series trajectory features; Abnormal sound events (aiding accident identification) require audio feature support; The GPS trajectory data is correlated with the vehicle speed and traffic density calculation dimensions; the audio data is correlated with the abnormal sound recognition dimensions such as collision sounds and warning sounds, and a mapping rule between data and road event types is established. Visible light image data → Road congestion, traffic accidents, and obstacle obstruction identification dimensions, extracting visual features such as vehicle outlines, traffic flow distribution, and obstacle shapes; Infrared thermal imaging data → Road congestion, traffic accidents, and obstacles blocking the road can be identified, supplementing the target outline and temperature difference features in nighttime / rainy / foggy scenes (such as the heated parts of accident vehicles). LiDAR point cloud data → Dimensional measurement of road surface damage and road obstacles, extracting three-dimensional dimensional features such as damage area / depth, obstacle length, width, height / distance; GPS trajectory data → Calculate vehicle speed and traffic density, and extract time-series features such as vehicle location time-series changes and the number of vehicles passing by per unit time; Audio data → Anomalous sound recognition dimensions such as collision sounds and warning sounds, extracting sound spectrum and time domain features, and matching anomalous sound templates; The effective acquisition time is the time during which data for each modality is continuously and completely acquired and the data quality meets the recognition requirements. Statistics are completed by checking the data integrity and clarity frame by frame. Visible light / infrared images: no missing frames, image signal-to-noise ratio ≥30dB, blur ≤0.2; LiDAR point cloud: No missing points, point cloud density ≥ 100 points / m 2 After distortion correction, the error is ≤3cm; GPS track: no coordinate jumps, sampling frequency ≥1Hz, time synchronization error ≤50ms; Audio data: no broken frames, signal-to-noise ratio ≥25dB, sampling rate ≥44.1kHz; Based on road topology information and multimodal data coverage, set regional division thresholds; For highways, the length of a single inspection area shall be ≤10km; For main urban roads, the length of a single inspection area shall be ≤5km; Secondary arterial roads in the city, with a single inspection area length of ≤2km Areas with continuous road segments and continuous data coverage are divided into the same inspection area to ensure that there are no blind spots in the multimodal data coverage within a single inspection area; Based on the road station number, candidate areas are initially divided according to the division threshold. The multimodal data coverage of the candidate areas is verified. If there is a data blind spot (no modal data coverage) in a candidate area, the area is split / merged into areas without blind spots to ensure that areas with continuous road segments and continuous data coverage are classified as the same inspection area. Selecting a baseline data node means choosing the multimodal data set with the longest effective acquisition time, the highest data integrity, and the best spatiotemporal alignment accuracy as the baseline data node for each inspection area.

[0020] S3. Integrate all related data to perform accuracy verification on the modal data corresponding to the baseline data node. Based on the verification results, retain the data that meets the accuracy standard and substitute them into the corresponding inspection area and baseline data node. In S3, all multimodal data matched to each road event type in S2 are summarized and grouped and classified according to the inspection area and the reference data node; Acquire industry standards, traffic management regulations, and historical valid inspection data for road event recognition. Set accuracy standard thresholds for each modality of data: event feature recognition error ≤5% for visible light image data, temperature recognition error ≤2℃ for infrared thermal imaging data, and size measurement error ≤3% for lidar point cloud data.

[0021] We retrieved and sorted out industry standards for road incident recognition (such as the "Highway Technical Condition Assessment Standard" and the "Road Traffic Accident Scene Investigation Specification"), and also retrieved the accuracy requirements for inspection data issued by local traffic management departments. We also statistically analyzed the accuracy distribution of historical valid inspection data (such as the average value of feature recognition error, temperature measurement error, and size measurement error of qualified data in the past year). Based on industry standards, management regulations, and historical data, set accuracy thresholds (best practice values): For visible light image data, the event feature recognition error is ≤5% (such as the recognition deviation of vehicle / obstacle outlines and positions). Infrared thermal imaging data, temperature identification error ≤2℃ (such as measurement deviation of heated parts of accident vehicles and abnormal temperature areas on the road surface). The size measurement error of lidar point cloud data is ≤3% (such as the measurement deviation of road surface damage area / depth, and the length, width and height of obstacles).

[0022] S4. In each inspection area, road event identification nodes are generated based on the benchmark data nodes' compliance data and multimodal data features. At the same time, event feature data of each identification node is predicted based on the benchmark data nodes' compliance data. In S4, within each inspection area, the benchmark data of the benchmark data node is used as the benchmark origin. Based on the differences in road event type characteristics, road event identification nodes are generated in a uniformly distributed manner according to the principle of equal intervals, so that the generated identification nodes cover all road sections and key locations within the inspection area. Read the benchmark data nodes of each inspection area, extract their UTM plane coordinates as the benchmark origin of the area, simultaneously retrieve the road topology information of the inspection area, determine the area boundary coordinates and the coordinates of key locations (intersections, sharp bends, ramps), and clarify the generation range of the identification nodes (covering all road sections + key locations within the area). Starting from the reference origin (X0, Y0), a sequence of coordinate points with equal intervals is generated along the road direction (X-axis) according to an interval threshold. For each direction coordinate point, cross-sectional coordinate points are generated perpendicular to the road direction (Y-axis) according to cross-sectional intervals to form a grid-like node matrix. Then, the coordinates of key locations are added to the node matrix, invalid nodes that exceed the boundary of the inspection area are removed, and a unique number (area number + node sequence number) is assigned to each valid node. The "List of Coordinates of Inspection Area and Identified Nodes" is output, marking the node location and whether it is a key location. A Transformer+CNN hybrid architecture model is constructed. Local detail features are extracted from the feature maps of visible light images and infrared thermal imaging through CNN, and global correlation features are extracted from the 3D features of LiDAR point clouds and the temporal features of GPS trajectories through Transformer. After fusion, multimodal joint features are obtained. The input layer receives multimodal data features (visible light / infrared image feature maps, lidar point cloud 3D features, GPS trajectory time series features); CNN Feature Extraction Layer (Local Details): Input: Normalized feature maps of visible light images and infrared thermal images (1920×1080×3). The structure uses 2 layers of convolution + pooling (Conv2d + MaxPool2d), with a convolution kernel size of 3×3, a stride of 1, and padding of 1. Output: Image local detail feature vector (dimensionality: 256) Transformer Feature Extraction Layer (Global Association): Inputs: 3D features of LiDAR point cloud (128-dimensional FPFH features) and time-series features of GPS trajectory (64-dimensional time-series features); The structure uses an Encoder layer (multi-head attention head count = 8, hidden layer dimension = 512). Output: Spatial-temporal global correlation feature vector (dimension: 512). The feature fusion layer concatenates the local features output by the CNN with the global features output by the Transformer, and maps them to a unified dimension (1024 dimensions) through a fully connected layer (FC) to obtain multimodal joint features; The output layer outputs the event feature prediction results (such as obstacle size, congestion level, and damaged area) through the Softmax layer. For the image feature maps input to the CNN: normalize to the [0,1] interval, and adjust the dimensions to the model fit size; For the point cloud / trajectory features input to the Transformer: standardize (mean 0, variance 1), sort by time step / spatial points, and construct sequence features; Set a uniform batch size (BatchSize=32) and define the dimension validation rules for the input data.

[0023] The model inputs are the benchmark data nodes' compliance data and road event type features, and the historically labeled road event feature data are used as training labels. After the model training is completed, the multimodal basic data of the identification nodes are input, and the predicted event feature data corresponding to each identification node is output.

[0024] S5. Based on the compliance data retained in S3, perform same-node verification on the identification node data predicted in S4 in each inspection area, and adjust the prediction method according to the verification results until all identification node data meet the compliance data standards. In step S5, the compliance data retained in step S3 is matched to the identification nodes of the corresponding inspection area according to road location and event type, and the actual event feature data corresponding to each identification node is obtained. Calculate the Euclidean distance between the coordinates of the qualified data and the coordinates of the identification node, and set a distance threshold of ≤5m (to ensure accurate location matching). Then, for each identification node, select qualified data that meets the distance threshold and is consistent with the event type, and use it as the actual event feature data for that node. If a node has no matching qualified data, mark it as "data missing" and supplement it by interpolation from the qualified data of adjacent nodes (only as a temporary reference, and subsequent supplementary collection is required). Then generate the "Identification Node and Actual Event Feature Data Matching Table", which marks the node number, the matching qualified data, the actual feature value, and the data source. Check whether the accuracy of the matching data meets the threshold set by S3 (e.g., visible light feature recognition error ≤ 5%), remove data with matching errors (e.g., event type mismatch) or insufficient accuracy, and re-supplement the matching until each node has valid actual data; The predicted event feature data of the identification node is compared with the actual event feature data to determine the deviation. A qualified threshold is set. The relative deviation of numerical features (size / area / temperature) is ≤5% or the absolute deviation is ≤ the threshold (e.g., absolute temperature deviation is ≤2℃). The recognition accuracy of categorical features (event type) is ≥95%. If the deviation value is lower than the qualified threshold, the identification node is considered to have been successfully verified; if the deviation value is higher than the qualified threshold, the identification node is considered to have failed to be verified. For identification nodes that fail verification, the deviation features between the predicted data and the actual event feature data are extracted, input into the multimodal fusion prediction model, the feature weights and training parameters of the model are re-optimized, and the predicted data for the identification node is generated again and verified. Repeat the prediction, verification, and optimization process until the prediction data of all identification nodes in the inspection area meets the compliant data standards. Then, synchronize the optimized model and qualified prediction data to the corresponding database.

[0025] S6. The newly collected multimodal road inspection data is matched with the corresponding inspection area and identification node to carry out event identification. The road event type and level are determined by combining the multimodal feature fusion score and event confidence. The event location, handling suggestions and early warning information are output simultaneously.

[0026] In step S6, the acquisition location, time features and modality type of the newly acquired multimodal data are extracted, the similarity between the new data and the reference data nodes of each inspection area is calculated, the inspection area with the highest similarity is selected, and then the identification node with the closest position in that area is matched. The similarity calculation dimensions include location distance (Euclidean distance between the new data coordinates and the coordinates of the baseline nodes in each inspection area), time similarity (the overlap between the new data collection period and the baseline node data collection period), and modal integrity (the number of matches between the new data modal type and the baseline node modal type / the total number of modalities). A comprehensive similarity is calculated with weights (location weight 0.6, time weight 0.2, modal weight 0.2). Then, the baseline data nodes of all inspection areas are traversed, and the comprehensive similarity between the new data and each node is calculated. After sorting, the inspection area with the highest similarity is selected, as shown in the following formula: ; in, The overall similarity between the new data and the baseline data nodes. For positional similarity, For time similarity, a value of 1 is assigned to overlapping peak periods, and a value of 0.5 is assigned to non-overlapping periods. For modal integrity similarity; Calculate the Euclidean distance between the new data acquisition coordinates and the coordinates of all identification nodes within the area, and select the identification node with the smallest distance as the matching node, as shown in the following formula: ; in, The Euclidean distance between the new data and the baseline / identification node. , For the UTM plane coordinates of the newly acquired data, , The UTM plane coordinates of the reference data node / identification node; A weighted fusion strategy is used to calculate the multimodal feature fusion score: weights are assigned based on the contribution of each modality to the recognition of different road events, and the event recognition scores of each modality are weighted and summed to obtain the multimodal feature fusion score, as shown in the following formula: ; in, For multimodal feature fusion scoring, Let i be the weight of the i-th mode. Let n be the event recognition score for the i-th modality, and n be the number of valid modalities. If data for a certain modality is missing, its weight is proportionally distributed to the remaining valid modalities (e.g., in traffic accident recognition where audio is missing, a weight of 0.2 is evenly distributed among visible light / infrared / LiDAR), as shown in the following formula: ; in, The weights of the j-th effective mode after redistribution. The original weights for the j-th effective mode are... The original weights for the missing modes. This is the sum of the original weights of all valid modes; Based on the consistency of multimodal features and the model prediction probability, event confidence is generated. The consistency of the identification results of the effective modalities is calculated as follows: if ≥80% of the modalities are identified as the same event, the consistency level = 1; 60-79% = 0.8; <60% = 0.5; if only a single modality is effective, the consistency level = 0.7 (the default is medium consistency). Retrieve the predicted probability (range 0-1) of the event type output by the model for the matching node, take the maximum value as the predicted probability, and then calculate the time confidence score, as shown in the following formula: ; in, For the confidence level of the event, To assess the consistency of multimodal features, Predict the probability of the event type output by the model; Event grading standards are set by combining multimodal feature fusion scores and event confidence, resulting in Level 1, Level 2, and Level 3 events. Level 1 events with a fusion score ≥ 85 and a confidence level ≥ 0.9 include major traffic accidents and widespread traffic congestion. The secondary event fusion score was 70-84, with a confidence level of 0.8-0.89, indicating a general accident or a large obstacle blocking the road. Level 3 event fusion score 60-69, confidence level 0.7-0.79, minor damage, small obstacles blocking the way; The system pinpoints the exact location of the incident, generates handling suggestions based on the incident type and severity, and outputs early warning information to the traffic management platform via a wireless communication module.

[0027] Based on the three-dimensional coordinates of the LiDAR point cloud, combined with the perspective projection transformation results, the coordinates of the new data acquisition are corrected, and then the precise location of the event (UTM coordinates + road station number) is output with a positioning error of ≤1m. The electronic map is linked to mark the road name and lane information of the event location, and then the warning information is output to the traffic management platform.

[0028] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A low-altitude intelligent inspection method based on multimodal data fusion, characterized in that: Includes the following steps: S1. Acquire multimodal data related to road inspection, preprocess and align the data spatiotemporally, and establish a multimodal data integration library and data association index based on the processing results; S2. The integrated multimodal data is associated and matched according to the road event type, the effective collection time of each modality is counted, the inspection area is divided in combination with the road topology information, and the reference data node of each inspection area is selected. S3. Integrate all related data to perform accuracy verification on the modal data corresponding to the baseline data node. Based on the verification results, retain the data that meets the accuracy standard and substitute them into the corresponding inspection area and baseline data node. S4. In each inspection area, road event identification nodes are generated based on the benchmark data nodes' compliance data and multimodal data features. At the same time, event feature data of each identification node is predicted based on the benchmark data nodes' compliance data. S5. Based on the compliance data retained in S3, perform same-node verification on the identification node data predicted in S4 in each inspection area, and adjust the prediction method according to the verification results until all identification node data meet the compliance data standards. S6. The newly collected multimodal road inspection data is matched with the corresponding inspection area and identification node to carry out event identification. The road event type and level are determined by combining the multimodal feature fusion score and event confidence. The event location, handling suggestions and early warning information are output simultaneously.

2. The low-altitude intelligent inspection method based on multimodal data fusion according to claim 1, characterized in that: In S1, a low-altitude inspection platform equipped with multimodal sensors synchronously collects multimodal data related to road inspection. The multimodal data includes visible light image data, infrared thermal imaging data, lidar point cloud data, GPS trajectory data, and audio data. Perform denoising, image enhancement, and size normalization operations on visible light image data and infrared thermal imaging data; perform distortion correction, ground point filtering, and feature extraction operations on lidar point cloud data. Perform coordinate calibration and time synchronization on GPS trajectory data; perform noise reduction and feature extraction on audio data; Based on the timestamp of GPS trajectory data, a cross-correlation algorithm is used to achieve time synchronization of visible light images, infrared thermal imaging and lidar point clouds. Based on the three-dimensional coordinates of lidar point clouds, spatial registration of image data and point cloud data is achieved through perspective projection transformation. The multimodal data integration library stores preprocessed multimodal data according to data type, collection time, and road segment. The data association index is built according to time, location, and modality type to realize fast association query of multimodal data.

3. The low-altitude intelligent inspection method based on multimodal data fusion according to claim 1, characterized in that: In S2, visible light image data and infrared thermal imaging data are associated with the dimensions of road congestion, traffic accidents, and obstacle-occupying events, while lidar point cloud data are associated with the dimensions of road surface damage and road obstacle size measurement. The GPS trajectory data is correlated with the vehicle speed and traffic density calculation dimensions; the audio data is correlated with the abnormal sound recognition dimensions such as collision sounds and warning sounds, and a mapping rule between data and road event types is established. The effective acquisition time is the time during which data for each modality is continuously and completely acquired and the data quality meets the recognition requirements. Statistics are completed by checking the data integrity and clarity frame by frame. Based on road topology information and multimodal data coverage, a regional division threshold is set to divide areas with continuous road segments and continuous data coverage into the same inspection area, ensuring that there are no blind spots in multimodal data coverage within a single inspection area. Selecting a baseline data node means choosing the multimodal data set with the longest effective acquisition time, the highest data integrity, and the best spatiotemporal alignment accuracy as the baseline data node for each inspection area.

4. The low-altitude intelligent inspection method based on multimodal data fusion according to claim 1, characterized in that: In S3, all multimodal data matched to each road event type in S2 are summarized and grouped and classified according to the inspection area and the reference data node; Acquire industry standards, traffic management regulations, and historical valid inspection data for road event recognition. Set accuracy standard thresholds for each modality of data: event feature recognition error ≤5% for visible light image data, temperature recognition error ≤2℃ for infrared thermal imaging data, and size measurement error ≤3% for lidar point cloud data.

5. The low-altitude intelligent inspection method based on multimodal data fusion according to claim 1, characterized in that: In step S3, the modal data of each reference data node are compared with the corresponding accuracy standard threshold, and the data deviation rate is calculated. If the deviation rate is lower than the accuracy standard threshold, the accuracy is determined to meet the standard; otherwise, the accuracy is determined to be lower than the standard. If the multimodal data of the reference data nodes in the inspection area does not meet the standard, then a new multimodal data group with the highest accuracy and longest effective acquisition time is selected in the inspection area and substituted into the corresponding inspection area as a new reference data node.

6. The low-altitude intelligent inspection method based on multimodal data fusion according to claim 1, characterized in that: In S4, within each inspection area, the benchmark data of the benchmark data node is used as the benchmark origin. Based on the differences in road event type characteristics, road event identification nodes are generated in a uniformly distributed manner according to the principle of equal intervals, so that the generated identification nodes cover all road sections and key locations within the inspection area. A Transformer+CNN hybrid architecture model is constructed. Local detail features are extracted from the feature maps of visible light images and infrared thermal imaging through CNN, and global correlation features are extracted from the 3D features of LiDAR point clouds and the temporal features of GPS trajectories through Transformer. After fusion, multimodal joint features are obtained. The model inputs are the benchmark data nodes' compliance data and road event type features, and the historically labeled road event feature data are used as training labels. After the model training is completed, the multimodal basic data of the identification nodes are input, and the predicted event feature data corresponding to each identification node is output.

7. The low-altitude intelligent inspection method based on multimodal data fusion according to claim 1, characterized in that: In step S5, the compliance data retained in step S3 is matched to the identification nodes of the corresponding inspection area according to road location and event type, and the actual event feature data corresponding to each identification node is obtained. The predicted event feature data of the identification node is compared with the actual event feature data to determine the deviation. A verification threshold is set. If the deviation value is lower than the verification threshold, the identification node is determined to be verified successfully. If the deviation value is higher than the verification threshold, the identification node is determined to be verified unsuccessfully. For identification nodes that fail verification, the deviation features between the predicted data and the actual event feature data are extracted, input into the multimodal fusion prediction model, the feature weights and training parameters of the model are re-optimized, and the predicted data for the identification node is generated again and verified. Repeat the prediction, verification, and optimization process until the prediction data of all identification nodes in the inspection area meets the compliant data standards. Then, synchronize the optimized model and qualified prediction data to the corresponding database.

8. The low-altitude intelligent inspection method based on multimodal data fusion according to claim 1, characterized in that: In step S6, the acquisition location, time features and modality type of the newly acquired multimodal data are extracted, the similarity between the new data and the reference data nodes of each inspection area is calculated, the inspection area with the highest similarity is selected, and then the identification node with the closest position in that area is matched. A weighted fusion strategy is used to calculate the multimodal feature fusion score: weights are set according to the contribution of each modality data to the recognition of different road events, and the event recognition scores of each modality are weighted and summed to obtain the multimodal feature fusion score; Based on the consistency of multimodal features and the model prediction probability, event confidence is generated. Event grading standards are set by combining multimodal feature fusion scores and event confidence, resulting in Level 1, Level 2, and Level 3 events. The system pinpoints the exact location of the incident, generates handling suggestions based on the incident type and severity, and outputs early warning information to the traffic management platform via a wireless communication module.