Highway holographic intelligent detection method and system based on multi-modal sensing and edge AI

By leveraging multimodal sensing and edge AI technologies, holographic intelligent detection of highways is achieved, solving the problems of the single nature and real-time performance of traditional detection technologies. This improves detection depth and management efficiency, enabling accurate diagnosis of the causes of road damage and real-time early warning, thus ensuring road safety.

CN121884577APending Publication Date: 2026-04-17浪潮智慧科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
浪潮智慧科技有限公司
Filing Date
2025-11-28
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional highway detection technologies are singular and isolated, lacking depth and real-time performance. Detection results fail to be efficiently integrated with management systems, leading to road safety hazards and low management efficiency.

Method used

Employing multimodal sensing and edge AI technologies, data is synchronously collected through a multimodal sensing system mounted on a mobile platform, processed in real time using edge computing, and analyzed through cross-modal fusion diagnostic AI models to generate a digital twin for visualization and early warning, achieving comprehensive and integrated intelligent detection.

Benefits of technology

It enables holographic rapid diagnosis and early warning of highways, improving the depth, efficiency and intelligence of detection, accurately locating the causes of defects, shortening emergency response time, and ensuring road safety and management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884577A_ABST
    Figure CN121884577A_ABST
Patent Text Reader

Abstract

The invention relates to the field of expressway intelligent detection, in particular to an expressway holographic intelligent detection method and system based on multi-modal sensing and edge AI, and the method comprises the steps: synchronously collecting the multi-modal data of the surface layer and the internal structure layer of a road through a multi-modal sensing system carried on a mobile platform; performing real-time processing on the multi-modal data by using an edge calculation unit, and generating a time-space aligned standardized data set through time calibration and space alignment; performing fusion analysis on the standardized data set by adopting a cross-modal fusion diagnosis AI model to realize associated diagnosis and cause inference of road diseases; and mapping the diagnosis analysis result to a three-dimensional road model to generate a digital twinborn body, performing visual display, triggering early warning based on the diagnosis result, and performing linkage with roadside facilities to form closed-loop management and control. Single and isolated detection is spanned to omnibearing and integrated intelligent detection, and the detection depth, efficiency and intelligent level are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent highway detection technology, specifically to a holographic intelligent detection method and system for highways based on multimodal sensing and edge AI. Background Technology

[0002] With the continuous improvement of the expressway network and the sustained growth of traffic load, efficient and precise maintenance of highway assets has become crucial to ensuring road safety and extending service life. However, traditional expressway inspection and management technologies and systems have revealed many limitations in practical applications that urgently need to be addressed.

[0003] First, traditional testing methods are singular and isolated. They are usually conducted in stages using methods such as flatness gauges, deflectometers, or manual visual inspection, resulting in fragmented data between different testing items. For example, when a road surface crack is discovered, it is impossible to simultaneously obtain data on the subgrade settlement or internal structural condition at that location, making it impossible to determine whether the crack is caused by base layer delamination, material aging, or overload.

[0004] Secondly, the depth of detection is severely insufficient. There is a heavy reliance on surface-level inspection technologies such as optical imaging, lacking effective non-destructive detection capabilities for hidden defects such as voids and loosening in the structural layers below the road surface (e.g., base course and subbase). These "invisible" hazards are the main causes of road structural damage, and their delayed detection poses a significant risk to road safety.

[0005] Furthermore, the timeline from detection to decision-making is too long, resulting in poor real-time performance. After data collection, it needs to be sent back to the laboratory for manual processing and analysis for several days or even weeks, leading to a significant delay in output. For sudden and severe damage such as potholes after heavy rain or frost heave in winter, rapid early warning and emergency response are impossible.

[0006] Finally, the test results are mostly presented in the form of reports, which fail to effectively link with the digital model of road assets, early warning issuance, and maintenance and disposal actions, resulting in low management efficiency.

[0007] Therefore, there is an urgent need in this field for an integrated intelligent detection technology that can achieve all-round perception, real-time analysis, accurate diagnosis and closed-loop management, so as to promote the digital and intelligent transformation and upgrading of highway maintenance management. Summary of the Invention

[0008] To address the aforementioned issues, this invention provides a holographic intelligent detection method and system for highways based on multimodal sensing and edge AI. The aim is to achieve rapid holographic diagnosis and early warning of highway health conditions, improve detection efficiency and accuracy, reduce maintenance costs, and promote the digital transformation of maintenance management.

[0009] In a first aspect, the present invention provides a holographic intelligent detection method for highways based on multimodal sensing and edge AI, comprising the following steps: Multimodal sensing system mounted on a mobile platform is used to simultaneously collect multimodal data of the road surface and internal structural layers; The multimodal data is processed in real time using edge computing units, and a standardized dataset with spatiotemporal alignment is generated through time calibration and spatial alignment. A cross-modal fusion diagnostic AI model is used to perform fusion analysis on the standardized dataset to achieve correlation diagnosis and cause inference of road defects; The diagnostic analysis results are mapped to a 3D road model to generate a digital twin and visualize it. At the same time, early warnings are triggered based on the diagnostic results, and a closed-loop management system is formed in conjunction with roadside facilities.

[0010] This represents a leap from traditional, isolated detection to comprehensive, integrated intelligent detection, enhancing the depth, efficiency, and intelligence of the detection process. Through edge processing and correlation diagnosis, it enables rapid and accurate location and causal analysis of diseases, improving the scientific nature and effectiveness of maintenance work.

[0011] As a preferred embodiment of the technical solution of the present invention, the multimodal sensing system includes a visual sensor, a three-dimensional lidar, a radar sensor, an infrared thermal imager, and a high-precision positioning and inertial navigation system. The steps of multimodal data acquisition include: Road surface images are acquired by a high-resolution line scan camera and a panoramic camera mounted on a mobile platform; wherein, the line scan camera continuously takes pictures at a first frequency to obtain road surface detail images while the mobile platform is moving, and the panoramic camera periodically takes pictures at a second frequency to obtain panoramic road view images. A high-precision point cloud model of the road surface is constructed by emitting laser beams and receiving reflected signals using a three-dimensional lidar with a scanning frequency within a set range. This model is used to calculate rut depth, smoothness, and road geometry parameters. The system uses millimeter-wave radar to monitor traffic flow around vehicles in real time, and uses ground-penetrating radar to emit high-frequency electromagnetic waves that penetrate the road surface structure layer and receive reflected signals to analyze the thickness, voids, and looseness of the structure layer within a set depth range below the road surface. Infrared images of the road surface are acquired by an infrared thermal imager, and the differences in the road surface temperature field are analyzed to identify areas of voids or water seepage under the asphalt pavement layer. High-precision positioning and inertial navigation systems provide spatiotemporal references for all collected data, enabling geocoding and preliminary time stamping of the data.

[0012] Through the collaborative work of multiple sensors, the system simultaneously acquires information on road surface conditions, internal structure, traffic flow environment, and precise geographical location, forming a complete data loop and laying a solid data foundation for holographic perception.

[0013] As a preferred embodiment of the technical solution of the present invention, the steps of using an edge computing unit to process the multimodal data in real time and generating a spatiotemporally aligned standardized dataset through time calibration and spatial alignment include: The precise time signal generated by the high-precision positioning and inertial navigation system is used as a unified hardware trigger signal and sent to each sensor to trigger each sensor to synchronously start or mark the data acquisition process. In the edge computing unit, a time correction is applied to the collected data based on the acquisition and transmission delay of each sensor, so that all data are aligned on a unified time axis. By utilizing the high-precision positioning and attitude data generated by the high-precision positioning and inertial navigation system, all sensor data are unified under the same global geographic coordinate system, thereby generating a standardized dataset with spatiotemporal alignment.

[0014] By combining hardware triggering and software calibration, the inconsistencies in time and space between multi-source heterogeneous sensor data were resolved, providing a reliable guarantee for subsequent high-precision data fusion and correlation analysis, which is a prerequisite for achieving correlation diagnosis. This ensures accurate registration between different sensor data, enabling advanced analyses such as using LiDAR data to verify visually identified cracks and using ground-penetrating radar data to investigate the condition of the base layer beneath cracks.

[0015] As a preferred embodiment of the technical solution of the present invention, the time correction amount is dynamically calculated based on the time difference between the arrival of the hardware trigger signal and the sensor data packet at the edge computing unit.

[0016] As a preferred embodiment of the technical solution of the present invention, the steps of using a cross-modal fusion diagnostic AI model to perform fusion analysis on the standardized dataset to achieve correlation diagnosis and cause inference of road defects include: The spatiotemporally aligned standardized dataset is input into a multi-branch neural network for feature extraction; wherein, visual image data is input into a fully convolutional neural network to extract pixel-level feature maps, and 3D point cloud and non-image data are input into a graph neural network to construct a graph structure with spatial location as nodes and extract relational features; Features from different neural network branches are concatenated or weighted and fused at the feature layer, or the preliminary diagnostic results of each branch are comprehensively reasoned at the decision layer. When the visual branch identifies road surface cracks, the model automatically associates and calls up lidar point cloud data at the same location based on the spatial coordinates of the cracks to determine whether there is subsidence, calls ground-penetrating radar data to assess the health status of the base layer, and generates a comprehensive diagnostic report that includes the type, location, severity, and cause inference of the disease.

[0017] By fusing feature-level and decision-level data, the complementary advantages of different modalities are fully utilized, improving the overall accuracy and reliability of diagnosis and effectively avoiding false alarms and false negatives.

[0018] As a preferred embodiment of the technical solution of the present invention, the associated diagnostic step specifically includes: The pixel coordinates of the road surface cracks identified by the visual branch are combined with the corresponding high-precision positioning data in the spatiotemporally aligned standardized dataset and converted into unified global geographic coordinates. Based on the global geographic coordinates, a spatial range query is performed in the standardized dataset to automatically retrieve and call the data associated with that coordinate location: Three-dimensional lidar point cloud data is used to calculate the road elevation change at that location in order to determine whether there is subsidence; Ground-penetrating radar reflected wave data is used to analyze the electromagnetic wave characteristics and phase axis continuity of the road structure layer below the location to assess whether there is voiding, loosening or abnormal moisture content in the base layer. By integrating crack characteristics, settlement assessment results, and basic health status evaluation results, and based on a pre-defined disease causal knowledge graph, a comprehensive diagnostic report is generated. The report includes at least the disease type, precise geographical location, severity level, and inferences about the potential causes of the crack.

[0019] As a preferred embodiment of the technical solution of this invention, the steps of mapping the diagnostic analysis results to a three-dimensional road model to generate a digital twin and displaying it visually, triggering an early warning based on the diagnostic results, and linking with roadside facilities to form a closed-loop management system include: The comprehensive diagnostic report is used as attribute information and mapped and bound to the corresponding geometric position in the 3D road model to dynamically update the state of the road digital twin; On the visualization terminal, the disease information in the digital twin is rendered with different colors, icons or transparency, and users can interact with the disease tags to view detailed diagnostic reports. When the severity of the disease in the comprehensive diagnostic report exceeds a preset threshold, or when the disease type belongs to the sudden high-risk category, a real-time warning is automatically triggered. Based on the triggered warning information, specific equipment control commands are generated and sent to the corresponding roadside facilities through the communication network; The roadside facilities include at least variable message signs, broadcasting systems, lighting systems, and signal control systems, which change their operating status through the instructions to actively guide and warn traffic flow.

[0020] Transforming abstract detection data into intuitive and visual three-dimensional dynamic models greatly enhances managers' situational awareness and decision-making efficiency.

[0021] It has achieved closed-loop management from passive response to proactive early warning and then to coordinated handling, shortening the emergency response time for sudden high-risk diseases, effectively preventing secondary traffic accidents, and ensuring road traffic safety.

[0022] As a preferred embodiment of the technical solution of the present invention, the method further includes: Record the trigger time of the warning information, the content of the handling instructions, and the response status of the roadside facilities; During subsequent inspections, when the mobile platform passes by the reported location of the defect again, it automatically detects whether the defect has been repaired and updates the repair results to the digital twin and management file, thus completing closed-loop management.

[0023] As a preferred embodiment of the technical solution of the present invention, the training steps of the cross-modal fusion diagnostic AI model include: The historical dataset synchronously collected by the multimodal sensing system is acquired, and the historical dataset is spatiotemporally aligned to generate a spatiotemporally aligned training sample set. Each training sample contains multimodal sensor data of the same road location point, and corresponding manually labeled data; the labels include at least: disease type, disease location, severity, and cause of disease. Construct a multi-branch deep learning network model that includes at least: A vision branch based on a fully convolutional neural network, used to process image data; A non-visual branch based on graph neural networks is used to process point cloud and radar data; A feature fusion module is used to fuse features from different branches; Using a spatiotemporally aligned training sample set as input and the corresponding labels as supervision signals, the multi-branch deep learning network model is trained end-to-end. Among them, the loss function of the visual branch is constructed based on pixel-level classification accuracy, the loss function of the non-visual branch is constructed based on graph node classification or regression error, and the overall loss function of the feature fusion module is a weighted sum of the loss functions of each branch, with the addition of a regularization term to promote cross-modal association learning. The performance of the trained model is evaluated using a reserved validation set. When the model’s comprehensive diagnostic accuracy for disease type, location and cause exceeds a preset threshold, a trained cross-modal fusion diagnostic AI model is obtained, and the model parameters are deployed to the edge computing unit.

[0024] Secondly, the present invention also provides a highway holographic intelligent detection system based on multimodal sensing and edge AI, used to implement the method described in the first aspect, the system comprising: The perception layer includes a multimodal sensing system mounted on a mobile platform, used to simultaneously collect multimodal data from the road surface and internal structural layers; The processing layer, including edge computing units, is configured as follows: The multimodal data is processed in real time, and a standardized dataset with spatiotemporal alignment is generated through time calibration and spatial alignment. A cross-modal fusion diagnostic AI model is used to perform fusion analysis on the standardized dataset to achieve correlation diagnosis and cause inference of road defects; The application layer is configured as follows: The diagnostic analysis results are mapped to a 3D road model to generate a digital twin and then visualized. Early warnings are triggered based on diagnostic results, and a closed-loop management system is formed in conjunction with roadside facilities.

[0025] As a preferred embodiment of the technical solution of the present invention, the multimodal sensing system of the sensing layer includes: High-resolution line scan cameras and panoramic cameras are used to acquire images of the road surface; 3D LiDAR is used to construct high-precision point cloud models of road surfaces at a set scanning frequency; Millimeter-wave radar and ground-penetrating radar are used to monitor traffic flow and detect defects in the structural layers below the road surface, respectively. Infrared thermal imagers are used to acquire infrared images of road surfaces to identify temperature field anomalies. A high-precision positioning and inertial navigation system is used to provide a spatiotemporal reference for all acquired data.

[0026] As a preferred embodiment of the technical solution of the present invention, the edge computing unit of the processing layer further includes: The multi-source data spatiotemporal synchronization module is configured to perform time calibration using hardware trigger signals from a high-precision positioning and inertial navigation system, and to perform spatial alignment using its positioning and attitude data to generate the standardized dataset. The cross-modal fusion diagnostic AI model is a multi-branch neural network structure, including: The visual branch based on fully convolutional neural networks is used to process image data; A non-visual branch of graph neural networks is used to process point cloud and radar data. The feature fusion module is used to fuse features from different branches and realize the correlation diagnosis of multimodal data based on unified geographic coordinates.

[0027] As a preferred embodiment of the technical solution of the present invention, the application layer further includes: A digital twin engine is used to bind diagnostic results as attribute information to the geometric position of a 3D model, enabling dynamic updates and visualization rendering; The early warning linkage controller is configured to automatically generate control commands and send them to at least one of the variable message signs, broadcasting systems, lighting systems, and signal control systems when the diagnostic results meet preset conditions.

[0028] As can be seen from the above technical solutions, this application has the following advantages: This invention breaks down the barriers of data fragmentation in traditional detection by using a cross-modal fusion diagnostic AI model. The system can not only identify individual defects but also reveal the inherent causal relationships between defects. For example, it can accurately determine whether surface cracks are caused by the underlying base layer delamination or are associated with material aging. This allows maintenance plans to be tailored to the root cause of defects, achieving a shift from treating symptoms to addressing the root cause, fundamentally preventing recurrence of defects, and extending the service life of roads.

[0029] By integrating data from sensors such as ground-penetrating radar and infrared thermal imagers, the limitations of traditional optical inspection have been overcome, enabling non-destructive perception of the condition and material properties of the subsurface structural layers. This allows for the early detection of major safety hazards, truly achieving preventative maintenance and ensuring the long-term safety of the road structure.

[0030] Leveraging edge computing capabilities, data analysis timelines have been reduced from days or even weeks to minutes or even seconds. For sudden and severe problems such as potholes and foundation settlement, near real-time early warnings can be achieved, significantly shortening emergency response time and providing crucial technical support for proactive traffic safety management. Attached Figure Description

[0031] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a flowchart illustrating the method provided in an embodiment of the present invention.

[0033] Figure 2 This is a diagram illustrating a fully convolutional neural network.

[0034] Figure 3 This is a diagram illustrating a graph neural network.

[0035] Figure 4 This is a system block diagram provided for an embodiment of the present invention. Detailed Implementation

[0036] To make the purpose, features, and advantages of this application more apparent and understandable, specific embodiments and accompanying drawings will be used to clearly and completely describe the technical solution protected by this application. Obviously, the embodiments described below are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0037] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this application and in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0038] like Figure 1 As shown, this embodiment of the invention provides a holographic intelligent detection method for highways based on multimodal sensing and edge AI, including the following steps: S1. Multimodal data of the road surface and internal structural layers are simultaneously collected through a multimodal sensing system mounted on a mobile platform; the multimodal sensing system includes a visual sensor, a three-dimensional lidar, a radar sensor, an infrared thermal imager, and a high-precision positioning and inertial navigation system. Visual sensors: High-resolution line-scan cameras and panoramic cameras are mounted on the top of the inspection vehicle or at appropriate locations on the drone. The line-scan camera continuously captures images of the road surface while the vehicle is in motion to obtain high-resolution images of road details. Its shooting frequency is dynamically adjusted according to the vehicle's speed and the required resolution. For example, when the vehicle is traveling at 80 km / h, the line-scan camera captures 50-100 images per second to ensure that even tiny cracks and damage are captured. The panoramic camera acquires road surface images with a wider field of view, assisting in the positioning and scene supplementation of the line-scan camera images. It can cover the road surface conditions within a 360° range centered on the inspection vehicle or drone, capturing a panoramic image every 3-5 seconds.

[0039] 3D LiDAR: A 3D LiDAR is mounted on the front of an inspection vehicle or the bottom of a drone. The LiDAR emits a laser beam and receives reflected signals to construct a point cloud model of the road. With a scanning frequency set between 10Hz and 20Hz, the LiDAR can accurately measure the distance information of various points on the road surface, thereby generating a high-precision 3D road model. Analysis of this model can calculate rut depth, inter-iron index (IRI), and road geometric parameters such as cross slope and longitudinal slope. For example, a road surface curve can be fitted using point cloud data, and the IRI can be calculated by comparing it with a standard curve.

[0040] Radar sensors: Millimeter-wave radar is installed at the front and rear of the inspection vehicle, as well as at the front and rear of the drone. Millimeter-wave radar monitors traffic flow around the vehicle in real time by transmitting and receiving millimeter-wave signals, including vehicle speed, distance, and acceleration. Its detection range is 200-300 meters in front and 100-200 meters behind. Ground-penetrating radar is installed on the bottom of the inspection vehicle near the road surface. It transmits high-frequency electromagnetic waves to penetrate the road surface structure and receives signals reflected from different media interfaces to analyze the thickness, voids, and looseness of the subsurface structure. The detection depth of ground-penetrating radar can be adjusted according to the road surface structure, generally detecting depths of 1-3 meters below the road surface.

[0041] Infrared thermal imager: An infrared thermal imager is mounted on the side of an inspection vehicle near the road surface or at a specific angle below a drone to detect the temperature field of the road surface. The imager acquires 10-20 frames of infrared images per second. By analyzing the temperature differences in different areas of the images, it identifies areas of voids and water seepage beneath the asphalt pavement. For example, due to air insulation, voided areas exhibit different temperature variations than normal road surfaces, appearing as specific temperature anomalies in the infrared thermal image.

[0042] High-precision positioning and inertial navigation: An IMU / GNSS integrated system is installed on inspection vehicles or drones. GNSS provides precise global positioning information, while the IMU measures the vehicle's or drone's acceleration and angular velocity in real time. By fusing the data from both, a precise spatiotemporal reference is provided for data collected by other sensors. The system's positioning accuracy can reach the centimeter level, ensuring the precise location and time stamp of the collected data in geospatial space, and achieving data synchronization and geocoding.

[0043] S2. Utilize edge computing units to process the multimodal data in real time, and generate a spatiotemporally aligned standardized dataset through time calibration and spatial alignment; this step specifically includes: The precise time signal generated by the high-precision positioning and inertial navigation system is used as a unified hardware trigger signal and sent to each sensor to trigger each sensor to synchronously start or mark the data acquisition process. In the edge computing unit, a time correction is applied to the collected data based on the acquisition and transmission delay of each sensor, so that all data are aligned on a unified time axis; the time correction is dynamically calculated based on the time difference between the hardware trigger signal and the arrival of the sensor data packet in the edge computing unit.

[0044] By utilizing the high-precision positioning and attitude data generated by the high-precision positioning and inertial navigation system, all sensor data are unified under the same global geographic coordinate system, thereby generating a standardized dataset with spatiotemporal alignment.

[0045] Edge computing units utilize embedded AI devices (such as the NVIDIA Jetson series) mounted on vehicles or roadside platforms to perform real-time computation locally, reducing data transmission latency and improving processing efficiency. Data collected by various sensors is transmitted to the edge computing unit in real time via high-speed data transmission lines, such as Gigabit Ethernet or high-speed USB interfaces. After receiving the data, the edge computing unit first performs preprocessing, including data format conversion and noise reduction, to ensure data quality and consistency. A multi-source data spatiotemporal synchronization module aligns data from different frequencies and sources within a unified spatiotemporal framework through hardware triggering and software algorithms. Taking data from two different sensors as an example, sensor A uses data from different frequencies... Data is collected by sensor B at a frequency Collect data. Assume... For data from sensor A, the time variable is used. and data from sensor B To achieve time and space synchronization, a hardware trigger signal is required. As a time reference, the data is time-calibrated in the software algorithm using the following formula: For data from sensor A, the calibrated time for:

[0046] in It is a time correction amount determined based on the hardware trigger signal and the characteristics of sensor A.

[0047] For data from sensor B, the calibrated time for:

[0048] in It is a time correction amount determined based on the hardware trigger signal and the characteristics of sensor B.

[0049] This calibration aligns data from different sensors on the same time scale, meeting the requirements of subsequent cross-modal fusion diagnostic AI models for spatiotemporal data consistency.

[0050] S3. Employ a cross-modal fusion diagnostic AI model to perform fusion analysis on the standardized dataset to achieve correlation diagnosis and cause inference of road defects; this step specifically includes: The spatiotemporally aligned standardized dataset is input into a multi-branch neural network for feature extraction; wherein, visual image data is input into a fully convolutional neural network to extract pixel-level feature maps, and 3D point cloud and non-image data are input into a graph neural network to construct a graph structure with spatial location as nodes and extract relational features; Features from different neural network branches are concatenated or weighted and fused at the feature layer, or the preliminary diagnostic results of each branch are comprehensively reasoned at the decision layer. When the visual branch identifies road surface cracks, the model automatically associates and calls up lidar point cloud data at the same location based on the spatial coordinates of the cracks to determine whether there is subsidence, calls ground-penetrating radar data to assess the health status of the base layer, and generates a comprehensive diagnostic report that includes the type, location, severity, and cause inference of the disease.

[0051] Deploy a deep learning-based cross-modal fusion diagnostic AI model in an edge computing unit. The model employs a fully convolutional neural network (FCN) (such as...). Figure 2 (as shown) and Graph Neural Networks (GNNs) (as shown) Figure 3 The architecture combines the two (as shown). Figure 2 The fully convolutional neural network in the image classification task consists of an input layer, a convolutional layer, a pooling layer, and a fully connected layer from left to right. The functions and structure of each part are as follows: The leftmost input layer is a 28×28 pixel input image, displayed in a grid format, containing the image information to be recognized.

[0052] The first convolutional layer performs convolution operations on the input image (using multiple convolutional kernels to extract local features), generating 6 feature maps (denoted as C1), each with a size of 28×28 (labeled "6@28×28"). By capturing local features of the image through convolutional kernels, the 6 kernels correspond to 6 different feature extraction methods, outputting 6 feature maps representing the responses of different features.

[0053] The first pooling layer S2 performs pooling operations (usually max pooling or average pooling) on ​​the feature map of C1. The purpose is to reduce dimensionality and computational cost, while preserving key features and enhancing translation invariance.

[0054] Output: Generate 6 feature maps, each with a size of 14×14 (labeled as "6@14×14" or "S2 feature map").

[0055] The second convolutional layer C3 performs convolution operations on the feature maps of S2 again, using more convolutional kernels (16) to extract more complex features (such as combined edges and shapes). This generates 16 feature maps (denoted as C3), each with a size of 10×10 (labeled as "16@10×10" or "C3 feature map").

[0056] The second aggregation layer S4 performs a second aggregation on the feature maps of C3, further reducing the dimensionality. This generates 16 feature maps, each with a size of 5×5 (labeled as "16@5×5" or "S4 feature map").

[0057] After two convolutions and two convergences, the features in the fully connected layer are flattened into a one-dimensional vector, which is then input into the fully connected layer for classification: The first fully connected layer F5 takes a 16×5×5=400-dimensional vector as input (all 5×5×16 feature maps are flattened) and outputs 120 neurons (labeled "120-F5"). The second fully connected layer F6 takes a 120-dimensional input and outputs 84 neurons (labeled "84-F6"). The output layer takes an 84-dimensional input and outputs 10 neurons (labeled "10-output"), corresponding to the number of categories in the classification task (e.g., the 10 digit categories in MNIST).

[0058] Figure 3 The graph neural network shown on the left is a complex network composed of nodes and edges. The one on the right is a typical multilayer feedforward neural network, which includes the following components: (1) The part marked "t" is the input layer, which is responsible for receiving external data. The number of nodes in the input layer is usually related to the dimension of the input data.

[0059] (2) Shows hidden layers such as "Hidden Layer 1", "Hidden Layer 2"... "Hidden Layer L-1", etc. These layers are the core computational layers of the neural network, responsible for performing nonlinear transformations and feature extraction on the input data. Each hidden layer contains multiple neurons, which are connected by edges (weights) to form a fully connected (or partially connected) structure. Neurons are the basic computational units of the neural network, responsible for performing weighted summation on the input and outputting the result after passing through activation functions (such as ReLU, Sigmoid, etc.).

[0060] (3) The rightmost layer is the output layer, which contains multiple neurons and is responsible for outputting the final result (e.g. Figure 3 The marked The number of neurons in the output layer is usually task-dependent. Classification tasks: It could be the number of categories (e.g., binary classification). =2, multiple categories ≥3).

[0061] Return mission: It is the dimension of the output (such as predicting multiple continuous values).

[0062] In this embodiment of the invention, a fully convolutional neural network (FCN) is used to extract features from image data acquired by a visual sensor, identifying features such as road surface cracks and damage. A graph neural network (GNN) is used to process non-image data acquired by 3D lidar, radar sensors, etc., constructing a graph structure from different sensor data to uncover spatial and logical relationships between the data. When the visual sensor identifies road surface cracks, the model uses spatial location information to call lidar data at that location to determine if subsidence exists, and ground-penetrating radar data to assess the health status of the subgrade, achieving feature-level and decision-level fusion of multimodal data, thereby comprehensively diagnosing road defects and their causes.

[0063] It should be noted here that the associated diagnostic steps specifically include: The pixel coordinates of the road surface cracks identified by the visual branch are combined with the corresponding high-precision positioning data in the spatiotemporally aligned standardized dataset and converted into unified global geographic coordinates. Based on the global geographic coordinates, a spatial range query is performed in the standardized dataset to automatically retrieve and call the data associated with that coordinate location: Three-dimensional lidar point cloud data is used to calculate the road elevation change at that location in order to determine whether there is subsidence; Ground-penetrating radar reflected wave data is used to analyze the electromagnetic wave characteristics and phase axis continuity of the road structure layer below the location to assess whether there is voiding, loosening or abnormal moisture content in the base layer. By integrating crack characteristics, settlement assessment results, and basic health status evaluation results, and based on a pre-defined disease causal knowledge graph, a comprehensive diagnostic report is generated. The report includes at least the disease type, precise geographical location, severity level, and inferences about the potential causes of the crack.

[0064] The disease causal knowledge graph pre-defines various disease patterns, including: Subbase void-settlement-surface crack pattern: When there are abnormal subbase void assessment results, positive settlement judgment results and linear cracks at the same time, it is inferred that the cracks are mainly caused by the void of the underlying subbase. Material aging-rutting-network cracking pattern: When there are no obvious base course abnormalities and settlement, but there are severe rutting and network cracks, it is inferred that the cracks are mainly caused by the aging of asphalt materials.

[0065] The standardized dataset includes the following dimensional information: Visual data providing apparent evidence of defects (provided by linear / panoramic cameras), three-dimensional geometric data providing precise quantitative information on road geometry (provided by three-dimensional lidar), underground structural data providing perspective information on the health status of structures below the road surface (provided by ground-penetrating radar), material property data providing indirect evidence of material thermophysical properties related to internal defects (provided by infrared thermal imagers), and precise spatiotemporal reference data (provided by an IMU / GNSS integrated system).

[0066] When the visual branch identifies a crack in an image, the model immediately uses the pixel's position in the image, combined with spatiotemporal reference data, to calculate the crack's precise geographic coordinates in the real world. At these same coordinates, it automatically consults data from other modalities. Check the lidar data: Has the road surface elevation at this location subsided? Are there any abnormal ruts? Check the ground-penetrating radar data: Is there any abnormality in the signal of the underlying layer at this location? Is there any void? Check the infrared data: Is the temperature at this location significantly different from that of the surrounding normal road surface?

[0067] The model integrates the aforementioned related information at the feature layer or decision layer and then proceeds to the inference stage.

[0068] Scenario 1: Determining the cause of the crack Input: Visual features (cracks) + LiDAR features (subsidence at this location) + Ground penetrating radar features (vacuum beneath this location).

[0069] Fusion reasoning: Based on the knowledge learned by the model, it is determined that the probability of cracks, settlement, and voids occurring simultaneously is extremely high. Therefore, a diagnostic report is generated: A longitudinal crack was found at coordinate XX, and the cause is inferred to be voids in the underlying base layer leading to insufficient structural bearing capacity, accompanied by uneven settlement.

[0070] Scenario 2: Comprehensive assessment of tire tracks Input: LiDAR features (severe rutting) + visual features (surface texture wear, but no cracks) + Ground penetrating radar features (intact structural layers).

[0071] Fusion reasoning: The model infers that ruts are mainly caused by the plastic flow of the asphalt surface layer under heavy load and high temperature, which is a material problem, and the base structure is still healthy.

[0072] S4. Map the diagnostic analysis results to a 3D road model to generate a digital twin and visualize it. Simultaneously, trigger early warnings based on the diagnostic results and link with roadside facilities to form a closed-loop management system. Specifically, this includes: The comprehensive diagnostic report is used as attribute information and mapped and bound to the corresponding geometric position in the 3D road model to dynamically update the state of the road digital twin; On the visualization terminal, the disease information in the digital twin is rendered with different colors, icons or transparency, and users can interact with the disease tags to view detailed diagnostic reports. When the severity of the disease in the comprehensive diagnostic report exceeds a preset threshold, or when the disease type belongs to the sudden high-risk category, a real-time warning is automatically triggered. Based on the triggered warning information, specific equipment control commands are generated and sent to the corresponding roadside facilities through the communication network; The roadside facilities include at least variable message signs, broadcasting systems, lighting systems, and signal control systems, which change their operating status through the instructions to actively guide and warn traffic flow. Record the trigger time of the warning information, the content of the handling instructions, and the response status of the roadside facilities; During subsequent inspections, when the mobile platform passes by the reported location of the defect again, it automatically detects whether the defect has been repaired and updates the repair results to the digital twin and management file, thus completing closed-loop management.

[0073] The triggering conditions for real-time alerts include: When a pothole is detected and its depth is greater than 5 cm or its area is greater than 0.1 square meters, the highest level of warning is immediately triggered. When a void is detected in the base layer and its area exceeds 1 square meter, an advanced warning is triggered. When cracks, settlement, and base layer delamination occur simultaneously in the same location, the overall risk level is raised and a corresponding early warning is triggered.

[0074] The steps for generating and executing linkage control commands include: Send instructions to the variable message sign 500 meters ahead to publish a graphic warning message containing road defects ahead and urging drivers to drive with caution. Send a command to the lighting system of the corresponding road section to increase the street light brightness to 100%; Send a command to the signal control system of the corresponding road section to switch the traffic lights to flashing yellow warning mode; Send maintenance work orders containing photos of the location of the damage and a detailed diagnostic report to the mobile devices of maintenance personnel.

[0075] Data analyzed by a cross-modal fusion diagnostic AI model is mapped onto a 1:1 three-dimensional digital road model using a specific algorithm. This 3D model, built on a high-precision map, includes information such as road geometry, pavement material, and roadside infrastructure. Detected defects, such as crack location, length, and width, rut depth, and areas of base layer delamination, are visually annotated on the 3D model, forming a holographic digital twin of the road. Managers can use dedicated software on computers or mobile devices to view the real-time health status of the road from different angles, achieving a visualized and holographic perception of the road.

[0076] When the cross-modal fusion diagnostic AI model detects severe potholes, the system immediately generates a warning message. This warning message is transmitted via wireless communication modules (such as 4G / 5G or V2X communication) to roadside information boards, broadcasting systems, and terminal equipment of relevant management departments. Simultaneously, the system can coordinate with lighting and signal control systems. For example, when a severe pothole is detected ahead, the system automatically increases the brightness of the lighting on the road ahead and adjusts the traffic lights to warning mode, guiding vehicles to slow down or avoid the pothole, forming a closed-loop management system of detection, analysis, and response, effectively ensuring road safety.

[0077] In some embodiments, the training steps of the cross-modal fusion diagnostic AI model include: The historical dataset synchronously collected by the multimodal sensing system is acquired, and the historical dataset is spatiotemporally aligned to generate a spatiotemporally aligned training sample set. Each training sample contains multimodal sensor data of the same road location point, and corresponding manually labeled data; the labels include at least: disease type, disease location, severity, and cause of disease. Construct a multi-branch deep learning network model that includes at least: A vision branch based on a fully convolutional neural network, used to process image data; A non-visual branch based on graph neural networks is used to process point cloud and radar data; A feature fusion module is used to fuse features from different branches; Using a spatiotemporally aligned training sample set as input and the corresponding labels as supervision signals, the multi-branch deep learning network model is trained end-to-end. Among them, the loss function of the visual branch is constructed based on pixel-level classification accuracy, the loss function of the non-visual branch is constructed based on graph node classification or regression error, and the overall loss function of the feature fusion module is a weighted sum of the loss functions of each branch, with the addition of a regularization term to promote cross-modal association learning. The performance of the trained model is evaluated using a reserved validation set. When the model’s comprehensive diagnostic accuracy for disease type, location and cause exceeds a preset threshold, a trained cross-modal fusion diagnostic AI model is obtained, and the model parameters are deployed to the edge computing unit.

[0078] Regularization terms used to facilitate cross-modal association learning are implemented by calculating the KL divergence between the attention distributions of different branches to the same disease cause, thus forcing the model to learn consistent causal features across modalities. Disease cause labels are determined by synthesizing results from manual surveys, core drilling, or damage repair records.

[0079] It should be noted that the specific process of end-to-end training of a multi-branch deep learning network model using a spatiotemporally aligned training sample set as input and the corresponding labels as supervision signals is as follows: 1. Forward propagation calculation Input: A batch of data taken from the spatiotemporally aligned training sample set. Each sample is a data packet containing multimodal data from the same location and time.

[0080] The vision branch takes image data as input to a fully convolutional neural network. This network outputs a pixel-level classification map, where each pixel is classified as background, crack, pit, etc.

[0081] Non-visual branch: Non-image data such as LiDAR point clouds and ground-penetrating radar signals are input into the graph neural network. The GNN treats the road structure as a graph, where nodes can be points in the point cloud or pre-divided grids, and edges represent spatial or logical relationships. This branch outputs node-level or graph-level feature representations, used to determine, for example, whether "this area has subsided" or "this unit's base layer is healthy".

[0082] The high-level features output from the two branches are fused. For example, the crack feature map from the visual branch and the basic health status feature map from the GNN branch are combined by feature stitching or attention-weighted fusion under a unified geographic coordinate system to form a unified joint feature representation containing multimodal information.

[0083] 2. Loss Calculation The model learns through a loss function. The total loss function consists of several parts, each corresponding to a different supervision label: Visual branch loss: Using the cross-entropy loss function, the pixel-level classification map output by FCN is compared with the manually labeled pixel-level disease tags, penalizing incorrect pixel classifications. This forces the visual branch to learn to accurately identify surface diseases.

[0084] Non-visual branch loss: Depending on the task, mean squared error loss (for regression problems, such as sedimentation) or binary cross-entropy loss (for classification problems, such as whether a void has been cleared) may be used. This loss function compares the output of the GNN with the corresponding label (such as "settlement value" or "void cleared").

[0085] Total loss after fusion: The sum of the losses from the above branches constitutes the overall objective of the model. However, to achieve correlation diagnosis, a crucial step must be added. item. During the design process, the correlation between the visual and non-visual lesion feature vectors output by the model is calculated and compared with the correlation of the true labels. If "crack" and "void" always appear simultaneously in the true labels, the loss function will encourage the model to extract the "void" feature when it extracts the "crack" feature.

[0086] 3. Backpropagation and parameter update Calculate total loss (in (This is a hyperparameter used to balance the weights of each task).

[0087] The gradient of the total loss with respect to all parameters (weights and biases) in the model is calculated using the backpropagation algorithm. The gradient indicates the direction and magnitude in which the parameters should be adjusted to reduce the loss.

[0088] Use an optimizer (such as Adam) to update all parameters in the model based on the calculated gradients.

[0089] 4. Iteration and Verification Repeat steps 1-3, iterating through all samples in the training set multiple times (Epochs) until the model loss converges, meaning the model performance no longer improves significantly.

[0090] After each training epoch or periodically, the performance of the current model is evaluated using a reserved validation set. The evaluation metric is not only to look at classification accuracy, but also to look at the accuracy of inferring the causes of diseases.

[0091] Ultimately, the model parameters that perform best on the validation set are selected and deployed to the edge computing unit.

[0092] like Figure 4 As shown, this embodiment of the invention also provides a highway holographic intelligent detection system based on multimodal sensing and edge AI, used to implement the method described in the first aspect, the system comprising: The perception layer includes a multimodal sensing system mounted on a mobile platform, used to simultaneously collect multimodal data from the road surface and internal structural layers; The processing layer, including edge computing units, is configured as follows: The multimodal data is processed in real time, and a standardized dataset with spatiotemporal alignment is generated through time calibration and spatial alignment. A cross-modal fusion diagnostic AI model is used to perform fusion analysis on the standardized dataset to achieve correlation diagnosis and cause inference of road defects; The application layer is configured as follows: The diagnostic analysis results are mapped to a 3D road model to generate a digital twin and then visualized. Early warnings are triggered based on diagnostic results, and a closed-loop management system is formed in conjunction with roadside facilities.

[0093] It should be noted that the multimodal sensing system of the perception layer includes: High-resolution line scan cameras and panoramic cameras are used to acquire images of the road surface; 3D LiDAR is used to construct high-precision point cloud models of road surfaces at a set scanning frequency; Millimeter-wave radar and ground-penetrating radar are used to monitor traffic flow and detect defects in the structural layers below the road surface, respectively. Infrared thermal imagers are used to acquire infrared images of road surfaces to identify temperature field anomalies. A high-precision positioning and inertial navigation system is used to provide a spatiotemporal reference for all acquired data.

[0094] The edge computing unit of the processing layer further includes: The multi-source data spatiotemporal synchronization module is configured to perform time calibration using hardware trigger signals from a high-precision positioning and inertial navigation system, and to perform spatial alignment using its positioning and attitude data to generate the standardized dataset. The cross-modal fusion diagnostic AI model is a multi-branch neural network structure, including: The visual branch based on fully convolutional neural networks is used to process image data; A non-visual branch of graph neural networks is used to process point cloud and radar data. The feature fusion module is used to fuse features from different branches and realize the correlation diagnosis of multimodal data based on unified geographic coordinates.

[0095] In some embodiments, the application layer further includes: A digital twin engine is used to bind diagnostic results as attribute information to the geometric position of a 3D model, enabling dynamic updates and visualization rendering; The early warning linkage controller is configured to automatically generate control commands and send them to at least one of the variable message signs, broadcasting systems, lighting systems, and signal control systems when the diagnostic results meet preset conditions.

[0096] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for highway holographic intelligent detection based on multi-modal sensing and edge AI, characterized in that, Includes the following steps: Multimodal sensing system mounted on a mobile platform is used to simultaneously collect multimodal data of the road surface and internal structural layers; The multimodal data is processed in real time using edge computing units, and a standardized dataset with spatiotemporal alignment is generated through time calibration and spatial alignment. A cross-modal fusion diagnostic AI model is used to perform fusion analysis on the standardized dataset to achieve correlation diagnosis and cause inference of road defects; The diagnostic analysis results are mapped to a 3D road model to generate a digital twin and visualize it. At the same time, early warnings are triggered based on the diagnostic results, and a closed-loop management system is formed in conjunction with roadside facilities.

2. The highway holographic intelligent detection method based on multimodal sensing and edge AI according to claim 1, characterized in that, The multimodal sensing system includes a visual sensor, a 3D lidar, a radar sensor, an infrared thermal imager, and a high-precision positioning and inertial navigation system. The steps of multimodal data acquisition include: Road surface images are acquired by a high-resolution line scan camera and a panoramic camera mounted on a mobile platform; wherein, the line scan camera continuously takes pictures at a first frequency to obtain road surface detail images while the mobile platform is moving, and the panoramic camera periodically takes pictures at a second frequency to obtain panoramic road view images. A high-precision point cloud model of the road surface is constructed by emitting laser beams and receiving reflected signals through a three-dimensional lidar with a scanning frequency within a set range. This model is used to calculate rut depth, smoothness, and road geometry parameters. The system uses millimeter-wave radar to monitor traffic flow around vehicles in real time, and uses ground-penetrating radar to emit high-frequency electromagnetic waves that penetrate the road surface structure layer and receive reflected signals to analyze the thickness, voids, and looseness of the structure layer within a set depth range below the road surface. Infrared images of the road surface are acquired by an infrared thermal imager, and the differences in the road surface temperature field are analyzed to identify areas of voids or water seepage under the asphalt pavement layer. High-precision positioning and inertial navigation systems provide spatiotemporal references for all collected data, enabling geocoding and preliminary time stamping of the data.

3. The highway holographic intelligent detection method based on multimodal sensing and edge AI according to claim 2, characterized in that, The steps of using edge computing units to process the multimodal data in real time and generating a spatiotemporally aligned standardized dataset through time calibration and spatial alignment include: The precise time signal generated by the high-precision positioning and inertial navigation system is used as a unified hardware trigger signal and sent to each sensor to trigger each sensor to synchronously start or mark the data acquisition process. In the edge computing unit, a time correction is applied to the collected data based on the acquisition and transmission delay of each sensor, so that all data are aligned on a unified time axis. By utilizing the high-precision positioning and attitude data generated by the high-precision positioning and inertial navigation system, all sensor data are unified under the same global geographic coordinate system, thereby generating a standardized dataset with spatiotemporal alignment.

4. The highway holographic intelligent detection method based on multimodal sensing and edge AI according to claim 3, characterized in that, The time correction amount is dynamically calculated based on the time difference between the arrival of the hardware trigger signal and the sensor data packet at the edge computing unit.

5. The highway holographic intelligent detection method based on multimodal sensing and edge AI according to claim 4, characterized in that, The steps for using a cross-modal fusion diagnostic AI model to perform fusion analysis on the standardized dataset to achieve correlation diagnosis and causal inference of road defects include: The spatiotemporally aligned standardized dataset is input into a multi-branch neural network for feature extraction; wherein, visual image data is input into a fully convolutional neural network to extract pixel-level feature maps, and 3D point cloud and non-image data are input into a graph neural network to construct a graph structure with spatial location as nodes and extract relational features; Features from different neural network branches are concatenated or weighted and fused at the feature layer, or the preliminary diagnostic results of each branch are comprehensively reasoned at the decision layer. When the visual branch identifies road surface cracks, the model automatically associates and calls up lidar point cloud data at the same location based on the spatial coordinates of the cracks to determine whether there is subsidence, calls ground-penetrating radar data to assess the health status of the base layer, and generates a comprehensive diagnostic report that includes the type, location, severity, and cause inference of the disease.

6. The highway holographic intelligent detection method based on multimodal sensing and edge AI according to claim 5, characterized in that, The specific steps of the associated diagnostic process include: The pixel coordinates of the road surface cracks identified by the visual branch are combined with the corresponding high-precision positioning data in the spatiotemporally aligned standardized dataset and converted into unified global geographic coordinates. Based on the global geographic coordinates, a spatial range query is performed in the standardized dataset to automatically retrieve and call the data associated with that coordinate location: Three-dimensional lidar point cloud data is used to calculate the road elevation change at that location in order to determine whether there is subsidence; Ground-penetrating radar reflected wave data is used to analyze the electromagnetic wave characteristics and phase axis continuity of the road structure layer below the location to assess whether there is voiding, loosening or abnormal moisture content in the base layer. By integrating crack characteristics, settlement assessment results, and basic health status evaluation results, and based on a pre-defined disease causal knowledge graph, a comprehensive diagnostic report is generated. The report includes at least the disease type, precise geographical location, severity level, and inferences about the potential causes of the crack.

7. The highway holographic intelligent detection method based on multimodal sensing and edge AI according to claim 6, characterized in that, The steps involved in mapping diagnostic analysis results to a 3D road model to generate a digital twin and visualize it, triggering early warnings based on the diagnostic results, and linking with roadside facilities to form a closed-loop management system include: The comprehensive diagnostic report is used as attribute information and mapped and bound to the corresponding geometric position in the 3D road model to dynamically update the state of the road digital twin; On the visualization terminal, the disease information in the digital twin is rendered with different colors, icons or transparency, and users can interact with the disease tags to view detailed diagnostic reports. When the severity of the disease in the comprehensive diagnostic report exceeds a preset threshold, or when the disease type belongs to the sudden high-risk category, a real-time warning is automatically triggered. Based on the triggered warning information, specific equipment control commands are generated and sent to the corresponding roadside facilities through the communication network; The roadside facilities include at least variable message signs, broadcasting systems, lighting systems, and signal control systems, which change their operating status through the instructions to actively guide and warn traffic flow.

8. The highway holographic intelligent detection method based on multimodal sensing and edge AI according to claim 7, characterized in that, The method further includes: Record the trigger time of the warning information, the content of the handling instructions, and the response status of the roadside facilities; During subsequent inspections, when the mobile platform passes by the reported location of the defect again, it automatically detects whether the defect has been repaired and updates the repair results to the digital twin and management file, thus completing closed-loop management.

9. The highway holographic intelligent detection method based on multimodal sensing and edge AI according to claim 8, characterized in that, The training steps for a cross-modal fusion diagnostic AI model include: Acquire historical datasets synchronously collected by a multimodal sensing system, and perform spatiotemporal alignment on the historical datasets to generate a spatiotemporally aligned training sample set; Each training sample contains multimodal sensor data of the same road location point, and corresponding manually labeled data; the labels include at least: disease type, disease location, severity, and disease cause. Construct a multi-branch deep learning network model that includes at least: A vision branch based on a fully convolutional neural network, used to process image data; A non-visual branch based on graph neural networks is used to process point cloud and radar data; A feature fusion module is used to fuse features from different branches; Using a spatiotemporally aligned training sample set as input and the corresponding labels as supervision signals, the multi-branch deep learning network model is trained end-to-end. Among them, the loss function of the visual branch is constructed based on pixel-level classification accuracy, the loss function of the non-visual branch is constructed based on graph node classification or regression error, and the overall loss function of the feature fusion module is a weighted sum of the loss functions of each branch, with the addition of a regularization term to promote cross-modal association learning. The performance of the trained model is evaluated using a reserved validation set. When the model’s comprehensive diagnostic accuracy for disease type, location and cause exceeds a preset threshold, a trained cross-modal fusion diagnostic AI model is obtained, and the model parameters are deployed to the edge computing unit.

10. A highway holographic intelligent detection system based on multimodal sensing and edge AI, used to implement the method of any one of claims 1 to 9, characterized in that, The system includes: The perception layer includes a multimodal sensing system mounted on a mobile platform, used to simultaneously collect multimodal data from the road surface and internal structural layers; The processing layer, including edge computing units, is configured as follows: The multimodal data is processed in real time, and a standardized dataset with spatiotemporal alignment is generated through time calibration and spatial alignment. A cross-modal fusion diagnostic AI model is used to perform fusion analysis on the standardized dataset to achieve correlation diagnosis and cause inference of road defects; The application layer is configured as follows: The diagnostic analysis results are mapped to a 3D road model to generate a digital twin and then visualized. Early warnings are triggered based on diagnostic results, and a closed-loop management system is formed in conjunction with roadside facilities.