Road surface collapse risk monitoring method and system
Through the combination of visible light and infrared images, lightweight convolutional neural networks and multimodal knowledge graphs, the accuracy and timeliness of urban road monitoring in the existing technology are solved, and a more accurate assessment of road collapse risk is achieved.
Patent Information
- Application Number
- CN202510804099.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
In the monitoring of urban road safety, it is difficult to achieve high-frequency and full-coverage dynamic monitoring, resulting in the initial subtle cracks and underground hidden dangers being easily ignored. The existing monitoring technology has limitations in accuracy and timeliness.
The cross-modal fusion of visible light and infrared images is used, combined with lightweight convolutional neural networks and multimodal knowledge graphs for hierarchical causal reasoning, and the risk level of road surface collapse is determined through reinforcement learning models.
It improves the accuracy and real-time monitoring of road collapse risk, can more accurately reflect the actual status of the road surface and quickly identify potential collapse factors.
Smart Images

Figure CN120339962A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of road safety monitoring, and particularly to a method and system for monitoring the risk of road surface collapse. Background Art
[0002] Roads are key infrastructure of cities, carrying out the daily traffic tasks of people and vehicles. Once a road collapses, serious accidents such as vehicle falling and casualties may be triggered instantly, and the consequences are unimaginable. Moreover, if the collapse is not detected in time, it will subsequently lead to a large-scale paralysis of traffic, the obstruction of surrounding commercial activities, and the lives of residents will also be greatly disturbed due to inconvenient travel. Therefore, it is crucial to monitor road collapses in real time.
[0003] Currently, in the field of urban road safety monitoring, single or simply combined monitoring technologies are mostly adopted, and these monitoring technologies have obvious defects. Among them, although manual regular inspections can detect potential hazards based on experience, limited by the limited human resources and fixed inspection cycles, they cannot conduct high-frequency and full-coverage dynamic monitoring of road conditions, resulting in the easy neglect of hidden risks such as initial fine cracks and underground holes during inspection intervals. Although single-sensor monitoring can collect some data in real time, it is difficult to comprehensively capture complex risk signals caused by the superposition of multiple factors such as groundwater level fluctuations, soil settlement, and pipeline leakage. Therefore, in the face of urban environments with variable geological conditions and complex underground structures, the limitations of existing monitoring technologies in terms of accuracy and timeliness need to be broken through by innovative technical means. Summary of the Invention
[0004] This application provides a method and system for monitoring the risk of road surface collapse to improve the accuracy and real-time performance of road surface collapse risk monitoring.
[0005] In a first aspect, this application provides a method for monitoring the risk of road surface collapse, including: Collect visible light images and infrared images of the target road area, and perform cross-modal fusion on the visible light images and the infrared images through a generative adversarial network constrained by a thermodynamic equation to generate a road surface image; Input the road surface image into a lightweight convolutional neural network to extract multi-scale image features. When the multi-scale image features meet the preset conditions, update the initial collapse factor set based on the feature distribution divergence difference within a preset sliding time window to obtain a target collapse factor set; Construct a multi-modal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data, and perform hierarchical causal reasoning on the multi-modal knowledge graph through a gated spatio-temporal graph convolutional network to obtain a first collapse feature, and determine a second collapse feature based on the first collapse feature and multi-source environmental data; Input the multi-scale image features and the second collapse feature into a reinforcement learning model to determine the road surface collapse risk level.
[0006] In the embodiment of the present application, by collecting visible light and infrared images, the morphological features and structural features of the road surface can be intuitively obtained. Through cross-modal fusion, the advantages of these two types of information can be complementary to more accurately reflect the actual state of the road surface, facilitating the subsequent precise and rapid extraction of features related to collapse; by inputting the road surface image into a lightweight convolutional neural network, information at different spatial scales of the road surface can be captured; by updating the collapse factor set, features related to road surface collapse can be accurately and rapidly determined, improving the accuracy and real-time performance of monitoring; by constructing a multi-modal knowledge graph, rich multi-source information is provided for subsequent causal reasoning; through hierarchical causal reasoning on the multi-modal knowledge graph, the internal connections and causal relationships between different data can be deeply explored. By determining the second collapse feature by integrating multi-source environmental data, the influence of the current actual environmental situation on road surface collapse is considered, and the second collapse feature that more comprehensively reflects the actual environmental conditions of the current road surface is obtained; by inputting the multi-scale image features and the second collapse feature into a reinforcement learning model, the associations and complementarities between different features can be fully exploited, and the influence of these two features on road surface collapse risk is comprehensively considered to generate a more accurate road surface collapse risk level. Compared with the prior art, the present application can improve the accuracy and real-time performance of road surface collapse risk monitoring.
[0007] Further, the generative adversarial network constrained by the thermodynamic equation performs cross-modal fusion on the visible light image and the infrared image to generate a road surface image, specifically as follows: Perform image preprocessing on the visible light image and the infrared image respectively to obtain a first preprocessing result and a second preprocessing result; Stitch the first preprocessing result and the second preprocessing result to obtain a stitched image, and input the stitched image into the generative adversarial network to adaptively enhance the stitched image in combination with network parameters to generate the road surface image. Among them, the generative adversarial network includes a generator and a discriminator. The network parameters calculate the matching error between the infrared channel temperature field of the generated image generated by the generator and the Fourier heat conduction equation, and backpropagate the matching error to optimize the generated image generated by the generator. After receiving the generated image, the discriminator determines the classification result of the generated image and the stitched image, and optimizes to obtain the loss function value according to the classification result.
[0008] In this way, through cross-modal fusion, the advantages of these two types of information can be complementary to more accurately reflect the actual state of the road surface, facilitating the subsequent precise and rapid extraction of features related to collapse.
[0009] Further, the relevant formula for calculating the matching error between the infrared channel temperature field of the generated image and the Fourier heat conduction equation is specifically as follows: ; In the formula, is the matching error; is a preset weight coefficient used to balance the matching error and other loss terms; is the temperature value at position in the infrared channel of the generated image; is the Laplace operator of the temperature field; is the rate of change of temperature with time; is the thermal diffusivity of the material, which characterizes the diffusion speed of heat in the material and can be dynamically adjusted according to real-time environmental temperature and humidity data.
[0010] Further, the process of inputting the road surface image into the lightweight convolutional neural network to extract multi-scale image features is specifically as follows: Input the road surface image into the lightweight convolutional neural network to extract an initial feature map, where the initial feature map includes a shallow feature map, a middle feature map, and a deep semantic feature map; Align and flatten each scale feature map in the initial feature map into a corresponding spatial sequence, perform position encoding on each position in each pair of the spatial sequences to obtain a first encoding result corresponding to each level, calculate the feature correlation of each first encoding result through a multi-head attention mechanism to generate a self-attention weight corresponding to each level, and obtain the multi-scale image features based on each self-attention weight and the initial feature map.
[0011] In this way, by inputting the road surface image into the lightweight convolutional neural network, information about the road surface at different spatial scales can be captured.
[0012] Further, the process of updating the initial collapse factor set based on the feature distribution divergence difference within a preset sliding time window to obtain a target collapse factor set is specifically as follows: Obtain the feature distribution of the sliding time window and calculate the feature distribution divergence difference between the feature distribution and the historical feature distribution; If the feature distribution divergence difference is greater than a preset difference threshold, extract the abnormal features within the sliding time window, perform incremental clustering based on the abnormal features to determine candidate collapse factors, calculate the cosine similarity between the candidate collapse factors and each collapse factor in the initial collapse factor set, and if the cosine similarity is greater than a preset similarity threshold, merge the candidate collapse factor with the nearest collapse factor, otherwise add the candidate collapse factor to the initial collapse factor set to obtain the updated target collapse factor set.
[0013] In this way, by updating the set of collapse factors, the features related to road collapse can be accurately and quickly determined, improving the accuracy and real-time performance of monitoring.
[0014] Furthermore, constructing a multimodal knowledge graph based on the set of collapse factors, historical collapse data, and historical environmental data is specifically as follows: Obtain the historical collapse data and corresponding historical environmental data of the target road area, and perform data preprocessing on the historical collapse data and the historical environmental data respectively to obtain the third preprocessing result and the fourth preprocessing result. Among them, the historical collapse data includes historical image data and historical text data, and the historical environmental data includes historical ground penetrating radar detection data, historical temperature and humidity sensor data, and historical traffic load time series data; Extract a number of entities from the third preprocessing result and the fourth preprocessing result, and determine the entity types corresponding to each entity and the relationships between the entities. Among them, the entity types include road structure entities, environmental entities, collapse event entities, and collapse factor entities; Map each entity to a node in the knowledge graph according to the entity type, and map the relationship to an edge in the knowledge graph to obtain a multimodal knowledge graph.
[0015] In this way, by constructing a multimodal knowledge graph, rich multi-source information is provided for subsequent causal reasoning, facilitating the subsequent in-depth exploration of the internal connections and causal relationships between different data to accurately determine the first collapse feature.
[0016] Furthermore, performing hierarchical causal reasoning on the multimodal knowledge graph through a gated spatio-temporal graph convolutional network to obtain the first collapse feature is specifically as follows: Based on the image features, query relevant road entities and road edges in the multimodal knowledge graph, and construct a subgraph based on the road entities and the road edges; Input the subgraph into a preset gated spatio-temporal graph convolutional network to determine node features, perform graph convolution on the node features to obtain a convolution result, perform multi-hop aggregation on the convolution result to obtain target node features, and predict the first collapse feature based on the target node features.
[0017] In this way, by performing hierarchical causal reasoning on the multimodal knowledge graph, the internal connections and causal relationships between different data can be quickly and deeply explored to obtain the first collapse feature.
[0018] Furthermore, determining the second collapse feature based on the first collapse feature and multi-source environmental data is specifically as follows: Obtain the multi-source environmental data of the target road area, and encode the multi-source environmental data to obtain a second encoding result, where the second encoding result includes a radar encoding result, a temperature and humidity encoding result, and a load encoding result; Calculate the correlation weight between the first collapse feature and the radar encoding result, determine a radar enhanced feature based on the correlation weight and the radar encoding result, and splice the temperature and humidity encoding result and the load encoding result to obtain an environmental time series feature; Constrain the radar enhanced feature and the environmental time series feature through the material linear elastic equation to obtain a fusion feature, and based on the first collapse feature and the fusion feature, generate a corresponding dynamic weight through a cross-attention mechanism, and weighted aggregate the first collapse feature and the fusion feature based on the dynamic weight to obtain a second collapse feature.
[0019] In this way, by comprehensively determining the second collapse feature from multi-source environmental data, the influence of the current actual environmental conditions on road surface collapse is considered, and a second collapse feature that more comprehensively reflects the actual environmental conditions of the current road surface is obtained.
[0020] Further, inputting the multi-scale image feature and the second collapse feature into a reinforcement learning model to determine the road surface collapse risk level specifically includes: Perform spatial alignment and cross-modal feature fusion on the image feature and the second collapse feature to obtain a comprehensive feature vector; Use the comprehensive feature vector and the multi-source environmental data as the state space of the reinforcement learning model, and based on the state space, a preset action space, and a preset reward function, use the preset reinforcement learning model to iteratively update the policy network parameters with the goal of maximizing the cumulative value expectation of the reward function, determine the optimal risk assessment strategy, and determine the road surface collapse risk level based on the risk assessment strategy.
[0021] In this way, by inputting the multi-scale image feature and the second collapse feature into the reinforcement learning model, the association and complementarity between different features can be fully explored, the influence of these two features on the road surface collapse risk can be comprehensively considered, and a more accurate road surface collapse risk level can be generated.
[0022] In a second aspect, the present application provides a road surface collapse risk monitoring system, including: a collection module, an update module, an inference module, and a generation module; The collection module is used to collect visible light images and infrared images of the target road area, and perform cross-modal fusion on the visible light image and the infrared image through a generative adversarial network constrained by a thermodynamic equation to generate a road surface image; The update module is configured to input the road surface image into a lightweight convolutional neural network to extract multi-scale image features. When the multi-scale image features meet the preset conditions, update the initial collapse factor set based on the feature distribution divergence difference within a preset sliding time window to obtain a target collapse factor set; The inference module is configured to construct a multi-modal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data, and perform hierarchical causal inference on the multi-modal knowledge graph through a gated spatio-temporal graph convolutional network to obtain a first collapse feature, and determine a second collapse feature based on the first collapse feature and multi-source environmental data; The generation module is configured to input the multi-scale image features and the second collapse feature into a reinforcement learning model to determine the road surface collapse risk level.
[0023] In the embodiment of the present application, by collecting visible light and infrared images, the morphological features and structural features of the road surface can be intuitively obtained. Through cross-modal fusion, the advantages of these two types of information can be complementary to more accurately reflect the actual state of the road surface, facilitating the subsequent accurate and rapid extraction of features related to collapse; by inputting the road surface image into a lightweight convolutional neural network, information on the road surface at different spatial scales can be captured; by updating the collapse factor set, features related to road surface collapse can be accurately and quickly determined, improving the accuracy and real-time performance of monitoring; by constructing a multi-modal knowledge graph, rich multi-source information is provided for subsequent causal inference; by performing hierarchical causal inference on the multi-modal knowledge graph, the internal connections and causal relationships between different data can be deeply explored. By determining the second collapse feature by integrating multi-source environmental data, the influence of the current actual environmental conditions on road surface collapse is considered, and a second collapse feature that more comprehensively reflects the actual environmental conditions of the current road surface is obtained; by inputting the multi-scale image features and the second collapse feature into a reinforcement learning model, the associations and complementarities between different features can be fully exploited, and the influence of these two types of features on the road surface collapse risk is comprehensively considered to generate a more accurate road surface collapse risk level. Compared with the prior art, the present application can improve the accuracy and real-time performance of road surface collapse risk monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic flowchart of an embodiment of the road surface collapse risk monitoring method provided by the present application; Figure 2 is a schematic structural diagram of an embodiment of the road surface collapse risk monitoring system provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts shall fall within the protection scope of the present application.
[0026] It should be understood that the step numbers used in the text are only for convenient description and do not limit the execution order of the steps.
[0027] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0028] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0029] The term "and / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0030] Roads are key urban infrastructure, carrying out daily traffic tasks. Road collapses can trigger serious accidents, causing casualties, traffic paralysis, business disruptions and inconvenience to residents' lives. Real-time monitoring is crucial. Currently, most urban road safety monitoring uses single or simple combined technologies, which have defects. Manual inspections are limited by manpower and cycle, difficult to dynamically monitor, and prone to overlooking initial hidden dangers; the monitoring data of single sensors is one-sided and difficult to capture complex risk signals. In an urban environment with diverse geology and complex underground structures, the accuracy and timeliness of existing technologies urgently need innovative breakthroughs.
[0031] Next, the nouns involved in the present application will be analyzed: Thermodynamic equation constraint means that in certain physical processes or models, it is required to satisfy the laws related to thermodynamics (such as the law of conservation of energy, the second law of thermodynamics, etc.) or their corresponding mathematical equations to restrict and standardize the construction of the process or model.
[0032] A generative adversarial network consists of two parts: a generator and a discriminator. The generator is responsible for generating fake data that is as close as possible to real data, while the discriminator is responsible for distinguishing whether the data is real or generated by the generator. These two parts compete against each other during the training process. The generator continuously learns to generate more realistic data to deceive the discriminator, and the discriminator continuously learns to more accurately identify the authenticity of the data, thus achieving the co-optimization of the generator and the discriminator.
[0033] A multi-modal knowledge graph integrates different types of data (such as text, images, audio, video, etc.) into a unified knowledge graph to represent the complex relationships and semantic information between entities. By fusing and correlating multi-modal data, it constructs a more comprehensive and rich knowledge representation framework, enabling a more accurate description of various things and phenomena in the real world.
[0034] Based on this, the embodiments of the present application provide a method and system for monitoring the risk of road surface collapse, which can improve the accuracy and real-time performance of road surface collapse risk monitoring.
[0035] The method and system for monitoring the risk of road surface collapse provided by the embodiments of the present application are specifically described through the following embodiments. First, the method for monitoring the risk of road surface collapse in the embodiments of the present application is described.
[0036] The method for monitoring the risk of road surface collapse provided by the embodiments of the present application relates to the field of road safety monitoring. The method for monitoring the risk of road surface collapse provided by the embodiments of the present application can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application for implementing a method for monitoring the risk of road surface collapse, etc., but is not limited to the above forms.
[0037] This application can also be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0038] Embodiment 1 Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the road surface collapse risk monitoring method provided by this application, including steps S101 to S104; Step S101: Collect visible light images and infrared images of the target road area, and perform cross-modal fusion on the visible light images and the infrared images through a generative adversarial network constrained by a thermodynamic equation to generate a road surface image; In some embodiments, collecting visible light images and infrared images of the target road area specifically includes: mounting a visible light camera and an infrared thermal imager on a suitable carrier, such as installing them on a road detection vehicle, to ensure that they can simultaneously and at the same angle align with the target road area. During installation, the positions and angles of the devices need to be precisely adjusted so that their fields of view are basically coincident. Then, by setting the parameters of the visible light camera and the infrared thermal imager, the road detection vehicle is made to travel and collect according to a preset route and speed to collect visible light images and infrared images of the target road area.
[0039] It should be noted that visible light images can show the texture of road surface cracks, and infrared images can show abnormal temperature distributions caused by rainwater evaporation. Fusing the two can accurately correspond the crack positions with the temperature anomaly areas, enabling an assessment of the development trend of road surface diseases and facilitating subsequent accurate prediction of road collapse risks.
[0040] In some embodiments, the step of performing cross-modal fusion on the visible light images and the infrared images through a generative adversarial network constrained by a thermodynamic equation to generate a road surface image includes steps S201 to S202: Step S201: Perform image preprocessing on the visible light image and the infrared image respectively to obtain a first preprocessing result and a second preprocessing result. In some embodiments, first, align the spatial coordinates of the infrared image and the visible light image through affine transformation. Then, image preprocessing operations such as resolution standardization, gray-scale normalization, and optical distortion correction can be performed on the visible light image (but not limited to these) to obtain the first preprocessing result. Since the specific image preprocessing operations are not the focus of this application, they will not be elaborated here. Then, obtain the ambient temperature and humidity data in real time through a temperature and humidity sensor, and consult relevant materials or determine the emissivity of the road material through experiments to determine the ambient radiation rate for each pixel point in the infrared image, the original temperature value is calculated, and according to the formula the corrected true temperature value is calculated thus completing the temperature calibration to obtain the second preprocessing result.
[0041] Step S202: Stitch the first preprocessing result and the second preprocessing result to obtain a stitched image, and input the stitched image into the generative adversarial network to adaptively enhance the stitched image in combination with network parameters to generate the road surface image. The generative adversarial network includes a generator and a discriminator. The network parameters are obtained by calculating the matching error between the infrared channel temperature field of the generated image generated by the generator and the Fourier heat conduction equation, and backpropagating the matching error to optimize the generated image generated by the generator. After receiving the generated image, the discriminator determines the classification result of the generated image and the stitched image, and optimizes to obtain the loss function value according to the classification result.
[0042] In some embodiments, first, perform channel stitching on the first preprocessing result and the second preprocessing result to obtain a stitched image. For example, since the visible light image is usually 3 channels (RGB) and the infrared image is 1 channel (temperature information), a 4-channel stitched image is obtained after stitching. Then, input the stitched image into the generative adversarial network, and the generator will output a generated image based on the generator parameters, thereby obtaining the road surface image after adaptive enhancement.
[0043] It should be noted that the network parameters include generator parameters and discriminator parameters. The generator parameters are obtained by extracting the temperature field from the infrared channel of the generated image generated by the generator and calculating the temperature field The matching error with the Fourier heat conduction equation, and backpropagate the calculated matching error to the generator to adjust the generator parameters until the preset conditions are met. Then, the generator can obtain a generated image that preliminarily conforms to the heat conduction law based on the generator parameters. The discriminator determines the classification results of the generated image and the spliced image (judging whether it is a real image or a generated image), and according to the classification results, in accordance with the principle of minimizing the adversarial loss function The relevant formula is: , where in the formula represents the expectation; represents the discriminator; represents the generator; represents the noise or initial image input to the generator; represents the discriminator's judgment result on the image generated by the generator, indicating the probability that the image is a real image; represents the gap between the discriminator's judgment result and the real image label (usually 1). After that, it is also necessary to calculate the pixel-level loss and the perceptual loss . Thus, the total loss function can be determined according to the preset weights. By adjusting the weight coefficients , and of different loss terms, it can be used to balance each term in the total loss function (the contributions of adversarial loss, matching error, pixel-level loss, and perceptual loss), so that the total loss value is minimized to optimize the network parameters of the generative adversarial network. Only at this time can the optimal generator parameters and discriminator parameters be determined.
[0044] It should be noted that in order to improve the generalization ability and robustness of the generative adversarial network, visible light and infrared images in different environments such as rain, fog, and night can be simulated through algorithms to increase the diversity of training data. At the same time, Gaussian noise (variance = 0.1) can be randomly added to the infrared images to simulate the noise interference in actual measurements.
[0045] In some embodiments, the relevant formula for calculating the matching error between the infrared channel temperature field of the generated image and the Fourier heat conduction equation is specifically: ; where in the formula is the matching error; is the preset weight coefficient for balancing the matching error and other loss terms; is the temperature value at position in the infrared channel of the generated image; is the Laplacian operator of the temperature field; is the rate of change of temperature with time; is the thermal diffusivity of the material, which characterizes the diffusion rate of heat in the material and can be dynamically adjusted according to real-time environmental temperature and humidity data.
[0046] It should be noted that the thermal diffusivity is dynamically adjusted according to real-time environmental temperature and humidity data. Among them, the adjustment formula is , where is the reference thermal diffusivity; and are environmental coefficients; and are the temperature and humidity of the current environment respectively; and are the reference temperature and reference humidity respectively.
[0047] It should be noted that the generator adopts a U-Net structure as the backbone network, which includes 5 layers of downsampling (implemented through convolutional layers and LeakyReLU activation functions) and 5 layers of upsampling (implemented through transposed convolutions and skip connections), and a physical constraint module, that is, a thermodynamic consistency branch, is added at the end of the decoder.
[0048] In this way, through cross-modal fusion, the advantages of these two types of information can be complementary to more accurately reflect the actual state of the road surface and facilitate the subsequent accurate and rapid extraction of features related to subsidence.
[0049] Step S102: Input the road surface image into a lightweight convolutional neural network to extract multi-scale image features. When the multi-scale image features meet the preset conditions, update the initial subsidence factor set based on the divergence difference of the feature distribution within a preset sliding time window to obtain the target subsidence factor set; In some embodiments, the inputting the road surface image into a lightweight convolutional neural network to extract multi-scale image features includes steps S301 to S302; Step S301: Input the road surface image into a lightweight convolutional neural network to extract an initial feature map, where the initial feature map includes a shallow feature map, a middle feature map, and a deep semantic feature map; In some embodiments, before inputting the road surface image into the lightweight convolutional neural network, it is necessary to perform normalization processing on the stitched image to adapt to the requirements of network input. After that, the road surface image is input into the lightweight convolutional neural network (such as MobileNetV3-Small), and low-level texture features (such as edges and crack contours) are extracted through the initial convolutional layer, and a shallow feature map is output. Then, local details and channel importance information are fused through depthwise separable convolution and squeeze-and-excitation block (SE Block) to generate a middle-level feature map. After multiple convolutional and downsampling operations, a low-resolution and high-semantic feature map (such as 14×14×512) is output, which characterizes the global road structure and potential abnormal areas.
[0050] Step S302: Align the feature maps of each scale in the initial feature map and flatten them into corresponding spatial sequences, and perform position encoding on each position in each of the spatial sequences to obtain the first encoding result corresponding to each level. Then, calculate the feature correlation of each of the first encoding results through the multi-head attention mechanism to generate the self-attention weight corresponding to each level, and based on each of the self-attention weights and the initial feature map, obtain the multi-scale image features.
[0051] In some embodiments, aligning the feature maps of each scale in the initial feature map and flattening them into corresponding spatial sequences specifically means: Since the initial feature map contains feature maps of different scales (such as shallow, middle, and deep feature maps), and there are differences in size (resolution) and number of channels among these feature maps, therefore, for subsequent unified processing, it is necessary to first align them in the spatial dimension, for example, by upsampling or downsampling operations to make them have the same size specification. After alignment, the feature map of each scale is flattened from a three-dimensional tensor (height × width × number of channels) into a two-dimensional spatial sequence (number of pixels × number of channels) to obtain a shallow feature sequence, a middle-level feature sequence, and a deep semantic feature sequence, so as to facilitate the operation and analysis of the features at each position. For example, the original middle-level feature map is 56×56×128, and after flattening, it becomes (56×56)×128, and each position represents a pixel point and its corresponding channel feature information.
[0052] In some embodiments, position encoding is performed on each position in each pair of the spatial sequences to obtain a first encoding result corresponding to each level. Specifically, since after flattening the feature map into a spatial sequence, each position is just a numerical vector and loses the original spatial position information in the image, therefore, through learnable 2D sine encoding or other means, spatial coordinate information is embedded into each position in the spatial sequences (shallow feature sequence, middle feature sequence, and deep semantic feature sequence) to obtain a first encoding result corresponding to each level (including shallow encoding result, middle encoding result, and deep encoding result), so as to facilitate subsequent understanding of the image structure and feature correlation.
[0053] In some embodiments, the feature correlation of each of the first encoding results is calculated through a multi-head attention mechanism to generate self-attention weights corresponding to each level. Specifically, first, the first encoding results (i.e., feature sequences, including shallow encoding results, middle encoding results, and deep encoding results) are divided into multiple attention heads, and each head independently learns the information in different feature subspaces to learn the correlation between features from multiple different perspectives; then, for each attention head, the feature sequence is respectively mapped into Query (query vector), Key (key vector), and Value (value vector) through linear projection; then, the matrix multiplication of Query (query vector) and the transpose of Key (key vector) is used, and then divided by a scaling factor , where is the dimension of each attention head in the attention mechanism to obtain a calculation result. Finally, the calculation result is converted into a probability distribution through the Softmax function to obtain the self-attention weight, where this weight represents the degree of correlation between the feature at the current position and the features at all other positions. The higher the degree of correlation, the greater the weight.
[0054] In some embodiments, based on each of the self-attention weights and the initial feature map, the multi-scale image features are obtained. Specifically, the self-attention weights calculated by each attention head are used to perform weighted aggregation on Value (value vector), that is, for the feature values at each position in Value, weighted summation is performed according to the degree of correlation represented by the self-attention weight. Exemplarily, if the weight of a certain position in the self-attention weight is larger, then the feature value at this position in the corresponding Value contributes more in the weighted aggregation. In this way, the feature representation after weighted aggregation of each attention head is obtained. Then, the feature representations after weighted aggregation of all attention heads are concatenated to obtain a feature representation that fuses multi-head attention information. Finally, the feature representation that fuses multi-head attention information is added to the initial feature map to obtain multi-scale image features.
[0055] It should be noted that Query is used to find information related to itself, Key is used to represent the information available for query, and Value is the eigenvalue that is actually weighted and aggregated.
[0056] By inputting the road surface image into the lightweight convolutional neural network in this way, information on the road surface at different spatial scales can be captured.
[0057] In some embodiments, when the multi-scale image features meet the preset conditions, specifically, calculate the local variance of the multi-scale image features within the sliding window, and compare the local variance with the dynamic threshold, where the dynamic threshold is determined according to the variance of the historical normal data features. If the local variance of the multi-scale image features is greater than the dynamic threshold and the number of times exceeding the dynamic threshold is greater than 3 times, it is determined to be an abnormal state, and the collapse factor set needs to be updated.
[0058] In some embodiments, updating the initial collapse factor set based on the divergence difference of the feature distribution within the preset sliding time window to obtain the target collapse factor set includes steps S401 to S402: Step S401, obtain the feature distribution of the sliding time window, and calculate the divergence difference of the feature distribution between the current feature distribution and the historical feature distribution; In some embodiments, by setting a sliding time window (such as the window length is 30 frames, the time span is 30 seconds, and the sliding step is 5 frames), the image feature vectors within the window are cached in real time to capture the dynamic changes of the image features in the time dimension. As time goes by, the window slides continuously, and new feature vectors are continuously collected, so as to obtain the distribution of the image features within the current window, that is, to determine the feature distribution. Exemplarily, when monitoring the road surface collapse risk, the road surface images at different time points will contain feature information such as crack expansion and vehicle load. The sliding window can gather these features that change with time for analysis. Then, the Wasserstein distance is used to measure the difference between the current sliding window feature distribution P and the historical normal distribution Q, that is, the divergence difference of the feature distribution. The relevant formula is: , where is the Wasserstein distance; P is the current sliding window feature distribution; Q is the historical normal distribution; belongs to all possible joint distributions , which is used to describe how to combine the sliding window feature distribution P and the historical normal distribution Q; , are randomly selected samples; represents in the joint distribution Under the above conditions, the expected value of the Euclidean distance between the sample x drawn from the distribution P and the sample y drawn from the distribution Q; and the Sinkhorn algorithm is used for approximate solution to quantify the deviation degree of the current feature distribution P from the historical normal state Q to obtain the feature distribution divergence difference. If an abnormal situation occurs on the current road surface (such as a sudden increase in cracks or an abnormal change in the groundwater level), then its feature distribution P will have a large difference from the historical normal distribution Q, and the Wasserstein distance will increase.
[0059] Step S402, if the feature distribution divergence difference is greater than a preset difference threshold, extract the abnormal features within the sliding time window, and perform incremental clustering based on the abnormal features to determine candidate collapse factors. Calculate the cosine similarity between the candidate collapse factors and each collapse factor in the initial collapse factor set. If the cosine similarity is greater than the preset similarity threshold, merge the candidate collapse factor with the nearest collapse factor; otherwise, add the candidate collapse factor to the initial collapse factor set to obtain the updated target collapse factor set.
[0060] In some embodiments, if the feature distribution divergence difference is greater than a preset difference threshold, extract the abnormal features within the sliding time window. Specifically: when the calculated Wasserstein distance is greater than the preset difference threshold (which can be understood as the preset Wasserstein distance threshold for determining whether the feature distribution is significantly abnormal), it indicates that the feature distribution within the current sliding window is significantly different from the historical normal distribution, that is, the current feature distribution has undergone a significant shift, that is, some abnormal situations have occurred. At this time, it is necessary to screen out those abnormal features that cause the distribution shift from the numerous features within the current window. For example, in the scenario of road surface collapse risk monitoring, originally the width of the road surface cracks fluctuates within a certain range. If it is found in a certain monitoring that the crack width exceeds the historical normal range, this crack width feature that exceeds the range belongs to the abnormal feature and needs to be extracted for further analysis of the potential relationship between these abnormal features and road surface collapse.
[0061] In some embodiments, perform incremental clustering based on the abnormal features to determine candidate collapse factors. Specifically: after extracting the abnormal features, in order to analyze these abnormal features more systematically, an incremental clustering algorithm (such as the StreamKM++ algorithm) is used to cluster these abnormal features to group similar abnormal features into one category. Each category represents a possible pattern or factor related to road surface collapse, and these clustering results form the candidate collapse factors. , for example, group similar abnormal features such as abnormal crack expansion and abnormal local settlement into one category, and the comprehensive feature represented by this category may be a candidate collapse factor. , which reflects a potential combination of factors leading to road surface collapse.
[0062] It should be noted that incremental clustering can dynamically adjust the clustering results while continuously receiving new data (i.e., new abnormal features).
[0063] In some embodiments, the cosine similarity between the candidate collapse factor and each collapse factor in the initial collapse factor set is calculated. Specifically: Since the initial collapse factor set is a set of factors related to road surface collapse that have been determined previously (which can be determined through the analysis of historical collapse events in the early stage), at this time, the candidate collapse factor and each collapse factor in the initial collapse factor set The cosine similarity is calculated to measure the candidate collapse factor The similarity degree with the existing collapse factor If the candidate collapse factor is similar in characteristics to an existing collapse factor It indicates that they may represent similar collapse risk factors or patterns. For example, the newly discovered candidate collapse factor "rapid expansion of cracks in sections frequently passed by heavy vehicles recently" may have a high cosine similarity with "structural damage to the road surface caused by heavy vehicle loads" in the initial collapse factor set because they are both related to the impact of heavy vehicle loads on the road surface.
[0064] In some embodiments, if the cosine similarity is greater than the preset similarity threshold, the candidate collapse factor is merged with the nearest collapse factor, otherwise the candidate collapse factor is added to the initial collapse factor set to obtain the updated target collapse factor set. Specifically: When the candidate collapse factor and a certain collapse factor in the initial collapse factor set The cosine similarity is greater than the preset similarity threshold (such as 0.9), indicating that they are very similar in characteristics and the collapse risk patterns they represent. At this time, the candidate collapse factor is merged with the nearest (highest similarity) collapse factor , where the merging method usually adopts methods such as weighted average to integrate the feature information of the two to update the representation of the existing collapse factor so that it can more comprehensively reflect the relevant collapse risk factors. If the cosine similarity does not meet the preset threshold, it means that the candidate collapse factor represents a new collapse risk factor or pattern different from the existing collapse factor , then it is added to the initial collapse factor set Thereby enriching the content of the collapse factor set, enabling it to cover more factors that may cause road surface collapse, and providing a more comprehensive basis for more accurately evaluating the road surface collapse risk in the subsequent stage. After the above processing (merging or adding) of the candidate collapse factors, the initial collapse factor set is updated to form the target collapse factor set as . This updated set includes both the stable collapse patterns that have been determined historically (i.e., the original factors in the initial collapse factor set) and the newly discovered collapse risk factors that reflect the current environmental changes (i.e., the newly added or merged candidate collapse factors). Such a target collapse factor set can better meet the requirements of road surface collapse risk assessment at different times and in different environments, providing more accurate and comprehensive basic data for subsequent risk assessment based on these collapse factors.
[0065] It should be noted that both the preset difference threshold and the preset similarity threshold can be set in advance according to the results of historical data analysis, and this application does not make any restrictions.
[0066] It should be noted that if there are requirements for the number of the collapse factor set, it can be sorted according to the most recent update time to eliminate the oldest or least active factors.
[0067] In this way, by updating the collapse factor set, the characteristics related to road surface collapse can be accurately and quickly determined, improving the accuracy and real-time performance of monitoring.
[0068] Step S103: Construct a multi-modal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data, and perform hierarchical causal reasoning on the multi-modal knowledge graph through a gated spatio-temporal graph convolutional network to obtain the first collapse feature, and determine the second collapse feature based on the first collapse feature and multi-source environmental data; In some embodiments, the constructing of the multi-modal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data includes steps S501 to S503: Step S501: Obtain the historical collapse data and the corresponding historical environmental data of the target road area, and perform data preprocessing on the historical collapse data and the historical environmental data respectively to obtain the third preprocessing result and the fourth preprocessing result, where the historical collapse data includes historical image data and historical text data, and the historical environmental data includes historical ground penetrating radar detection data, historical temperature and humidity sensor data, and historical traffic load time series data; In some embodiments, since the historical collapse data and historical environmental data of the target road area are the basis for constructing a multi-modal knowledge graph. Among them, the historical collapse data can reflect the actual situation of past collapses in the area, such as location, time, collapse degree, etc., which can be obtained by retrieving archives; while the historical environmental data covers the historical information of various environmental factors related to the road, including geological conditions, temperature and humidity changes, traffic load conditions, etc., which can be obtained through sensors and other means. By obtaining these data, it is convenient to comprehensively understand the possible correlations between road collapses and environmental factors subsequently, providing data support for subsequent analysis and graph construction. For example, by analyzing the historical collapse data, it can be known which sections are more prone to collapse, and combined with the historical environmental data, it can be explored which environmental factors led to these collapse events. Subsequently, since the original historical collapse data and historical environmental data may have problems such as noise, missing values, and inconsistent formats, directly using them will affect the accuracy and reliability of subsequent analysis. Therefore, they need to be preprocessed. Among them, for the historical collapse data (including historical image data and historical text data), image denoising, enhancement, etc. need to be performed on the images, and cleaning, format normalization, etc. need to be performed on the text data; while for the historical environmental data (including historical ground penetrating radar detection data, historical temperature and humidity sensor data, and historical traffic load time series data), missing values and outliers of the data may need to be processed, and data standardization, etc. need to be performed. Through these preprocessing operations, relatively clean and standardized data, that is, the third preprocessing result and the fourth preprocessing result, can be obtained to more effectively extract valuable information from them.
[0069] Step S502, extract a number of entities from the third preprocessing result and the fourth preprocessing result, and determine the entity type corresponding to each entity and the relationship between each entity, where the entity type includes a road structure entity, an environmental entity, a collapse event entity, and a collapse factor entity; In some embodiments, by analyzing the data and applying relevant algorithms (such as causal discovery algorithms), entities are mined from the third preprocessing result and the fourth preprocessing result, and the entities are classified to determine the entity type and the relationships between the entities. By determining the entity type (such as road structure entities, environmental entities, collapse event entities, and collapse factor entities, etc.), it helps to classify and manage and understand the entities. Different types of entities have different roles and attributes in the knowledge graph. At the same time, the relationships between entities reflect their interconnections in the real world. For example, there may be an influence relationship between a road structure entity (such as a road segment unit) and an environmental entity (such as temperature and humidity), and there may be a causal relationship between a collapse event entity (such as a certain collapse) and a collapse factor entity (such as crack density). Therefore, by mining the entity type and the internal connections between entities, relationship information can be provided for the construction of the knowledge graph.
[0070] Step S503: Map each entity to a node in the knowledge graph according to the entity type, and map the relationship to an edge in the knowledge graph to obtain a multimodal knowledge graph.
[0071] In some embodiments, since the knowledge graph represents knowledge in the form of nodes and edges, where entities are the carriers of knowledge, and different types of entities correspond to different types of nodes. Therefore, map the corresponding entities to different node types in the knowledge graph according to the entity type. For example, map the road structure entity to the physical topology node, the environmental entity to the environmental state node, the collapse event entity to the historical collapse node, and the collapse factor entity to the collapse factor node. In this way, the nodes in the knowledge graph represent various entities in the real world, laying a foundation for expressing the relationships between entities in the graph subsequently. Additionally, in the knowledge graph, edges are used to represent the relationships between nodes (entities). Then, map the previously determined relationships between entities (such as spatial proximity edges, causal dependence edges, dynamic interaction edges, etc.) to the edges in the knowledge graph, thus establishing the connections between nodes, enabling the knowledge graph to completely express the association relationships between entities. In this way, present the information in the historical collapse data and historical environmental data in a structured form to form a multimodal knowledge graph. This graph integrates various types of data and relationships, and can represent the knowledge related to road collapse from multiple dimensions, providing a powerful tool for subsequent road collapse analysis, risk assessment, etc. using the knowledge graph.
[0072] It should be noted that after determining the multimodal knowledge graph, it is also necessary to update it regularly. The causal discovery algorithm can be re-run based on the latest multi-source environmental data for updating, and the attributes of the environmental state nodes (temperature and humidity, traffic load) are updated in real time.
[0073] In this way, by constructing the multimodal knowledge graph, rich multi-source information is provided for subsequent causal reasoning, facilitating the subsequent in-depth exploration of the internal connections and causal relationships between different data to accurately determine the first collapse feature.
[0074] In some embodiments, the hierarchical causal reasoning of the multimodal knowledge graph by the gated spatio-temporal graph convolutional network to obtain the first collapse feature includes steps S601 to S602; Step S601: Query the relevant road entities and road edges in the multimodal knowledge graph based on the image features, and construct a subgraph based on the road entities and the road edges; In some embodiments, first, since the multimodal knowledge graph integrates various information about roads, including road structure, environmental factors, historical collapse situations, etc., by matching and correlating image features with the information in the multimodal knowledge graph, road entities (such as specific road segment units, road materials, etc.) and road edges (i.e., the relationship edges between road entities, such as spatial proximity relationships, causal dependence relationships, etc.) related to the situation reflected in the image can be found, so as to filter out the local information closely related to the current image features, namely road entities and road edges, and provide a focused data range for subsequent analysis; then, the selected road entities and road edges are combined to construct a subgraph for subsequent targeted analysis. Among them, this subgraph contains the key information related to the current road condition and their mutual relationships, which is a simplification and focus of the complex knowledge graph. For example, the constructed subgraph may include an entity of a certain road section, an entity of the material of that road section, a recent traffic load entity, and the edges representing the influence relationship between them. By constructing the subgraph, the connections between these relevant information can be presented more clearly, providing a structured framework for further exploring the potential factors related to road collapse.
[0075] Step S602: Input the subgraph into a preset gated spatio-temporal graph convolutional network to determine node features, perform graph convolution on the node features to obtain a convolution result, perform multi-hop aggregation on the convolution result to obtain target node features, and predict the first collapse feature based on the target node features.
[0076] In some embodiments, after the constructed sub-graph is input into the gated spatio-temporal graph convolutional network, the network processes the initial features of each node and updates the feature representation of the nodes through operations in spatio-temporal graph convolution, combining the topological structure of the graph and node features. At the same time, the gating mechanism adjusts the spatio-temporal information flow according to information such as environmental data features, determines which information needs to be retained or filtered, so as to determine the node features that can better reflect the true state and importance of the nodes in the spatio-temporal dimension. After that, after determining the node features, graph convolution operations are performed to weighted aggregate the features of the nodes and their neighbor nodes to extract higher-level feature representations, that is, the convolution results are obtained. Specifically, the graph convolution operations transform and fuse the node features according to the structural information of the graph (such as the connection relationship between nodes) and learnable parameters to obtain convolution features. Subsequently, since the influence between nodes may not be limited to directly connected neighbor nodes, through multi-hop aggregation, the convolution results are iteratively processed multiple times to gradually fuse the feature information of multi-order neighbor nodes, so that the finally obtained target node features can more comprehensively reflect the structure and semantic information of the entire sub-graph, covering various relevant information from local to global, providing a more accurate and complete feature basis for subsequent collapse feature prediction. Finally, the target node features obtained after the above series of operations have integrated rich spatio-temporal information, relationship information between nodes, and global information after multi-hop aggregation. After that, based on these target node features, through subsequent reasoning and calculations of the network (for example, in the hierarchical causal reasoning layer, combining physical layer reasoning and data layer reasoning, considering the stress-strain relationship of road materials and focusing on key causal edges through the attention mechanism, etc.), the first collapse feature of the road can be predicted.
[0077] In this way, by performing hierarchical causal reasoning on the multi-modal knowledge graph, the internal connections and causal relationships between different data can be quickly and deeply mined to obtain the first collapse feature.
[0078] In some embodiments, determining the second collapse feature based on the first collapse feature and multi-source environmental data includes steps S701 to S703: Step S701, obtaining the multi-source environmental data of the target road area and performing data encoding on the multi-source environmental data to obtain a second encoding result, where the second encoding result includes a radar encoding result, a temperature and humidity encoding result, and a load encoding result; In some embodiments, by obtaining multi-source environmental data such as ground penetrating radar data, temperature and humidity sensor data, and traffic load data of a target road area, then, for the ground penetrating radar data, using a 3D convolutional neural network (3DCNN) to process the dielectric constant distribution data of the ground penetrating radar to output radar features such as the volume, depth, and shape complexity of underground cavities, and then mapping the radar features to road grid nodes (1-meter resolution) and filling in missing areas through bilinear interpolation to obtain a radar coding result; for the temperature and humidity sensor data, extracting time-dependent features in the temperature and humidity sensor data through a bidirectional LSTM and identifying temperature and humidity mutation points (such as before and after heavy rain) based on the time-dependent features to generate an event marker vector, that is, determining a temperature and humidity coding result; for the traffic load data, converting the traffic load data into an equivalent single axle load (ESAL) to uniformly measure the action intensity of different vehicles on the road, and statistically calculating the ESAL mean and peak per hour and mapping them to the corresponding road grid to present the traffic load distribution in the spatio-temporal dimension and determine a load coding result.
[0079] Step S702, calculate the correlation weight between the first collapse feature and the radar coding result, and determine a radar enhanced feature based on the correlation weight and the radar coding result, and splice the temperature and humidity coding result and the load coding result to obtain an environmental time series feature; In some embodiments, due to the first collapse feature contains existing collapse-related information, calculate the first collapse feature and the radar features in the radar coding result correlation weight , the relevant formula is: , where represents the query operation on the first collapse feature for extracting key information of the feature; represents the key operation on the radar feature for extracting key information of the feature; represents the dimension of the feature, used to normalize the dot product result to prevent gradient disappearance. Then, fuse the correlation weight with the radar features in the radar coding result, and the relevant formula is: , where represents the value operation on the radar feature to obtain the specific content of the feature; through this formula, a radar enhanced feature can be determined, which can highlight the underground information that has an important impact on collapse and make subsequent analysis more focused on key factors; then, the temperature and humidity time series feature Perform data splicing with the traffic load characteristic ESAL, integrate the information of two types of environmental factors, and adjust the information flow through a gating unit. The gating unit outputs a control information fusion method to dynamically determine the contribution of each information according to the actual road conditions, so as to obtain the environmental time series characteristics more in line with reality.
[0080] Step S703: Constrain the radar enhancement feature and the environmental time series feature through the material linear elastic equation to obtain a fusion feature. Based on the first collapse feature and the fusion feature, generate corresponding dynamic weights through a cross-attention mechanism, and perform weighted aggregation on the first collapse feature and the fusion feature based on the dynamic weights to obtain a second collapse feature.
[0081] In some embodiments, first, since road collapse is related to material mechanical properties, the radar enhancement feature is constrained according to the road material constitutive equation (such as a linear elastic model) and the environmental time series feature to obtain a fusion feature ; then, based on the first collapse feature and the fusion feature , the attention mechanism will capture the mutual relationship and importance difference between features and generate dynamic weights . Subsequently, according to the dynamic weights , each feature is weighted and aggregated to obtain a second collapse feature , and the relevant formula is: . Comprehensively consider the influence of each feature on road collapse, fuse according to the importance degree, so that the second collapse feature can more comprehensively and accurately reflect potential factors and patterns. Finally, perform counterfactual intervention on the second feature (such as assuming normal temperature and humidity), calculate the difference in causal effects, and clarify the actual impact of environmental factors on road collapse by comparing the features before and after the intervention, enhancing the causal interpretability of the features and providing a more reliable basis for risk assessment and prevention.
[0082] It should be noted that the first and second do not represent the order of precedence and can be understood as nouns. Among them, the first collapse feature is obtained through hierarchical causal reasoning of a multi-modal knowledge graph, which is the summary and reasoning result of past experience and data knowledge; while the second collapse feature is determined on the basis of the first collapse feature, combined with actual dynamic multi-source environmental data, considering the impact of the current actual environmental situation on road surface collapse, such as the current rainfall intensity, real-time change of the groundwater level, etc., and can more accurately reflect the real collapse risk situation faced by the current road surface.
[0083] In this way, by comprehensively determining the second collapse feature with multi-source environmental data, considering the impact of the current actual environmental situation on road surface collapse, a second collapse feature that more comprehensively reflects the actual environmental conditions of the current road surface is obtained.
[0084] Step S104: Input the multi-scale image features and the second collapse feature into a reinforcement learning model to determine the road surface collapse risk level.
[0085] In some embodiments, the step of inputting the multi-scale image features and the second collapse feature into a reinforcement learning model to determine the road surface collapse risk level is specifically as follows: spatially align the image features and the second collapse feature and perform cross-modal feature fusion to obtain a comprehensive feature vector; use the comprehensive feature vector and the multi-source environmental data as the state space of the reinforcement learning model, and based on the state space, a preset action space, and a preset reward function, use a preset reinforcement learning model to iteratively update the policy network parameters with the goal of maximizing the expected value of the cumulative value of the reward function, determine the optimal risk assessment strategy, and determine the road surface collapse risk level based on the risk assessment strategy. Specifically, first, unify the multi-scale image features (shallow / middle / deep) and the second collapse feature (environment fusion feature after physical correction) to the same road grid coordinate system, achieve resolution alignment through bilinear interpolation, and use a cross-modal attention mechanism to calculate the correlation weights between the image features and the second collapse feature, and generate a weighted comprehensive feature vector; subsequently, map the heterogeneous data of the comprehensive feature vector and the multi-source environmental data (ground penetrating radar void distribution, temperature and humidity time series, traffic load statistics) to a unified state vector through a fully connected layer to determine the state space, and define the dynamically adjusted classification threshold (such as the high-risk probability threshold), attention weight (controlling the contribution ratio of multi-source data), and physical constraint coefficient as the state space. At the same time, considering the model accuracy and risk control requirements, define the reward function: , where is the model prediction accuracy; is the recall rate of collapse samples; is the false alarm rate; is the physical constraint consistency score (0 - 1), which measures whether the prediction result conforms to the material mechanics equation; , and are all artificially set weight coefficients. Finally, use the Proximal Policy Optimization (PPO) algorithm to iteratively update the policy network parameters with the goal of maximizing the cumulative reward , where represents the policy network with parameters , which is used to determine the probability of taking a certain action in a given state; represents the expectation; represents the process of finding the parameters that maximize the objective function (here it is the expected cumulative reward); represents the cumulative reward, from the current time step to the termination time step The reward for each time step is discounted by the discount factor, and the sum of the discounted rewards; and based on the optimal policy classify the comprehensive feature vector and output the road surface collapse risk level (low / medium / high / urgent).
[0086] In this way, by inputting the multi-scale image features and the second collapse features into the reinforcement learning model, the association and complementarity between different features can be fully exploited, the influence of these two features on the road surface collapse risk can be comprehensively considered, and a more accurate road surface collapse risk level can be generated.
[0087] In the embodiment of the present application, by collecting visible light and infrared images, the morphological features and structural features of the road surface can be intuitively obtained. Through cross-modal fusion, the advantages of these two types of information can be complementary to more accurately reflect the actual state of the road surface, facilitating the subsequent accurate and rapid extraction of features related to collapse; by inputting the road surface image into a lightweight convolutional neural network, the information of the road surface at different spatial scales can be captured; by updating the collapse factor set, the features related to road surface collapse can be accurately and quickly determined, improving the accuracy and real-time performance of monitoring; by constructing a multi-modal knowledge graph, rich multi-source information is provided for subsequent causal reasoning; through hierarchical causal reasoning on the multi-modal knowledge graph, the internal connections and causal relationships between different data can be deeply explored, and the second collapse features are determined by synthesizing multi-source environmental data, considering the influence of the current actual environmental conditions on road surface collapse, and obtaining the second collapse features that more comprehensively reflect the actual environmental conditions of the current road surface; by inputting the multi-scale image features and the second collapse features into the reinforcement learning model, the association and complementarity between different features can be fully exploited, the influence of these two features on the road surface collapse risk can be comprehensively considered, and a more accurate road surface collapse risk level can be generated. Compared with the prior art, the present application can improve the accuracy and real-time performance of road surface collapse risk monitoring.
[0088] Embodiment 2 Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of an embodiment of the road surface collapse risk monitoring system provided by the present application, including a collection module 100, an update module 200, an inference module 300, and a generation module 400; The collection module 100 is used to collect visible light images and infrared images of the target road area, and perform cross-modal fusion on the visible light images and the infrared images through a generative adversarial network constrained by a thermodynamic equation to generate a road surface image; The update module 200 is configured to input the road surface image into a lightweight convolutional neural network to extract multi-scale image features. When the multi-scale image features meet the preset conditions, the initial collapse factor set is updated based on the divergence difference of the feature distribution within a preset sliding time window to obtain a target collapse factor set. The inference module 300 is configured to construct a multi-modal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data, and perform hierarchical causal inference on the multi-modal knowledge graph through a gated spatio-temporal graph convolutional network to obtain a first collapse feature, and determine a second collapse feature based on the first collapse feature and multi-source environmental data. The generation module 400 is configured to input the multi-scale image features and the second collapse feature into a reinforcement learning model to determine the road surface collapse risk level.
[0089] Regarding the information interaction, execution process, etc. among the modules in the above road surface collapse risk monitoring system, since they are based on the same concept as the embodiments of the road surface collapse risk monitoring method in the first aspect of the present invention, the achieved technical effects are basically the same. For specific content, reference can be made to the description in Embodiment 1 of the method of the present invention, and details will not be repeated here.
[0090] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the method in this embodiment.
[0091] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the road surface collapse risk monitoring method as described in Embodiment 1 above is implemented.
[0092] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0093] The above specific embodiments have further elaborated on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application.
[0094] It is particularly pointed out that, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A method for monitoring the risk of road surface collapse, characterized in that, Including: Collect visible light images and infrared images of the target road area, and perform cross-modal fusion on the visible light images and the infrared images through a generative adversarial network constrained by a thermodynamic equation to generate a road surface image; Input the road surface image into a lightweight convolutional neural network to extract multi-scale image features. When the multi-scale image features meet preset conditions, update the initial collapse factor set based on the feature distribution divergence difference within a preset sliding time window to obtain a target collapse factor set; Construct a multi-modal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data, and perform hierarchical causal reasoning on the multi-modal knowledge graph through a gated spatio-temporal graph convolutional network to obtain a first collapse feature, and determine a second collapse feature based on the first collapse feature and multi-source environmental data; Input the multi-scale image features and the second collapse feature into a reinforcement learning model to determine the road surface collapse risk level.
2. The road surface subsidence risk monitoring method according to claim 1, characterized in that, The step of performing cross-modal fusion on the visible light images and the infrared images through a generative adversarial network constrained by a thermodynamic equation to generate a road surface image is specifically as follows: Perform image preprocessing on the visible light image and the infrared image respectively to obtain a first preprocessing result and a second preprocessing result; Stitch the first preprocessing result and the second preprocessing result to obtain a stitched image, and input the stitched image into the generative adversarial network to adaptively enhance the stitched image in combination with network parameters to generate the road surface image, where the generative adversarial network includes a generator and a discriminator, and the network parameters calculate the matching error between the infrared channel temperature field of the generated image and the Fourier heat conduction equation through the generator, and backpropagate the matching error to optimize the generated image generated by the generator. After receiving the generated image, the discriminator determines the classification result of the generated image and the stitched image, and optimizes to obtain the loss function value according to the classification result.
3. The road surface collapse risk monitoring method according to claim 2, wherein, The relevant formula for calculating the matching error between the infrared channel temperature field of the generated image and the Fourier heat conduction equation is specifically as follows: ; In the formula, is the matching error; is a preset weight coefficient for balancing the matching error and other loss terms; is the temperature value at position in the infrared channel of the generated image; is the Laplacian operator of the temperature field; is the rate of change of temperature with time; is the thermal diffusivity of the material, which characterizes the diffusion speed of heat in the material and can be dynamically adjusted according to real-time environmental temperature and humidity data.
4. The road surface subsidence risk monitoring method according to claim 1, characterized in that The step of inputting the road surface image into a lightweight convolutional neural network to extract multi-scale image features is specifically as follows: Input the road surface image into a lightweight convolutional neural network to extract an initial feature map, where the initial feature map includes a shallow feature map, a middle feature map, and a deep semantic feature map; Align the feature maps of each scale in the initial feature map and flatten them into corresponding spatial sequences, perform position encoding on each position in each pair of the spatial sequences to obtain a first encoding result corresponding to each level, calculate the feature correlation of each of the first encoding results through a multi-head attention mechanism to generate self-attention weights corresponding to each level, and obtain the multi-scale image features based on each of the self-attention weights and the initial feature map.
5. The road surface subsidence risk monitoring method according to claim 1, wherein The step of updating the initial collapse factor set based on the feature distribution divergence difference within a preset sliding time window to obtain a target collapse factor set is specifically as follows: Obtain the feature distribution of the sliding time window, and calculate the feature distribution divergence difference between the feature distribution and the historical feature distribution; If the difference in the characteristic distribution divergence is greater than a preset divergence threshold, extract the abnormal characteristics within the sliding time window, and perform incremental clustering based on the abnormal characteristics to determine candidate collapse factors. Calculate the cosine similarity between the candidate collapse factors and each collapse factor in the initial collapse factor set. If the cosine similarity is greater than the preset similarity threshold, merge the candidate collapse factor with the nearest collapse factor; otherwise, add the candidate collapse factor to the initial collapse factor set to obtain the updated target collapse factor set.
6. The road surface collapse risk monitoring method according to claim 1, characterized in that The construction of the multimodal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data is specifically as follows: Obtain the historical collapse data and corresponding historical environmental data of the target road area, and perform data preprocessing on the historical collapse data and the historical environmental data respectively to obtain the third preprocessing result and the fourth preprocessing result. Among them, the historical collapse data includes historical image data and historical text data, and the historical environmental data includes historical ground penetrating radar detection data, historical temperature and humidity sensor data, and historical traffic load time series data; Extract a number of entities from the third preprocessing result and the fourth preprocessing result, and determine the entity type corresponding to each entity and the relationship between each entity. Among them, the entity types include road structure entities, environmental entities, collapse event entities, and collapse factor entities; Map each entity to a node in the knowledge graph according to the entity type, and map the relationship to an edge in the knowledge graph to obtain a multimodal knowledge graph.
7. The road surface subsidence risk monitoring method according to claim 1, wherein The hierarchical causal reasoning of the multimodal knowledge graph through the gated spatio-temporal graph convolutional network to obtain the first collapse feature is specifically as follows: Based on the image features, query the relevant road entities and road edges in the multimodal knowledge graph, and construct a subgraph based on the road entities and the road edges; Input the subgraph into a preset gated spatio-temporal graph convolutional network to determine the node features, perform graph convolution on the node features to obtain a convolution result, perform multi-hop aggregation on the convolution result to obtain the target node features, and predict the first collapse feature based on the target node features.
8. The road surface collapse risk monitoring method according to claim 1, wherein The determination of the second collapse feature based on the first collapse feature and multi-source environmental data is specifically as follows: Obtain the multi-source environmental data of the target road area, and encode the multi-source environmental data to obtain a second encoding result. Among them, the second encoding result includes a radar encoding result, a temperature and humidity encoding result, and a load encoding result; Calculate the correlation weight between the first collapse feature and the radar encoding result, and determine the radar enhanced feature based on the correlation weight and the radar encoding result. Concatenate the temperature and humidity encoding result and the load encoding result to obtain the environmental time series feature; Constraining the radar enhancement feature and the environmental time series feature through the material linear elastic equation to obtain a fusion feature, and based on the first collapsed feature and the fusion feature, generating corresponding dynamic weights through a cross-attention mechanism, and aggregating the first collapsed feature and the fusion feature with weights based on the dynamic weights to obtain a second collapsed feature.
9. The road surface collapse risk monitoring method according to claim 1, wherein Inputting the multi-scale image feature and the second collapsed feature into a reinforcement learning model to determine the road surface collapse risk level, specifically: Performing spatial alignment and cross-modal feature fusion on the image feature and the second collapsed feature to obtain a comprehensive feature vector; Taking the comprehensive feature vector and the multi-source environmental data as the state space of the reinforcement learning model, and based on the state space, a preset action space, and a preset reward function, using a preset reinforcement learning model, iteratively updating the policy network parameters with the goal of maximizing the cumulative value expectation of the reward function, determining the optimal risk assessment strategy, and determining the road surface collapse risk level based on the risk assessment strategy.
10. A road surface subsidence risk monitoring system, characterized in that, Including: A collection module, an update module, an inference module, and a generation module; The collection module is used to collect visible light images and infrared images of the target road area, and perform cross-modal fusion on the visible light image and the infrared image through a generative adversarial network constrained by a thermodynamic equation to generate a road surface image; The update module is used to input the road surface image into a lightweight convolutional neural network to extract multi-scale image features, and when the multi-scale image features meet preset conditions, update the initial collapse factor set based on the feature distribution divergence difference within a preset sliding time window to obtain a target collapse factor set; The inference module is used to construct a multi-modal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data, and perform hierarchical causal inference on the multi-modal knowledge graph through a gated spatio-temporal graph convolutional network to obtain a first collapsed feature, and determine a second collapsed feature based on the first collapsed feature and multi-source environmental data; The generation module is used to input the multi-scale image feature and the second collapsed feature into a reinforcement learning model to determine the road surface collapse risk level.
Citation Information
Patent Citations
Complex network link prediction method and system based on logical reasoning and graph convolution
CN113190688A
Medical entity relationship extraction method based on hierarchical reasoning
CN113553440A
Road defect detection method based on deep learning
CN118918551A
Visible light and infrared light image fusion method and device for road crack detection
CN119067867A
Power backbone transmission network fault diagnosis method and system based on knowledge graph
CN119172221A
Cited By
Embedded image enhancement method and system for large-target-surface image sensor
CN120725933A
Road property state monitoring data analysis method and system applying deep learning
CN120763825A
Iron tower structure defect detection method and device
CN120976700A
High and steep slope operation state non-contact visual settlement displacement monitoring and operation state evaluation method
CN121191101A
Weather radar reflectivity synthesis method and system based on synchronous stationary satellite
CN121559504A