A road collapse risk monitoring method and system
By fusing visible light and infrared images to generate road surface images and using lightweight convolutional neural networks and multimodal knowledge graphs for hierarchical causal reasoning, the limitations of monitoring accuracy and timeliness in existing technologies are overcome, achieving more efficient road collapse risk monitoring.
Patent Information
- Application Number
- CN202510804099.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing technologies make it difficult to achieve high-frequency, full-coverage dynamic monitoring in urban road safety monitoring, resulting in the neglect of early-stage fine cracks and underground hidden risks. Existing monitoring technologies are limited in accuracy and timeliness.
Visible light and infrared images are fused to generate road surface images, and multi-scale image features are extracted through a lightweight convolutional neural network. Combined with a multimodal knowledge graph and a reinforcement learning model, hierarchical causal reasoning is performed to determine the risk level of road collapse.
It improves the accuracy and real-time performance of road collapse risk monitoring, can more accurately reflect the actual state of the road surface, quickly extract collapse-related features, and generate more accurate collapse risk levels.
Smart Images

Figure CN120339962B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of road safety monitoring, and in particular to a road subsidence risk monitoring method and system. Background Art
[0002] Roads are critical urban infrastructure, carrying the daily necessities of people and vehicles. A road collapse can instantly cause serious accidents, including vehicle falls and casualties, with devastating consequences. Furthermore, if a collapse goes undetected, it can lead to widespread traffic disruptions, hinder surrounding businesses, and significantly disrupt residents' lives due to travel inconvenience. Therefore, real-time monitoring of road collapses is crucial.
[0003] Currently, urban road safety monitoring primarily relies on single or simply combined monitoring technologies, which have significant drawbacks. While regular manual inspections can identify potential hazards based on experience, limited human resources and fixed inspection cycles prevent high-frequency, comprehensive dynamic monitoring of road conditions. This results in hidden risks such as early-stage microcracks and underground holes being easily overlooked between inspections. While single-sensor monitoring can collect some data in real time, it struggles to fully capture the complex risk signals caused by multiple factors, such as groundwater level fluctuations, soil subsidence, and pipeline leakage. Therefore, in urban environments with variable geological conditions and complex underground structures, existing monitoring technologies face limitations in accuracy and timeliness, necessitating breakthroughs through innovative technological approaches. Summary of the Invention
[0004] The present application provides a road subsidence risk monitoring method and system to improve the accuracy and real-time performance of road subsidence risk monitoring.
[0005] In a first aspect, the present application provides a method for monitoring road subsidence risk, comprising:
[0006] Collecting visible light images and infrared images of a target road area, and cross-modally fusing the visible light images and the infrared images through a generative adversarial network constrained by thermodynamic equations to generate a road surface image;
[0007] Inputting the road surface image into a lightweight convolutional neural network to extract multi-scale image features, and when the multi-scale image features meet preset conditions, updating an initial collapse factor set based on feature distribution divergence differences within a preset sliding time window to obtain a target collapse factor set;
[0008] Constructing a multimodal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data, performing hierarchical causal reasoning on the multimodal knowledge graph through a gated spatiotemporal graph convolutional network to obtain a first collapse feature, and determining a second collapse feature based on the first collapse feature and multi-source environmental data;
[0009] The multi-scale image features and the second collapse features are input into a reinforcement learning model to determine a road collapse risk level.
[0010] By collecting visible light and infrared images, the embodiments of the present application can intuitively obtain the morphological and structural characteristics of the road surface. By performing cross-modal fusion, the advantages of these two types of information can be complemented to more accurately reflect the actual state of the road surface, facilitating the subsequent accurate and rapid extraction of collapse-related features. By inputting the road surface images into a lightweight convolutional neural network, information about the road surface at different spatial scales can be captured. By updating the collapse factor set, features related to road collapse can be accurately and quickly determined, improving the accuracy and real-time performance of monitoring. By constructing a multimodal knowledge graph, rich multi-source information is provided for subsequent causal reasoning. By performing hierarchical causal reasoning on the multimodal knowledge graph, the intrinsic connections and causal relationships between different data can be deeply explored. By integrating multi-source environmental data to determine the second collapse feature, the impact of the current actual environmental conditions on road collapse is considered, and a second collapse feature that more comprehensively reflects the actual environmental conditions of the current road surface is obtained. By inputting the multi-scale image features and the second collapse features into the reinforcement learning model, the correlation and complementarity between different features can be fully explored, and the impact of these two features on road collapse risk can be comprehensively considered to generate a more accurate road collapse risk level. Compared with the existing technology, this application can improve the accuracy and real-time performance of road subsidence risk monitoring.
[0011] Furthermore, the generative adversarial network constrained by thermodynamic equations performs cross-modal fusion of the visible light image and the infrared image to generate a road surface image, specifically:
[0012] performing image preprocessing on the visible light image and the infrared image respectively to obtain a first preprocessing result and a second preprocessing result;
[0013] The first preprocessing result and the second preprocessing result are spliced together to obtain a spliced image, and the spliced image is input into the generative adversarial network to adaptively enhance the spliced image in combination with network parameters to generate the road surface image, wherein the generative adversarial network includes a generator and a discriminator. The network parameters are calculated by the generator to obtain a matching error between the infrared channel temperature field of the generated image and the Fourier heat conduction equation, and the matching error is back-propagated to optimize the generated image generated by the generator. After receiving the generated image, the discriminator determines the classification results of the generated image and the spliced image, and optimizes the result after determining the loss function value based on the classification result.
[0014] In this way, through cross-modal fusion, the advantages of these two types of information can be complemented to more accurately reflect the actual state of the road surface, making it easier to extract features related to collapse accurately and quickly.
[0015] Furthermore, the formula for the matching error between the infrared channel temperature field of the calculated image and the Fourier heat conduction equation is specifically:
[0016] ;
[0017] Where, is the matching error; is a preset weight coefficient used to balance the matching error and other loss terms; To generate the position in the infrared channel of the image The temperature value at is the Laplace operator of the temperature field; is the rate of change of temperature with time; The thermal diffusivity of the material characterizes the diffusion rate of heat in the material and can be dynamically adjusted according to real-time ambient temperature and humidity data.
[0018] Furthermore, the road surface image is input into a lightweight convolutional neural network to extract multi-scale image features, specifically:
[0019] Inputting the road surface image into a lightweight convolutional neural network to extract an initial feature map, wherein the initial feature map includes a shallow feature map, a middle feature map, and a deep semantic feature map;
[0020] The scale feature maps in the initial feature map are aligned and flattened into corresponding spatial sequences, and position encoding is performed on each position in the spatial sequences to obtain a first encoding result corresponding to each level, and the feature correlation of each first encoding result is calculated through a multi-head attention mechanism to generate a self-attention weight corresponding to each level, and the multi-scale image feature is obtained based on each self-attention weight and the initial feature map.
[0021] In this way, by inputting the road surface image into a lightweight convolutional neural network, information about the road surface at different spatial scales can be captured.
[0022] Furthermore, the initial collapse factor set is updated based on the feature distribution divergence difference within the preset sliding time window to obtain the target collapse factor set, specifically:
[0023] Obtaining the feature distribution of the sliding time window, and calculating the feature distribution divergence difference between the feature distribution and the historical feature distribution;
[0024] If the feature distribution divergence difference is greater than a preset difference threshold, the abnormal features within the sliding time window are extracted, and incremental clustering is performed based on the abnormal features to determine candidate collapse factors. The cosine similarity between the candidate collapse factors and each collapse factor in the initial collapse factor set is calculated. If the cosine similarity is greater than a preset similarity threshold, the candidate collapse factor is merged with the nearest collapse factor; otherwise, the candidate collapse factor is added to the initial collapse factor set to obtain the updated target collapse factor set.
[0025] In this way, by updating the collapse factor set, the characteristics related to road collapse can be accurately and quickly determined, improving the accuracy and real-time performance of monitoring.
[0026] Furthermore, the multimodal knowledge graph is constructed based on the collapse factor set, historical collapse data, and historical environmental data, specifically:
[0027] Acquire historical collapse data and corresponding historical environmental data of a target road area, and perform data preprocessing on the historical collapse data and the historical environmental data, respectively, to obtain a third preprocessing result and a fourth preprocessing result, wherein the historical collapse data includes historical image data and historical text data, and the historical environmental data includes historical geological radar detection data, historical temperature and humidity sensor data, and historical traffic load time series data;
[0028] Extracting a plurality of entities from the third preprocessing result and the fourth preprocessing result, and determining entity types corresponding to the entities and relationships between the entities, wherein the entity types include road structure entities, environment entities, collapse event entities, and collapse factor entities;
[0029] According to the entity type, each entity is mapped to a node in the knowledge graph, and the relationship is mapped to an edge in the knowledge graph to obtain a multimodal knowledge graph.
[0030] In this way, by constructing a multimodal knowledge graph, rich multi-source information is provided for subsequent causal reasoning, which facilitates in-depth exploration of the intrinsic connections and causal relationships between different data to accurately determine the first collapse feature.
[0031] Furthermore, the multimodal knowledge graph is subjected to hierarchical causal reasoning through a gated spatiotemporal graph convolutional network to obtain a first collapsed feature, specifically:
[0032] Based on the image features, querying the multimodal knowledge graph for relevant road entities and road edges, and constructing a subgraph based on the road entities and the road edges;
[0033] The subgraph is input into a preset gated spatiotemporal graph convolutional network to determine node features, and graph convolution is performed on the node features to obtain convolution results, and multi-hop aggregation is performed on the convolution results to obtain target node features, and the first collapsed features are predicted based on the target node features.
[0034] In this way, by performing hierarchical causal reasoning on the multimodal knowledge graph, we can quickly and deeply explore the intrinsic connections and causal relationships between different data and obtain the first collapsed feature.
[0035] Furthermore, the determining of the second collapse feature based on the first collapse feature and multi-source environmental data is specifically:
[0036] Acquire the multi-source environmental data of the target road area, and encode the multi-source environmental data to obtain a second encoding result, wherein the second encoding result includes a radar encoding result, a temperature and humidity encoding result, and a load encoding result;
[0037] Calculating an association weight between the first collapse feature and the radar coding result, determining a radar enhancement feature based on the association weight and the radar coding result, and performing data splicing on the temperature and humidity coding result and the load coding result to obtain an environmental time series feature;
[0038] The radar enhancement feature and the environmental time series feature are constrained by the material linear elastic equation to obtain a fusion feature, and based on the first collapsed feature and the fusion feature, a corresponding dynamic weight is generated through a cross-attention mechanism. The first collapsed feature and the fusion feature are weightedly aggregated based on the dynamic weight to obtain a second collapsed feature.
[0039] In this way, the second collapse feature is determined by integrating multi-source environmental data, taking into account the impact of the current actual environmental conditions on road collapse, and obtaining a second collapse feature that more comprehensively reflects the actual environmental conditions of the current road surface.
[0040] Furthermore, the multi-scale image features and the second collapse features are input into a reinforcement learning model to determine the road collapse risk level, specifically:
[0041] spatially aligning and cross-modal fusing the image features and the second collapsed features to obtain a comprehensive feature vector;
[0042] The comprehensive feature vector and the multi-source environmental data are used as the state space of the reinforcement learning model. Based on the state space, the preset action space and the preset reward function, the preset reinforcement learning model is used to iteratively update the strategy network parameters with the goal of maximizing the expected cumulative value of the reward function, determine the optimal risk assessment strategy, and determine the road collapse risk level based on the risk assessment strategy.
[0043] In this way, by inputting the multi-scale image features and the second collapse features into the reinforcement learning model, the correlation and complementarity between different features can be fully explored, and the impact of these two features on the road collapse risk can be comprehensively considered to generate a more accurate road collapse risk level.
[0044] In a second aspect, the present application provides a road subsidence risk monitoring system, comprising: an acquisition module, an update module, an inference module, and a generation module;
[0045] The acquisition module is configured to acquire a visible light image and an infrared image of a target road area, and perform cross-modal fusion of the visible light image and the infrared image through a generative adversarial network constrained by thermodynamic equations to generate a road surface image;
[0046] The updating module is configured to input the road surface image into a lightweight convolutional neural network to extract multi-scale image features, and when the multi-scale image features meet preset conditions, update the initial collapse factor set based on the feature distribution divergence difference within a preset sliding time window to obtain a target collapse factor set;
[0047] The reasoning module is used to construct a multimodal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data, and perform hierarchical causal reasoning on the multimodal knowledge graph through a gated spatiotemporal graph convolutional network to obtain a first collapse feature, and determine a second collapse feature based on the first collapse feature and multi-source environmental data;
[0048] The generation module is used to input the multi-scale image features and the second collapse features into a reinforcement learning model to determine the road collapse risk level.
[0049] By collecting visible light and infrared images, the embodiments of the present application can intuitively obtain the morphological and structural characteristics of the road surface. By performing cross-modal fusion, the advantages of these two types of information can be complemented to more accurately reflect the actual state of the road surface, facilitating the subsequent accurate and rapid extraction of collapse-related features. By inputting the road surface images into a lightweight convolutional neural network, information about the road surface at different spatial scales can be captured. By updating the collapse factor set, features related to road collapse can be accurately and quickly determined, improving the accuracy and real-time performance of monitoring. By constructing a multimodal knowledge graph, rich multi-source information is provided for subsequent causal reasoning. By performing hierarchical causal reasoning on the multimodal knowledge graph, the intrinsic connections and causal relationships between different data can be deeply explored. By integrating multi-source environmental data to determine the second collapse feature, the impact of the current actual environmental conditions on road collapse is considered, and a second collapse feature that more comprehensively reflects the actual environmental conditions of the current road surface is obtained. By inputting the multi-scale image features and the second collapse features into the reinforcement learning model, the correlation and complementarity between different features can be fully explored, and the impact of these two features on road collapse risk can be comprehensively considered to generate a more accurate road collapse risk level. Compared with the existing technology, this application can improve the accuracy and real-time performance of road subsidence risk monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 This is a flow chart of an embodiment of a road subsidence risk monitoring method provided by the present application;
[0051] Figure 2 It is a structural diagram of an embodiment of the road subsidence risk monitoring system provided in this application. DETAILED DESCRIPTION
[0052] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0053] It should be understood that the step numbers used herein are only for convenience of description and are not intended to limit the order in which the steps are executed.
[0054] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0055] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0056] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items.
[0057] Roads are critical urban infrastructure, carrying daily traffic. Road collapses can cause serious accidents, resulting in casualties, traffic disruptions, commercial disruptions, and inconvenience for residents, making real-time monitoring crucial. Current urban road safety monitoring methods often rely on a single or a combination of technologies, which presents drawbacks. Manual inspections are limited by manpower and time constraints, making dynamic monitoring difficult and prone to overlooking early-stage hazards. Single sensors provide incomplete data and struggle to capture complex risk signals. In urban environments with diverse geology and complex underground structures, the accuracy and timeliness of existing technologies urgently require innovative breakthroughs.
[0058] Next, the nouns involved in this application are analyzed:
[0059] Thermodynamic equation constraints refer to the requirement that in certain physical processes or models, thermodynamic-related laws (such as the law of conservation of energy, the second law of thermodynamics, etc.) or their corresponding mathematical equations be satisfied to limit and regulate the construction of the process or model.
[0060] Generative adversarial networks consist of two parts: a generator and a discriminator. The generator is responsible for generating fake data that is as close to real data as possible, while the discriminator is responsible for distinguishing whether the data is real or generated by the generator. These two parts compete with each other during training. The generator continuously learns to generate more realistic data to deceive the discriminator, while the discriminator continuously learns to more accurately identify the authenticity of the data, thus achieving joint optimization of the generator and discriminator.
[0061] A multimodal knowledge graph integrates different types of data (such as text, images, audio, and video) into a unified knowledge graph to represent the complex relationships and semantic information between entities. By fusing and linking multimodal data, it constructs a more comprehensive and rich knowledge representation framework, enabling more accurate descriptions of various objects and phenomena in the real world.
[0062] Based on this, the embodiments of the present application provide a road subsidence risk monitoring method and system, which can improve the accuracy and real-time performance of road subsidence risk monitoring.
[0063] A road subsidence risk monitoring method and system provided in an embodiment of the present application are specifically illustrated through the following embodiments. First, the road subsidence risk monitoring method in an embodiment of the present application is described.
[0064] The road subsidence risk monitoring method provided in the embodiment of the present application relates to the field of road safety monitoring. The road subsidence risk monitoring method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements a road subsidence risk monitoring method, etc., but is not limited to the above forms.
[0065] The present application can also be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0066] Example 1
[0067] Please refer to Figure 1 , Figure 1 This is a flow chart of an embodiment of the road subsidence risk monitoring method provided by the present application, including steps S101 to S104;
[0068] Step S101: collecting a visible light image and an infrared image of a target road area, and cross-modally fusing the visible light image and the infrared image through a generative adversarial network constrained by thermodynamic equations to generate a road surface image;
[0069] In some embodiments, visible light images and infrared images of the target road area are collected, specifically by mounting a visible light camera and an infrared thermal imager on a suitable carrier, such as a road inspection vehicle, to ensure that they can be aimed at the target road area at the same time and at the same angle. During installation, the position and angle of the equipment must be accurately adjusted so that their fields of view basically overlap. Thereafter, the parameters of the visible light camera and the infrared thermal imager are set so that the road inspection vehicle travels along a preset route and speed to collect visible light images and infrared images of the target road area.
[0070] It should be noted that visible light images can show the texture of road cracks, and infrared images can show abnormal temperature distribution caused by rainwater evaporation. By fusing the two, the crack locations can be accurately matched with temperature abnormality areas, and the development trend of road diseases can be evaluated, facilitating subsequent accurate prediction of road collapse risks.
[0071] In some embodiments, the generative adversarial network constrained by thermodynamic equations performs cross-modal fusion of the visible light image and the infrared image to generate a road surface image, including steps S201 to S202:
[0072] Step S201, performing image preprocessing on the visible light image and the infrared image respectively to obtain a first preprocessing result and a second preprocessing result;
[0073] In some embodiments, first, the spatial coordinates of the infrared image and the visible light image are aligned through affine transformation; then, the visible light image may be subjected to, but is not limited to, image preprocessing operations such as resolution normalization, grayscale normalization, and optical distortion correction to obtain a first preprocessing result. The specific image preprocessing operations are not the focus of this application and will not be expanded here; then, the ambient temperature and humidity data are obtained in real time through the temperature and humidity sensor, and the emissivity of the road material is determined by consulting relevant information or through experiments. , to determine the ambient emissivity , the original temperature value of each pixel in the infrared image Calculate and use the formula , calculate the corrected true temperature value , thereby completing the temperature calibration to obtain the second preprocessing result.
[0074] Step S202: splice the first preprocessing result and the second preprocessing result to obtain a spliced image, and input the spliced image into the generative adversarial network to adaptively enhance the spliced image in combination with network parameters to generate the road surface image, wherein the generative adversarial network includes a generator and a discriminator. The network parameters are calculated by the generator to obtain a matching error between the infrared channel temperature field of the generated image and the Fourier heat conduction equation, and the matching error is back-propagated to optimize the generated image generated by the generator. After receiving the generated image, the discriminator determines the classification results of the generated image and the spliced image, and optimizes the result after determining the loss function value based on the classification result.
[0075] In some embodiments, first, the first preprocessing result and the second preprocessing result are channel-stitched to obtain a stitched image. For example, since visible light images usually have 3 channels (RGB) and infrared images have 1 channel (temperature information), a 4-channel stitched image is obtained after stitching; then, the stitched image is input into the generative adversarial network, and the generator will generate an image based on the generator parameter output, thereby obtaining an adaptively enhanced road surface image.
[0076] It should be noted that the network parameters include generator parameters and discriminator parameters. The generator parameters are obtained by extracting the temperature field of the infrared channel of the generated image generated by the generator. , and calculate the temperature field The matching error with the Fourier heat conduction equation is calculated, and the calculated matching error is back-propagated to the generator to adjust the generator parameters until the preset conditions are met. The generator can then obtain a generated image that preliminarily conforms to the heat conduction law based on the generator parameters. The discriminator determines the classification results of the generated image and the spliced image (determines whether it is a real image or a generated image), and based on the classification results, minimizes the adversarial loss function. The principle is: , where express expectations; represents the discriminator; Represents a generator; represents the noise or initial image input to the generator; Represents the judgment result of the discriminator on the image generated by the generator, indicating the probability that the image is a real image; Indicates the gap between the discriminator's judgment result and the real image label (usually 1). After that, it is necessary to calculate the pixel-level loss and perceptual loss , thus the total loss function can be determined according to the preset weights , by adjusting the weight coefficients of different loss terms 、 and , which can be used to balance the total loss function By combining the contributions of each item (adversarial loss, matching error, pixel-level loss, and perceptual loss), the total loss value can be minimized to optimize the network parameters of the generative adversarial network. Only then can the optimal generator parameters and discriminator parameters be determined.
[0077] It should be noted that in order to improve the generalization ability and robustness of the generative adversarial network, the algorithm can be used to simulate visible light and infrared images in different environments such as rain, fog, and night to increase the diversity of training data. At the same time, Gaussian noise (variance = 0.1) can be randomly added to the infrared image to simulate the noise interference in actual measurements.
[0078] In some embodiments, the formula for the matching error between the infrared channel temperature field of the calculated image and the Fourier heat conduction equation is specifically:
[0079] ;
[0080] Where, is the matching error; is a preset weight coefficient used to balance the matching error and other loss terms; To generate the position in the infrared channel of the image The temperature value at is the Laplace operator of the temperature field; is the rate of change of temperature with time; The thermal diffusivity of the material characterizes the diffusion rate of heat in the material and can be dynamically adjusted according to real-time ambient temperature and humidity data.
[0081] It should be noted that the thermal diffusivity It is dynamically adjusted according to the real-time ambient temperature and humidity data, where the adjustment formula is: , where is the base thermal diffusivity; and is the environmental factor; and are the temperature and humidity of the current environment respectively; and are the reference temperature and reference humidity respectively.
[0082] It should be noted that the generator uses the U-Net structure as the backbone network, which contains 5 layers of downsampling (implemented by convolutional layers and LeakyReLU activation functions) and 5 layers of upsampling (implemented by deconvolution and jump connections), and adds a physical constraint module at the end of the decoder, namely the thermodynamic consistency branch.
[0083] In this way, through cross-modal fusion, the advantages of these two types of information can be complemented to more accurately reflect the actual state of the road surface, making it easier to extract features related to collapse accurately and quickly.
[0084] Step S102: Input the road surface image into a lightweight convolutional neural network to extract multi-scale image features. When the multi-scale image features meet preset conditions, the initial collapse factor set is updated based on the feature distribution divergence difference within a preset sliding time window to obtain a target collapse factor set.
[0085] In some embodiments, inputting the road surface image into a lightweight convolutional neural network to extract multi-scale image features includes steps S301 to S302;
[0086] Step S301: input the road surface image into a lightweight convolutional neural network to extract an initial feature map, wherein the initial feature map includes a shallow feature map, a middle feature map, and a deep semantic feature map;
[0087] In some embodiments, before the road surface image is input into a lightweight convolutional neural network, the spliced image needs to be standardized to adapt to the network input requirements. Afterwards, the road surface image is input into a lightweight convolutional neural network (such as MobileNetV3-Small), and low-level texture features (such as edges and crack contours) are extracted through the initial convolution layer. A shallow feature map is output, and local details and channel importance information are fused through depthwise separable convolution and a squeeze excitation module (SE Block) to generate a mid-level feature map. After multiple layers of convolution and downsampling operations, a low-resolution, high-semantic feature map (such as 14×14×512) is output to characterize the global road structure and potential abnormal areas.
[0088] Step S302: align the scale feature maps in the initial feature map and flatten them into corresponding spatial sequences, and perform position encoding on each position in the spatial sequence to obtain a first encoding result corresponding to each level, and calculate the feature correlation of each first encoding result through a multi-head attention mechanism to generate a self-attention weight corresponding to each level, and obtain the multi-scale image feature based on each self-attention weight and the initial feature map.
[0089] In some embodiments, the feature maps of each scale in the initial feature map are aligned and flattened into corresponding spatial sequences. Specifically, since the initial feature map contains feature maps of different scales (such as shallow, middle, and deep feature maps), and these feature maps differ in size (resolution) and number of channels, in order to be able to be processed uniformly later, they need to be aligned in the spatial dimension first, such as by upsampling or downsampling operations to make them have the same size specifications. After alignment, the feature map of each scale is flattened from a three-dimensional tensor (height × width × number of channels) to a two-dimensional spatial sequence (number of pixels × number of channels) to obtain a shallow feature sequence, a middle feature sequence, and a deep semantic feature sequence, so as to facilitate the operation and analysis of the features at each position. For example, the original middle feature map is 56×56×128, which becomes (56×56)×128 after flattening. Each position represents a pixel point and its corresponding channel feature information.
[0090] In some embodiments, position encoding is performed on each position in each pair of the spatial sequences to obtain a first encoding result corresponding to each level. Specifically, since each position is only a numerical vector after the feature map is flattened into a spatial sequence, the original spatial position information in the image is lost. Therefore, spatial coordinate information is embedded in each position in the spatial sequence (shallow feature sequence, middle feature sequence and deep semantic feature sequence) through learnable 2D sinusoidal coding and other methods to obtain the first encoding result corresponding to each level (including shallow encoding results, middle encoding results and deep encoding results) to facilitate subsequent understanding of image structure and feature correlation.
[0091] In some embodiments, the feature correlation of each of the first encoding results is calculated through a multi-head attention mechanism to generate the self-attention weights corresponding to each level, specifically: first, the first encoding result (i.e., the feature sequence, including shallow encoding results, middle encoding results, and deep encoding results) is divided into multiple attention heads, and each head independently learns the information of different feature subspaces to learn the correlation between features from multiple different angles; then, for each attention head, the feature sequence is mapped to Query (query vector), Key (key vector), and Value (value vector) respectively by linear projection; then, the transpose of Query (query vector) and Key (key vector) is used for matrix multiplication, and then divided by a scaling factor ,in, The dimension of each attention head in the attention mechanism is calculated, and the calculation result is converted into a probability distribution through the Softmax function to obtain the self-attention weight. This weight represents the degree of correlation between the feature of the current position and the features of all other positions. The higher the correlation, the greater the weight.
[0092] In some embodiments, based on each of the self-attention weights and the initial feature map, the multi-scale image features are obtained, specifically: the self-attention weights calculated by each attention head are weightedly aggregated on the Value (value vector), that is, the eigenvalues of each position in the Value are weightedly summed according to the degree of association represented by the self-attention weight. For example, if the weight of a certain position in the self-attention weight is larger, then the eigenvalue of the corresponding position in the Value contributes more during weighted aggregation. In this way, the feature representation of each attention head after weighted aggregation is obtained, and then the feature representations of all attention heads after weighted aggregation are spliced to obtain a feature representation that integrates multi-head attention information. Finally, the feature representation that integrates multi-head attention information is added to the initial feature map to obtain multi-scale image features.
[0093] It should be noted that Query is used to find information related to itself, Key is used to represent the information available for query, and Value is the actual feature value to be weighted and aggregated.
[0094] In this way, by inputting the road surface image into a lightweight convolutional neural network, information about the road surface at different spatial scales can be captured.
[0095] In some embodiments, when the multi-scale image feature meets the preset conditions, specifically, the local variance of the multi-scale image feature in the sliding window is calculated, and the local variance is compared with a dynamic threshold, wherein the dynamic threshold is determined based on the variance of the distribution of historical normal data features. If the local variance of the multi-scale image feature is greater than the dynamic threshold, and the number of times it exceeds the dynamic threshold is greater than 3 times, it is determined to be an abnormal state and the collapse factor set needs to be updated.
[0096] In some embodiments, the updating of the initial collapse factor set based on the feature distribution divergence difference within the preset sliding time window to obtain the target collapse factor set includes steps S401 to S402:
[0097] Step S401, obtaining the feature distribution of the sliding time window, and calculating the feature distribution divergence difference between the feature distribution and the historical feature distribution;
[0098] In some embodiments, by setting a sliding time window (e.g., a window length of 30 frames, a time span of 30 seconds, and a sliding step of 5 frames), the image feature vectors within the window are cached in real time to capture the dynamic changes of image features in the time dimension. As time goes by, the window continues to slide and continuously collects new feature vectors, thereby obtaining the distribution of image features within the current window, that is, determining the feature distribution. For example, when monitoring the risk of road collapse, road images at different time points will contain feature information such as crack expansion and vehicle load. The sliding window can aggregate these time-varying features for analysis. Subsequently, the Wasserstein distance is used to measure the difference between the current sliding window feature distribution P and the historical normal distribution Q, that is, the feature distribution divergence difference. The relevant formula is: , where is the Wasserstein distance; P is the current sliding window feature distribution; Q is the historical normal distribution; Belongs to all possible joint distributions , used to describe how to combine the sliding window feature distribution P and the historical normal distribution Q; 、 It is a randomly selected sample; In the joint distribution The expected value of the Euclidean distance between a sample x drawn from distribution P and a sample y drawn from distribution Q is obtained by the Sinkhorn algorithm. The Sinkhorn algorithm is then used to approximate the solution, quantifying the degree of deviation of the current characteristic distribution P from the historical normal state Q, to obtain the characteristic distribution divergence difference. If the current road surface has an abnormal condition (such as a sudden increase in cracks or an abnormal change in the groundwater level), then its characteristic distribution P will differ significantly from the historical normal distribution Q, and the Wasserstein distance will increase.
[0099] Step S402: If the feature distribution divergence difference is greater than a preset difference threshold, the abnormal features within the sliding time window are extracted, and incremental clustering is performed based on the abnormal features to determine candidate collapse factors. The cosine similarity between the candidate collapse factors and each collapse factor in the initial collapse factor set is calculated. If the cosine similarity is greater than a preset similarity threshold, the candidate collapse factor is merged with the nearest collapse factor; otherwise, the candidate collapse factor is added to the initial collapse factor set to obtain the updated target collapse factor set.
[0100] In some embodiments, if the feature distribution divergence difference is greater than a preset difference threshold, the abnormal features within the sliding time window are extracted, specifically: when the calculated Wasserstein distance Greater than the preset difference threshold When (which can be understood as the preset Wasserstein distance threshold, used to determine whether the feature distribution is significantly abnormal), it means that the feature distribution in the current sliding window is significantly different from the historical normal distribution, that is, the current feature distribution has shifted significantly, that is, some abnormal conditions have occurred. At this time, it is necessary to screen out the abnormal features that cause the distribution shift from the many features in the current window. For example, in the pavement collapse risk monitoring scenario, the original pavement crack width fluctuates within a certain range. If a certain monitoring finds that the crack width exceeds the historical normal range, this crack width feature that exceeds the range is an abnormal feature and needs to be extracted to further analyze the potential relationship between these abnormal features and pavement collapse.
[0101] In some embodiments, incremental clustering is performed based on the abnormal features to determine candidate collapse factors. Specifically, after the abnormal features are extracted, in order to more systematically analyze these abnormal features, an incremental clustering algorithm (such as the StreamKM++ algorithm) is used to cluster these abnormal features so as to classify similar abnormal features into one category. Each category represents a pattern or factor that may be related to road collapse. These clustering results form candidate collapse factors. For example, similar abnormal features such as crack expansion anomaly and local settlement anomaly are grouped into one category. The comprehensive features represented by this category may be a candidate collapse factor. , which reflects a potential combination of factors that lead to road collapse.
[0102] It should be noted that incremental clustering can dynamically adjust the clustering results while continuously receiving new data (i.e., new abnormal features).
[0103] In some embodiments, the cosine similarity between the candidate collapse factor and each collapse factor in the initial collapse factor set is calculated, specifically: since the initial collapse factor set It is a set of factors related to road collapse that have been previously determined (can be determined by analyzing historical collapse events in the early stage). At this time, the candidate collapse factors are calculated. and each collapse factor in the initial collapse factor set The cosine similarity of the candidate collapse factor With the existing collapse factor The similarity between the candidate collapse factors With an existing collapse factor Similar characteristics suggest that they may represent similar collapse risk factors or patterns. For example, the newly discovered candidate collapse factor, "rapid crack expansion on sections with recent heavy vehicle traffic," may have a high cosine similarity to "heavy vehicle loads causing pavement structural damage" in the initial collapse factor set, as both are related to the impact of heavy vehicle loads on the pavement.
[0104] In some embodiments, if the cosine similarity is greater than a preset similarity threshold, the candidate collapse factor is merged with the nearest collapse factor, otherwise the candidate collapse factor is added to the initial collapse factor set to obtain the updated target collapse factor set, specifically: when the candidate collapse factor and a collapse factor in the initial collapse factor set When the cosine similarity of is greater than the preset similarity threshold (such as 0.9), it means that they are very similar in features and the collapse risk patterns they represent. At this time, the candidate collapse factor is compared with the collapse factor with the closest neighbor (highest similarity). Merge, where the merging method usually adopts weighted average and other methods to integrate the feature information of the two to update the representation of the existing collapse factor so that it can more comprehensively reflect the relevant collapse risk factors. If the cosine similarity does not meet the preset threshold, it means that the candidate collapse factor Represents a new, different collapse factor from the existing Different collapse risk factors or patterns are added to the initial collapse factor set In this way, the content of the collapse factor set is enriched, so that it can cover more factors that may cause road collapse, and provide a more comprehensive basis for subsequent more accurate assessment of road collapse risks. After the above processing (merging or adding) of the candidate collapse factors, the initial collapse factor set is updated to form the target collapse factor set as shown in the figure. This updated set includes both historically determined stable collapse patterns (i.e., the original factors in the initial collapse factor set) and newly discovered collapse risk factors that reflect current environmental changes (i.e., newly added or merged candidate collapse factors). This target collapse factor set can better adapt to the needs of road collapse risk assessment in different times and environments, providing more accurate and comprehensive basic data for subsequent risk assessments based on these collapse factors.
[0105] It should be noted that both the preset difference threshold and the preset similarity threshold can be set in advance based on the historical data analysis results, and this application does not impose any restrictions.
[0106] It should be noted that if there is a requirement for the number of collapse factor sets, they can be sorted by the most recent update time to eliminate the oldest or least active factors.
[0107] In this way, by updating the collapse factor set, the characteristics related to road collapse can be accurately and quickly determined, improving the accuracy and real-time performance of monitoring.
[0108] Step S103: constructing a multimodal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data, performing hierarchical causal reasoning on the multimodal knowledge graph through a gated spatiotemporal graph convolutional network to obtain a first collapse feature, and determining a second collapse feature based on the first collapse feature and multi-source environmental data;
[0109] In some embodiments, the multimodal knowledge graph is constructed based on the collapse factor set, historical collapse data, and historical environment data, including steps S501 to S503:
[0110] Step S501: Acquire historical collapse data and corresponding historical environmental data of a target road area, and perform data preprocessing on the historical collapse data and the historical environmental data, respectively, to obtain a third preprocessing result and a fourth preprocessing result, wherein the historical collapse data includes historical image data and historical text data, and the historical environmental data includes historical geological radar detection data, historical temperature and humidity sensor data, and historical traffic load time series data;
[0111] In some embodiments, since the historical collapse data and historical environmental data of the target road area are the basis for constructing a multimodal knowledge graph, the historical collapse data can reflect the actual situation of the collapse in the area in the past, such as location, time, degree of collapse, etc., which can be obtained by retrieving archives; and the historical environmental data covers the historical information of various environmental factors related to the road, including geological conditions, temperature and humidity changes, traffic load conditions, etc., which can be obtained through sensors and other means. By obtaining these data, it is convenient to fully understand the possible relationship between road collapse and environmental factors in the future, and provide data support for subsequent analysis and map construction. For example, by analyzing historical collapse data, you can know which sections of road are more likely to collapse, and combined with historical environmental data, you can explore which environmental factors caused these collapse events. Afterwards, since the original historical collapse data and historical environmental data may have problems such as noise, missing values, and inconsistent formats, direct use will affect the accuracy and reliability of subsequent analysis. Therefore, they need to be preprocessed. Among them, for historical collapse data (including historical image data and historical text data), the images need to be denoised and enhanced, and the text data needs to be cleaned and format normalized. For historical environmental data (including historical geological radar detection data, historical temperature and humidity sensor data, and historical traffic load time series data), it may be necessary to deal with missing values and outliers, and perform data standardization. Through these preprocessing operations, relatively clean and standardized data are obtained, namely the third preprocessing results and the fourth preprocessing results, so that valuable information can be extracted from them more effectively.
[0112] Step S502: extracting a plurality of entities from the third preprocessing result and the fourth preprocessing result, and determining entity types corresponding to the respective entities and relationships between the respective entities, wherein the entity types include road structure entities, environment entities, collapse event entities, and collapse factor entities;
[0113] In some embodiments, by analyzing data and applying relevant algorithms (such as causal discovery algorithms), entities are mined from the third preprocessing results and the fourth preprocessing results, and the entities are classified to determine the entity types and the relationships between the entities. Determining the entity types (such as road structure entities, environmental entities, collapse event entities, and collapse factor entities) helps to classify, manage, and understand the entities. Different types of entities have different roles and attributes in the knowledge graph. At the same time, the relationships between entities reflect their mutual connections in the real world. For example, there may be an influence relationship between road structure entities (such as road segment units) and environmental entities (such as temperature and humidity), and there may be a causal relationship between collapse event entities (such as a certain collapse) and collapse factor entities (such as crack density). Therefore, by mining the entity types and the intrinsic connections between entities, relationship information can be provided for the construction of the knowledge graph.
[0114] Step S503: Map each entity to a node in the knowledge graph according to the entity type, and map the relationship to an edge in the knowledge graph to obtain a multimodal knowledge graph.
[0115] In some embodiments, since the knowledge graph represents knowledge in the form of nodes and edges, where entities are the carriers of knowledge and different types of entities correspond to different types of nodes, corresponding entities are mapped to different node types in the knowledge graph according to the entity type. For example, road structure entities are mapped to physical topology nodes, environmental entities are mapped to environmental state nodes, collapse event entities are mapped to historical collapse nodes, and collapse factor entities are mapped to collapse factor nodes. In this way, the nodes in the knowledge graph represent various entities in the real world, laying the foundation for the subsequent expression of relationships between entities in the graph. In addition, in the knowledge graph, edges are used to represent relationships between nodes (entities). Subsequently, the relationships between previously determined entities (such as spatial proximity edges, causal dependency edges, dynamic interaction edges, etc.) are mapped to edges in the knowledge graph, establishing connections between nodes, allowing the knowledge graph to fully express the association relationships between entities. In this way, the information in historical collapse data and historical environmental data is presented in a structured form, forming a multimodal knowledge graph. This graph integrates various types of data and relationships, and can represent knowledge related to road collapse from multiple dimensions, providing a powerful tool for subsequent use of knowledge graphs for road collapse analysis, risk assessment, etc.
[0116] It should be noted that after the multimodal knowledge graph is determined, it needs to be updated regularly. The causal discovery algorithm can be rerun based on the latest multi-source environmental data for updating, and the environmental status node attributes (temperature, humidity, traffic load) can be updated in real time.
[0117] In this way, by constructing a multimodal knowledge graph, rich multi-source information is provided for subsequent causal reasoning, which facilitates in-depth exploration of the intrinsic connections and causal relationships between different data to accurately determine the first collapse feature.
[0118] In some embodiments, performing hierarchical causal reasoning on the multimodal knowledge graph through a gated spatiotemporal graph convolutional network to obtain a first collapsed feature includes steps S601 to S602;
[0119] Step S601: Based on the image features, relevant road entities and road edges are searched in the multimodal knowledge graph, and a subgraph is constructed based on the road entities and the road edges.
[0120] In some embodiments, first, because the multimodal knowledge graph integrates a variety of road information, including road structure, environmental factors, and historical collapse conditions, by matching and associating image features with information in the multimodal knowledge graph, it is possible to find road entities (such as specific road segment units, road materials, etc.) and road edges (i.e., relationship edges between road entities, such as spatial proximity and causal dependencies) related to the situation reflected in the image. This allows the selection of local information closely related to the current image features, namely road entities and road edges, to provide a focused data range for subsequent analysis. The selected road entities and road edges are then combined to construct a subgraph for subsequent targeted analysis. This subgraph contains key information related to the current road conditions and their interrelationships, simplifying and focusing the complex knowledge graph. For example, the constructed subgraph may include a road segment entity, a material entity for that segment, a recent traffic load entity, and edges representing the impact relationship between them. By constructing a subgraph, the connections between these related information can be more clearly presented, providing a structured framework for further exploring potential factors related to road collapse.
[0121] Step S602: input the subgraph into a preset gated spatiotemporal graph convolutional network to determine node features, perform graph convolution on the node features to obtain convolution results, perform multi-hop aggregation on the convolution results to obtain target node features, and predict the first collapsed features based on the target node features.
[0122] In some embodiments, after the constructed subgraph is input into the gated spatiotemporal graph convolutional network, the network will process the initial features of each node and update the feature representation of the node through operations in the spatiotemporal graph convolution, combined with the topological structure and node features of the graph. At the same time, the gating mechanism will adjust the spatiotemporal information flow based on information such as environmental data characteristics, and decide which information needs to be retained or filtered, so as to determine the node features that can better reflect the true state and importance of the node in the spatiotemporal dimension; after that, after determining the node features, the graph convolution operation is used to perform weighted aggregation on the features of the node and its neighboring nodes to extract a higher-level feature representation, that is, to obtain the convolution result. Specifically, the graph convolution operation will transform and fuse the node features according to the structural information of the graph (such as the connection relationship between nodes) and learnable parameters to obtain convolution features. Subsequently, because the influence between nodes may not be limited to directly connected neighboring nodes, the convolution results are iteratively processed multiple times through multi-hop aggregation, gradually integrating the feature information of multi-order neighboring nodes. This allows the resulting target node features to more comprehensively reflect the structure and semantics of the entire subgraph, encompassing a wide range of relevant information from local to global perspectives, providing a more accurate and complete feature foundation for subsequent collapse feature prediction. Finally, the target node features obtained through the previous series of operations have integrated rich spatiotemporal information, inter-node relationship information, and global information after multi-hop aggregation. Based on these target node features, subsequent network reasoning and computation (for example, in the hierarchical causal inference layer, combining physical and data layer reasoning, considering the stress-strain relationship of road materials, and focusing on key causal edges through attention mechanisms) can predict the first collapse feature of the road.
[0123] In this way, by performing hierarchical causal reasoning on the multimodal knowledge graph, we can quickly and deeply explore the intrinsic connections and causal relationships between different data and obtain the first collapsed feature.
[0124] In some embodiments, determining the second collapse feature based on the first collapse feature and multi-source environmental data includes steps S701 to S703:
[0125] Step S701: Acquire the multi-source environmental data of the target road area, and perform data encoding on the multi-source environmental data to obtain a second encoding result, wherein the second encoding result includes a radar encoding result, a temperature and humidity encoding result, and a load encoding result;
[0126] In some embodiments, multi-source environmental data such as geological radar data, temperature and humidity sensor data, and traffic load data of the target road area are obtained. Then, for the geological radar data, a 3D convolutional neural network (3DCNN) is used to process the dielectric constant distribution data of the geological radar to output radar features such as volume, depth, and shape complexity of underground cavities. The radar features are then mapped to road grid nodes (1-meter resolution), and missing areas are filled by bilinear interpolation to obtain radar coding results. For the temperature and humidity sensor data, a bidirectional LSTM is used to extract the time-dependent features in the temperature and humidity sensor data, and based on the time-dependent features, the temperature and humidity mutation points (such as before and after heavy rain) are identified to generate event marker vectors, that is, to determine the temperature and humidity coding results. For the traffic load data, the traffic load data is converted into equivalent axle load (ESAL) to uniformly measure the intensity of the impact of different vehicles on the road, and the ESAL mean and peak values are counted hourly and mapped to the corresponding road grid to present the traffic load distribution in the time and space dimensions and determine the load coding results.
[0127] Step S702: Calculate the association weight between the first collapse feature and the radar coding result, determine the radar enhancement feature based on the association weight and the radar coding result, and perform data splicing on the temperature and humidity coding result and the load coding result to obtain an environmental time series feature;
[0128] In some embodiments, due to the first collapse feature Contains existing collapse related information and calculates the first collapse feature Radar characteristics in the radar coding results The association weight , the relevant formula is: , where Indicates the first collapse feature Query operation is used to extract key information of features; Indicates radar signature Key operation is used to extract key information of features; Represents the dimension of the feature, which is used to normalize the dot product result to prevent the gradient from disappearing. The radar characteristics in the radar coding result are weighted and fused. The relevant formula is: , where Indicates radar signature The value operation is used to obtain the specific content of the feature; the radar enhancement feature can be determined by this formula, which can highlight the underground information that has an important impact on the collapse, so that the subsequent analysis can focus more on the key factors; then, the temperature and humidity time series features Data splicing is performed with traffic load characteristics ESAL to integrate two types of environmental factors information, and the information flow is adjusted through the gate control unit. The gate control unit outputs the control information fusion mode to dynamically determine the contribution of each information according to the actual road conditions, so that the obtained environmental time series characteristics More realistic.
[0129] Step S703: Constrain the radar enhancement feature and the environmental time series feature through the material linear elastic equation to obtain a fusion feature, and generate corresponding dynamic weights based on the first collapsed feature and the fusion feature through a cross-attention mechanism. The first collapsed feature and the fusion feature are weightedly aggregated based on the dynamic weight to obtain a second collapsed feature.
[0130] In some embodiments, first, since road collapse is related to material mechanical properties, the radar enhancement feature is constrained according to the road material constitutive equation (such as a linear elastic model). and the environmental timing characteristics Constraints are performed to obtain fusion features ; Then, based on the first collapse feature and the fusion features The attention mechanism captures the relationship between features and the difference in importance, generating dynamic weights , then, according to the dynamic weight Weighted aggregation of each feature to obtain the second collapsed feature , the relevant formula is: By comprehensively considering the impact of each characteristic on road collapse and integrating them according to their importance, the second collapse characteristic more comprehensively and accurately reflects the underlying factors and patterns. Finally, a counterfactual intervention is performed on the second characteristic (e.g., assuming normal temperature and humidity), and the difference in causal effect is calculated. By comparing the characteristics before and after the intervention, the actual impact of environmental factors on road collapse is clarified, the causal explanatory power of the characteristics is enhanced, and a more reliable basis for risk assessment and prevention is provided.
[0131] It should be noted that the first and the second do not indicate a sequence of events, but can be understood as nouns. The first collapse feature is obtained by performing hierarchical causal reasoning on the multimodal knowledge graph, which is a summary and reasoning result of past experience and data knowledge. The second collapse feature is determined on the basis of the first collapse feature, combined with actual dynamic multi-source environmental data, taking into account the impact of the current actual environmental conditions on road collapse, such as the current rainfall intensity, real-time changes in groundwater levels, etc., and can more accurately reflect the actual collapse risk situation faced by the current road surface.
[0132] In this way, the second collapse feature is determined by integrating multi-source environmental data, taking into account the impact of the current actual environmental conditions on road collapse, and obtaining a second collapse feature that more comprehensively reflects the actual environmental conditions of the current road surface.
[0133] Step S104: input the multi-scale image features and the second collapse features into a reinforcement learning model to determine the road collapse risk level.
[0134] In some embodiments, the multi-scale image features and the second collapse features are input into a reinforcement learning model to determine the road collapse risk level, specifically: the image features and the second collapse features are spatially aligned and cross-modal feature fused to obtain a comprehensive feature vector; the comprehensive feature vector and the multi-source environmental data are used as the state space of the reinforcement learning model, and based on the state space, the preset action space and the preset reward function, the preset reinforcement learning model is used to iteratively update the policy network parameters with the goal of maximizing the expected cumulative value of the reward function, determine the optimal risk assessment strategy, and determine the road collapse risk level based on the risk assessment strategy. Specifically, first, the multi-scale image features (shallow / middle / deep layers) and the second collapse features (physically corrected environmental fusion features) are unified into the same road grid coordinate system, and resolution alignment is achieved through bilinear interpolation. The cross-modal attention mechanism is used to calculate the association weights between the image features and the second collapse features, and a weighted comprehensive feature vector is generated. Subsequently, the comprehensive feature vector and multi-source environmental data (geological radar cavity distribution, temperature and humidity time series, traffic load statistics) are mapped to a unified state vector through a fully connected layer, which is determined as the state space. The dynamically adjusted classification threshold (such as the high-risk probability threshold), attention weight (controlling the contribution ratio of multi-source data), and physical constraint coefficient are defined as the state space. At the same time, the reward function is defined based on the comprehensive model accuracy and risk control requirements: , where Prediction accuracy for the model; is the recall rate of collapsed samples; is the false alarm rate; A score (0-1) is assigned to the consistency of the physical constraints, which measures whether the predictions are consistent with the material mechanics equations. 、 and The weight coefficients are all set manually. Finally, the Proximal Policy Optimization (PPO) algorithm is used to iteratively update the policy network parameters with the goal of maximizing the cumulative reward. , where represents the policy network, and its parameters are , used to determine the probability of taking a certain action in a given state; express expectations; Represents finding the parameters that maximize the objective function (here, the expected cumulative reward) process; Represents the cumulative reward, starting from the current time step To the end time step , the reward at each time step Discount factor The discounted sum; and based on the optimal strategy The comprehensive feature vector is classified and the road collapse risk level (low / medium / high / urgent) is output.
[0135] In this way, by inputting the multi-scale image features and the second collapse features into the reinforcement learning model, the correlation and complementarity between different features can be fully explored, and the impact of these two features on the road collapse risk can be comprehensively considered to generate a more accurate road collapse risk level.
[0136] By collecting visible light and infrared images, the embodiments of the present application can intuitively obtain the morphological and structural characteristics of the road surface. By performing cross-modal fusion, the advantages of these two types of information can be complemented to more accurately reflect the actual state of the road surface, facilitating the subsequent accurate and rapid extraction of collapse-related features. By inputting the road surface images into a lightweight convolutional neural network, information about the road surface at different spatial scales can be captured. By updating the collapse factor set, features related to road collapse can be accurately and quickly determined, improving the accuracy and real-time performance of monitoring. By constructing a multimodal knowledge graph, rich multi-source information is provided for subsequent causal reasoning. By performing hierarchical causal reasoning on the multimodal knowledge graph, the intrinsic connections and causal relationships between different data can be deeply explored. By integrating multi-source environmental data to determine the second collapse feature, the impact of the current actual environmental conditions on road collapse is considered, and a second collapse feature that more comprehensively reflects the actual environmental conditions of the current road surface is obtained. By inputting the multi-scale image features and the second collapse features into the reinforcement learning model, the correlation and complementarity between different features can be fully explored, and the impact of these two features on road collapse risk can be comprehensively considered to generate a more accurate road collapse risk level. Compared with the existing technology, this application can improve the accuracy and real-time performance of road subsidence risk monitoring.
[0137] Example 2
[0138] Please refer to Figure 2 , Figure 2 1 is a schematic structural diagram of an embodiment of a road subsidence risk monitoring system provided by the present application, comprising an acquisition module 100, an update module 200, an inference module 300, and a generation module 400;
[0139] The acquisition module 100 is configured to acquire a visible light image and an infrared image of a target road area, and perform cross-modal fusion of the visible light image and the infrared image using a generative adversarial network constrained by thermodynamic equations to generate a road surface image.
[0140] The updating module 200 is configured to input the road surface image into a lightweight convolutional neural network to extract multi-scale image features. When the multi-scale image features meet a preset condition, the initial collapse factor set is updated based on the feature distribution divergence difference within a preset sliding time window to obtain a target collapse factor set.
[0141] The reasoning module 300 is configured to construct a multimodal knowledge graph based on the collapse factor set, historical collapse data, and historical environmental data, perform hierarchical causal reasoning on the multimodal knowledge graph through a gated spatiotemporal graph convolutional network to obtain a first collapse feature, and determine a second collapse feature based on the first collapse feature and multi-source environmental data;
[0142] The generating module 400 is configured to input the multi-scale image features and the second collapse features into a reinforcement learning model to determine a road collapse risk level.
[0143] The information interaction, execution process, etc. between the modules within the above-mentioned road subsidence risk monitoring system are based on the same concept as the embodiment of the road subsidence risk monitoring method of the first aspect of the present invention, and the technical effects achieved are basically the same. For specific contents, please refer to the description in the first embodiment of the method of the present invention, and will not be repeated here.
[0144] The apparatus embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separate, i.e., they may be located in one location or distributed across multiple network elements. Some or all of these elements may be selected based on actual needs to achieve the objectives of the methods of this embodiment.
[0145] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the road subsidence risk monitoring method as described in the first embodiment above is implemented.
[0146] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-monitorable storage medium. When executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0147] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the scope of protection of the present application.
[0148] It is particularly pointed out that for those skilled in the art, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A road subsidence risk monitoring method, characterized in that: include: Collecting visible light images and infrared images of a target road area, and cross-modally fusing the visible light images and the infrared images using a generative adversarial network constrained by thermodynamic equations to generate a road surface image; Inputting the road surface image into a lightweight convolutional neural network to extract multi-scale image features, and when the multi-scale image features meet preset conditions, updating an initial collapse factor set based on feature distribution divergence differences within a preset sliding time window to obtain a target collapse factor set; Constructing a multimodal knowledge graph based on the target collapse factor set, historical collapse data, and historical environmental data, performing hierarchical causal reasoning on the multimodal knowledge graph through a gated spatiotemporal graph convolutional network to obtain a first collapse feature, and determining a second collapse feature based on the first collapse feature and the multi-source environmental data; Inputting the multi-scale image features and the second collapse features into a reinforcement learning model to determine a road collapse risk level; The generative adversarial network constrained by thermodynamic equations performs cross-modal fusion of the visible light image and the infrared image to generate a road surface image, specifically: performing image preprocessing on the visible light image and the infrared image respectively to obtain a first preprocessing result and a second preprocessing result; The first preprocessing result and the second preprocessing result are stitched together to obtain a stitched image, and the stitched image is input into the generative adversarial network to adaptively enhance the stitched image in combination with network parameters to generate the road surface image, wherein the generative adversarial network includes a generator and a discriminator, the network parameters are calculated by the generator to calculate the matching error between the infrared channel temperature field of the generated image and the Fourier heat conduction equation, and the matching error is back-propagated to optimize the generated image generated by the generator, and the discriminator determines the classification result of the generated image and the stitched image after receiving the generated image, and optimizes the result after determining the loss function value based on the classification result; The initial collapse factor set is updated based on the feature distribution divergence difference within the preset sliding time window to obtain the target collapse factor set, specifically: Obtaining the feature distribution of the sliding time window, and calculating the feature distribution divergence difference between the feature distribution and the historical feature distribution; If the feature distribution divergence difference is greater than a preset difference threshold, the abnormal features within the sliding time window are extracted, and incremental clustering is performed based on the abnormal features to determine candidate collapse factors. The cosine similarity between the candidate collapse factors and each collapse factor in the initial collapse factor set is calculated. If the cosine similarity is greater than a preset similarity threshold, the candidate collapse factor is merged with the nearest collapse factor; otherwise, the candidate collapse factor is added to the initial collapse factor set to obtain the updated target collapse factor set.
2. The road subsidence risk monitoring method according to claim 1, characterized in that: The formula for the matching error between the infrared channel temperature field of the calculated image and the Fourier heat conduction equation is specifically: ; Where, is the matching error; is a preset weight coefficient used to balance the matching error and other loss terms; To generate the position in the infrared channel of the image The temperature value at is the Laplace operator of the temperature field; is the rate of change of temperature with time; The thermal diffusivity of the material characterizes the diffusion rate of heat in the material and is dynamically adjusted according to real-time ambient temperature and humidity data.
3. The road subsidence risk monitoring method according to claim 1, characterized in that: The road surface image is input into a lightweight convolutional neural network to extract multi-scale image features, specifically: Inputting the road surface image into a lightweight convolutional neural network to extract an initial feature map, wherein the initial feature map includes a shallow feature map, a middle feature map, and a deep semantic feature map; The scale feature maps in the initial feature map are aligned and flattened into corresponding spatial sequences, and position encoding is performed on each position in the spatial sequences to obtain a first encoding result corresponding to each level, and the feature correlation of each first encoding result is calculated through a multi-head attention mechanism to generate a self-attention weight corresponding to each level, and the multi-scale image feature is obtained based on each self-attention weight and the initial feature map.
4. The road subsidence risk monitoring method according to claim 1, characterized in that: The multimodal knowledge graph is constructed based on the collapse factor set, historical collapse data, and historical environmental data, specifically: Acquire historical collapse data and corresponding historical environmental data of a target road area, and perform data preprocessing on the historical collapse data and the historical environmental data, respectively, to obtain a third preprocessing result and a fourth preprocessing result, wherein the historical collapse data includes historical image data and historical text data, and the historical environmental data includes historical geological radar detection data, historical temperature and humidity sensor data, and historical traffic load time series data; Extracting a plurality of entities from the third preprocessing result and the fourth preprocessing result, and determining entity types corresponding to the entities and relationships between the entities, wherein the entity types include road structure entities, environment entities, collapse event entities, and collapse factor entities; According to the entity type, each entity is mapped to a node in the knowledge graph, and the relationship is mapped to an edge in the knowledge graph to obtain a multimodal knowledge graph.
5. The road subsidence risk monitoring method according to claim 1, characterized in that: The gated spatiotemporal graph convolutional network is used to perform hierarchical causal reasoning on the multimodal knowledge graph to obtain the first collapsed feature, specifically: Based on the image features, querying the multimodal knowledge graph for relevant road entities and road edges, and constructing a subgraph based on the road entities and the road edges; The subgraph is input into a preset gated spatiotemporal graph convolutional network to determine node features, and graph convolution is performed on the node features to obtain convolution results, and multi-hop aggregation is performed on the convolution results to obtain target node features, and the first collapsed features are predicted based on the target node features.
6. The road subsidence risk monitoring method according to claim 1, characterized in that: The determining of the second collapse feature based on the first collapse feature and multi-source environmental data is specifically: Acquire the multi-source environmental data of the target road area, and encode the multi-source environmental data to obtain a second encoding result, wherein the second encoding result includes a radar encoding result, a temperature and humidity encoding result, and a load encoding result; Calculating an association weight between the first collapse feature and the radar coding result, determining a radar enhancement feature based on the association weight and the radar coding result, and performing data splicing on the temperature and humidity coding result and the load coding result to obtain an environmental time series feature; The radar enhancement feature and the environmental time series feature are constrained by the material linear elastic equation to obtain a fusion feature, and based on the first collapsed feature and the fusion feature, a corresponding dynamic weight is generated through a cross-attention mechanism. The first collapsed feature and the fusion feature are weightedly aggregated based on the dynamic weight to obtain a second collapsed feature.
7. The road subsidence risk monitoring method according to claim 1, characterized in that: The multi-scale image features and the second collapse features are input into a reinforcement learning model to determine the road collapse risk level, specifically: Performing spatial alignment and cross-modal feature fusion on the image features and the second collapsed features to obtain a comprehensive feature vector; The comprehensive feature vector and the multi-source environmental data are used as the state space of the reinforcement learning model. Based on the state space, the preset action space and the preset reward function, the preset reinforcement learning model is used to iteratively update the strategy network parameters with the goal of maximizing the expected cumulative value of the reward function, determine the optimal risk assessment strategy, and determine the road collapse risk level based on the risk assessment strategy.
8. A road subsidence risk monitoring system, characterized in that: include: Acquisition module, update module, reasoning module and generation module; The acquisition module is used to acquire visible light images and infrared images of the target road area, and perform cross-modal fusion of the visible light images and the infrared images through a generative adversarial network constrained by thermodynamic equations to generate a road surface image; The updating module is configured to input the road surface image into a lightweight convolutional neural network to extract multi-scale image features, and when the multi-scale image features meet preset conditions, update the initial collapse factor set based on the feature distribution divergence difference within a preset sliding time window to obtain a target collapse factor set; The reasoning module is used to construct a multimodal knowledge graph based on the target collapse factor set, historical collapse data, and historical environmental data, and perform hierarchical causal reasoning on the multimodal knowledge graph through a gated spatiotemporal graph convolutional network to obtain a first collapse feature, and determine a second collapse feature based on the first collapse feature and multi-source environmental data; The generating module is configured to input the multi-scale image features and the second collapse features into a reinforcement learning model to determine a road collapse risk level; The generative adversarial network constrained by thermodynamic equations performs cross-modal fusion of the visible light image and the infrared image to generate a road surface image, specifically: performing image preprocessing on the visible light image and the infrared image respectively to obtain a first preprocessing result and a second preprocessing result; The first preprocessing result and the second preprocessing result are stitched together to obtain a stitched image, and the stitched image is input into the generative adversarial network to adaptively enhance the stitched image in combination with network parameters to generate the road surface image, wherein the generative adversarial network includes a generator and a discriminator, the network parameters are calculated by the generator to calculate the matching error between the infrared channel temperature field of the generated image and the Fourier heat conduction equation, and the matching error is back-propagated to optimize the generated image generated by the generator, and the discriminator determines the classification result of the generated image and the stitched image after receiving the generated image, and optimizes the result after determining the loss function value based on the classification result; The initial collapse factor set is updated based on the feature distribution divergence difference within the preset sliding time window to obtain the target collapse factor set, specifically: Obtaining the feature distribution of the sliding time window, and calculating the feature distribution divergence difference between the feature distribution and the historical feature distribution; If the feature distribution divergence difference is greater than a preset difference threshold, the abnormal features within the sliding time window are extracted, and incremental clustering is performed based on the abnormal features to determine candidate collapse factors. The cosine similarity between the candidate collapse factors and each collapse factor in the initial collapse factor set is calculated. If the cosine similarity is greater than a preset similarity threshold, the candidate collapse factor is merged with the nearest collapse factor; otherwise, the candidate collapse factor is added to the initial collapse factor set to obtain the updated target collapse factor set.
Citation Information
Patent Citations
Visible light and infrared light image fusion method and device for road crack detection
CN119067867A
Road condition monitoring identification method, device and equipment based on machine vision and medium
CN119445503A
Road hidden danger sensitive element feature factorization method based on numerical mapping
CN119691951A
Road collapse risk comprehensive assessment method based on multi-source data fusion
CN119740124A