Data Analysis Method and System for Road Asset Condition Monitoring Using Deep Learning
By acquiring surface image data and structural sensing data of road assets, cross-modal feature association mapping and feature evolution dependency network construction are carried out, which solves the shortcomings of intrinsic correlation and dynamic dependency in road asset condition monitoring and realizes comprehensive and accurate monitoring and prediction of road asset condition.
Patent Information
- Application Number
- CN202511256879.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-04
AI Technical Summary
Existing technologies are insufficient to effectively establish the intrinsic relationship between the surface morphology and internal structural features of road assets, and static models cannot capture the dynamic dependencies and coupling mechanisms of features over time, resulting in insufficient accuracy and foresight in road asset condition monitoring.
By acquiring road surface image data and road structure sensing data, cross-modal feature association mapping is performed to generate a cross-modal association feature matrix, and a feature evolution dependency network is constructed to capture the dynamic dependency relationship of cross-modal feature dimensions over time and the coupling effect between features.
It has improved the comprehensiveness and accuracy of road asset status, enabling integrated current diagnosis and future trend prediction, thus avoiding the one-sidedness of monitoring.
Smart Images

Figure CN120763825B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis, and more specifically, to a method and system for analyzing road asset status monitoring data using deep learning. Background Technology
[0002] Road asset condition monitoring data analysis is a key technology for ensuring the safe operation of road infrastructure. Through continuous monitoring and data analysis of road asset structure status, potential risks can be identified in a timely manner and the evolution trend of structural performance can be assessed. Currently, methods typically rely on image data to identify road surface damage features or only on sensor data to monitor structural stress changes. Then, static feature splicing or simple correlation analysis is used to construct a condition assessment model and output the current condition classification result. These methods often fail to effectively establish the intrinsic relationship between the surface morphology and internal structural features of road assets, resulting in insufficient semantic consistency between feature dimensions. At the same time, static models cannot capture the dynamic dependencies and coupling mechanisms of features over time, making it difficult for the condition assessment results to fully reflect the overall performance and potential evolution trend of the road asset structure, thus affecting the accuracy and foresight of monitoring. Summary of the Invention
[0003] This invention provides a method and system for analyzing road asset condition monitoring data using deep learning.
[0004] In a first aspect, embodiments of the present invention provide a method for analyzing road asset status monitoring data using deep learning, comprising: acquiring a road asset data set, the road asset data set including road asset surface image data and road asset structure sensing data; performing cross-modal feature association mapping on the road asset data set to generate a cross-modal association feature matrix, wherein the row vectors of the cross-modal association feature matrix correspond to a single heterogeneous data sample, and the column vectors correspond to the fused cross-modal feature dimensions; wherein a single heterogeneous data sample is obtained by associating road asset surface image data and road asset structure sensing data at the same time period and spatial location; constructing a feature evolution dependency network based on the cross-modal association feature matrix, the feature evolution dependency network being used to characterize the dynamic dependency relationship of cross-modal feature dimensions over time and the coupling effect between features; and outputting road asset status inference results through the feature evolution dependency network.
[0005] Secondly, embodiments of the present invention provide a computer system, comprising: a memory storing a computer program; and a processor for loading the computer program to implement the road property status monitoring data analysis method using deep learning as described above.
[0006] The road asset condition monitoring data analysis method provided by this invention utilizes deep learning. By acquiring road asset surface image data containing surface morphology and texture information and road asset structural sensing data containing internal stress and strain information, a multi-source heterogeneous road asset data set is formed. This method fully leverages the complementary information of different modal data. By performing cross-modal feature correlation mapping on the multi-source heterogeneous road asset data set to generate a cross-modal correlation feature matrix, an intrinsic correlation relationship between surface texture features and internal stress features can be established, forming a correlation feature dimension that integrates dual-modal information. This overcomes the limitations of conventional multi-modal feature simple physical splicing. The problem of semantic fragmentation is addressed by constructing a feature evolution dependency network based on a cross-modal correlation feature matrix. This network can capture the dynamic dependencies of cross-modal feature dimensions over time and the coupling mechanism between features, overcoming the limitation of static feature correlation models that can only analyze the correlation of features at the current moment. By outputting road asset status inference results that include the current structural integrity category and the potential risk evolution direction of road assets through the feature evolution dependency network, the current diagnosis of road asset status and future trend prediction can be integrated, avoiding the one-sidedness of monitoring caused by only outputting a single current status classification result, thereby effectively improving the comprehensiveness and accuracy of road asset status monitoring. Attached Figure Description
[0007] Figure 1 This is a flowchart of a road asset status monitoring data analysis method using deep learning, provided by an embodiment of the present invention.
[0008] Figure 2 This is a schematic diagram of the composition of a computer system provided in an embodiment of the present invention. Detailed Implementation
[0009] Please see Figure 1 The flowchart below illustrates a method for analyzing road asset status monitoring data using deep learning, as provided in an embodiment of the present invention. This method can be executed by a computer system and includes the following steps:
[0010] Step S100: Obtain a road property data set, which includes road property surface image data and road property structure sensing data. The road property surface image data contains road property surface morphology and texture information, and the road property structure sensing data contains road property internal stress and strain information.
[0011] The road asset dataset is a fundamental data set used for monitoring and analyzing the condition of road assets. It consists of two different types of data: road asset surface image data and road asset structure sensor data. Road asset surface image data is image information of the road asset surface acquired through image acquisition equipment. These images contain morphological and textural information about the road asset surface, such as the texture of cracks in the road surface and the texture of rust on the bridge surface. This texture information can intuitively reflect the condition of the road asset surface. Road asset structure sensor data, on the other hand, is data collected by sensors installed inside the road asset. It includes stress and strain information within the road asset, such as stress changes within the roadbed and strain conditions within the bridge structure.
[0012] When acquiring road asset data sets, for road asset surface image data, drones equipped with high-definition cameras can be used to take aerial photos of the road assets, flying along a predetermined route and altitude to obtain complete road asset surface images. For road asset structural sensing data, strain gauges, pressure sensors, and other sensing devices can be installed on existing road assets during road construction to collect stress and strain data inside the road assets in real time.
[0013] Step S200: Perform cross-modal feature association mapping on the road production data set to generate a cross-modal association feature matrix. The row vectors of the cross-modal association feature matrix correspond to a single heterogeneous data sample, and the column vectors correspond to the dimensions of the fused cross-modal features.
[0014] Cross-modal feature association mapping is the process of associating and mapping data from different modalities (road property surface image data and road property structure sensor data) in a road property dataset. The aim is to find the intrinsic connections between different modalities and extract more valuable information. The cross-modal association feature matrix is a matrix generated after cross-modal feature association mapping. Its row vectors correspond to a single heterogeneous data sample; that is, each row represents a road property data sample containing comprehensive information from both road property surface image data and road property structure sensor data. Specifically, road property surface image data and road property structure sensor data from the same time period and spatial location can be associated as a single heterogeneous data sample, forming a sample sequence. The column vectors correspond to the fused cross-modal feature dimensions; that is, each column represents a fused feature dimension, which is obtained by integrating the features of both road property surface image data and road property structure sensor data.
[0015] As one implementation method, step S200 can be specifically implemented as the following steps S210~S270:
[0016] Step S210: Divide the road property surface image data in the road property data set into regions to obtain a road property image region set. Each road property image region in the road property image region set contains representative morphological texture information of the road property surface.
[0017] Region segmentation involves dividing this large image into multiple smaller regions, forming a road property image region set. The road property image region set is a collection of multiple road property image regions, each containing representative morphological and textural information about the road property surface. This information can be used for subsequent feature extraction and analysis.
[0018] When performing region segmentation, image segmentation algorithms can be used. For example, threshold-based segmentation algorithms can be used, setting a threshold based on the grayscale or color values of the image, and dividing pixels with grayscale or color values within the threshold range into regions. Edge detection-based segmentation algorithms can also be used, detecting edges in the image and dividing the areas enclosed by those edges into road image regions. Taking a road surface image as an example, the road surface image can be divided into multiple small blocks according to certain rules; each small block is a road image region. These regions may contain different texture information, such as crack textures, wear textures, etc.
[0019] Step S220: Perform texture enhancement on each road property image region in the road property image region set to obtain an enhanced texture region.
[0020] Texture enhancement is a process of strengthening and highlighting texture information in road property image regions. Its purpose is to make the texture clearer and more prominent, facilitating subsequent feature extraction and analysis. The enhanced texture region is the region obtained after texture enhancement processing; its texture information is more prominent, better reflecting the morphological characteristics of the road property surface.
[0021] Various image processing algorithms can be used for texture enhancement. For example, histogram equalization can be used to adjust the image's histogram, making the grayscale distribution more uniform, thereby enhancing image contrast and highlighting texture information. Filtering algorithms, such as Gaussian filtering and Laplacian filtering, can also be used to filter the image, removing noise while enhancing the edge information of the texture. For a road infrastructure image region containing rust texture on a bridge surface, histogram equalization can be used to process it, making the rust texture more clearly visible; the result is the enhanced texture region.
[0022] Step S230: Perform spatiotemporal correlation anchoring on the road property structure sensing data in the road property data set to obtain spatiotemporal anchored sensing data. The spatiotemporal anchored sensing data carries spatial location identifiers and temporal correlation identifiers corresponding to the road property image areas.
[0023] Road asset spatiotemporal correlation anchoring is the process of spatially and temporally associating and matching road asset structure sensing data with road asset surface image data. The aim is to establish a connection between the road asset structure sensing data and the corresponding road asset image region, ensuring temporal correlation. Spatiotemporally anchored sensing data is obtained after road asset spatiotemporal correlation anchoring, carrying spatial location identifiers and temporal correlation identifiers corresponding to the road asset image region. This allows us to determine the road asset image region corresponding to each sensing data point and the time of acquisition.
[0024] As one implementation method, step S230 can be specifically implemented as the following steps S231~S236:
[0025] Step S231: Extract the original acquisition information of the road property structure sensor data and the imaging area information of the road property surface image data from the road property data set. The original acquisition information includes the sensor data acquisition location and acquisition time, and the imaging area information includes the road property entity location and imaging time corresponding to the image area.
[0026] Raw acquisition information consists of the basic information recorded during the acquisition of road asset structure sensor data, including the sensor data acquisition location and acquisition time. The sensor data acquisition location can be determined by the sensor's installation location; for example, if a sensor is installed on a bridge pier, its acquisition location is the specific location of that pier. The acquisition time is the point in time when the sensor acquires the data. Imaging area information is the information recorded during the imaging process of road asset surface image data, including the location of the road asset entity corresponding to the image area and the imaging time. The location of the road asset entity corresponding to the image area can be determined by the positioning information of the image acquisition device and the image's shooting range. The imaging time is the time when the image was captured.
[0027] When extracting raw acquisition information and imaging area information, this information can be obtained from the metadata of the road property dataset. The metadata of the road property dataset records basic data information, typically including the acquisition time, acquisition location, and data source. For example, in the database storing the road property dataset, each piece of road property structure sensing data and road property surface image data has a corresponding metadata record. By querying this metadata, the raw acquisition information and imaging area information can be extracted.
[0028] Step S232: Construct a spatial association graph of road assets. The spatial association graph of road assets uses road asset entity segments as nodes and the topological connection relationship between nodes as edges. It stores the spatial coordinate range and road asset type attributes corresponding to each node.
[0029] A road property spatial relationship map is a graph structure used to represent the spatial relationships of road assets. It uses road asset entity segments as nodes, with each node representing a segment of a road asset entity, such as a section of a road or a pier of a bridge. The topological connections between nodes are represented by edges, indicating the connections between road asset entity segments, such as the connection between adjacent road segments or the connection between a bridge pier and its beam. The road property spatial relationship map also stores the spatial coordinate range and road asset type attribute corresponding to each node. The spatial coordinate range determines the spatial location of the road asset entity segment, while the road asset type attribute indicates the type of the road asset entity segment, such as road, bridge, or tunnel. When constructing a road property spatial relationship map, a Geographic Information System (GIS) can be used. First, spatial data of road asset entities is collected, including information such as the location, shape, and size of the road asset entities. Then, according to the segmentation rules of the road asset entities, the road asset entities are divided into multiple nodes. Next, the topological connections between the nodes are determined, and the edges of the graph are constructed. Finally, the spatial coordinate range and road asset type attribute of each node are stored in the graph.
[0030] Step S233: Based on the road property spatial association map, spatially match the sensor data acquisition location in the original acquisition information with the road property entity location in the imaging area information to determine the road property image area to which each sensor data acquisition location belongs, and obtain the spatial association result.
[0031] Spatial matching is the process of comparing and matching the locations of sensor data acquisition points with the locations of road property entities. Its purpose is to determine the road property image region to which each sensor data acquisition point belongs. The spatial association result is obtained after spatial matching; it records the correspondence between the sensor data acquisition location and its corresponding road property image region.
[0032] As one implementation method, step S233 can be specifically implemented as the following steps S2331~S2336:
[0033] Step S2331: Extract the spatial coordinate range of each node from the road property spatial association map to obtain the node coordinate range set.
[0034] The node coordinate range set is a collection of the spatial coordinate ranges of each node in the road property spatial association graph. A spatial coordinate range can be represented by a coordinate interval; for example, a rectangular area on a two-dimensional plane can be represented by the coordinates of its lower left and upper right corners. When extracting node coordinate ranges, this information can be obtained from the data structure of the road property spatial association graph. Road property spatial association graphs are typically stored in the form of a graph database or files, where each node has corresponding attribute information, including its spatial coordinate range. By querying this attribute information, the spatial coordinate ranges of each node can be extracted, forming the node coordinate range set. For example, in a road property spatial association graph stored in a graph database, each node has an attribute field recording its spatial coordinate range. By writing a query statement, the spatial coordinate ranges of all nodes can be extracted, resulting in the node coordinate range set.
[0035] Step S2332: For each sensor data acquisition location in the original acquisition information, extract the coordinate information of that sensor data acquisition location.
[0036] The coordinate information of the sensor data acquisition location is the information that determines the specific location of that location in space, and is usually represented by latitude and longitude or Cartesian coordinates.
[0037] When extracting coordinate information of sensor data acquisition locations, it can be obtained from the raw acquisition information. The raw acquisition information typically contains detailed information about the sensor data acquisition locations, including coordinate information. For example, during the acquisition of road structure sensor data, the sensors record their own position information, which is stored in the raw acquisition information. By reading the raw acquisition information, the coordinate information of each sensor data acquisition location can be extracted.
[0038] Step S2333: Traverse each node coordinate range in the set of node coordinate ranges and determine whether the coordinate information falls within the coordinate range of that node.
[0039] Traversing the set of node coordinate ranges involves checking each node coordinate range in the set sequentially. Determining whether coordinate information falls within a node's coordinate range involves comparing the coordinates of the sensor data acquisition location with the node's coordinate range. If the coordinates are within the defined interval of the node's coordinate range, they are considered to fall within that range. Coordinate comparison methods can be used for this determination. For example, for coordinates on a two-dimensional plane and a rectangular node coordinate range, we check whether the x-coordinate of the sensor data acquisition location is within the x-coordinate interval of the node's coordinate range, and whether the y-coordinate is within the y-coordinate interval of the node's coordinate range. If both conditions are met, the coordinates are considered to fall within the node's coordinate range. Taking a road infrastructure spatial association map as an example, each node coordinate range represents a segment of the road. For the coordinates of a sensor data acquisition location, we sequentially check whether it falls within the node coordinate range of each road segment.
[0040] Step S2334: When the coordinate information falls within the coordinate range of a certain node, determine the road property image area corresponding to the coordinate range of that node as the road property image area to which the sensor data acquisition location belongs.
[0041] When the coordinates of the sensor data acquisition location fall within the coordinate range of a certain node, it indicates that the sensor data acquisition location and the road property entity segment corresponding to that node are spatially related. Therefore, it can be determined that the road property image area corresponding to the coordinate range of that node is the road property image area to which the sensor data acquisition location belongs.
[0042] For example, in the spatial correlation map of road assets of a bridge, there is a node that represents a pier of the bridge. The coordinate range of the node includes the coordinate information of the acquisition location of a sensor installed on the pier. Then, the road asset image area corresponding to the node (such as the image area on the surface of the pier) is determined as the road asset image area to which the sensor data acquisition location belongs.
[0043] Step S2335: Record the correspondence between the sensor data acquisition location and the corresponding road property image area as a spatial association unit.
[0044] A spatial association unit is a unit that records the correspondence between a sensor data acquisition location and its associated road property image area. It can be represented by a data structure, such as a binary tuple (sensor data acquisition location, road property image area). When recording a spatial association unit, the coordinate information of the sensor data acquisition location and the identification information of the road property image area can be combined to form a spatial association unit. For example, in a database table, create one field to record the coordinates of the sensor data acquisition location and another field to record the identification of the road property image area; combining these into a single record constitutes a spatial association unit.
[0045] Step S2336: Summarize the spatial association units corresponding to all sensor data acquisition locations to obtain the spatial association results.
[0046] The spatial correlation result is a summary of the spatial correlation units corresponding to all sensor data acquisition locations, which can be represented as a list or set. When summarizing spatial correlation units, the spatial correlation units corresponding to each sensor data acquisition location can be added to a list or set, ultimately yielding the spatial correlation result. For example, for multiple sensor data acquisition locations in a road asset dataset, the spatial correlation units corresponding to each location are added sequentially to a list; this list is the spatial correlation result.
[0047] Step S234: Based on the spatial correlation results, the acquisition time in the original acquisition information is temporally correlated with the imaging time of the corresponding road property image area, and the correlation strength value between the acquisition time and the imaging time is calculated. The correlation strength value decreases as the time interval increases.
[0048] Temporal correlation is the process of linking the time of sensor data acquisition with the time of imaging of the road property image area to determine the temporal relationship between the two. The correlation strength value represents the degree of correlation between the acquisition time and the imaging time; it decreases as the time interval increases, meaning the larger the time interval, the smaller the correlation strength value. In performing temporal correlation, the road property image area corresponding to each sensor data acquisition location is first determined based on the spatial correlation results. Then, the acquisition time of that sensor data acquisition location is compared with the imaging time of the corresponding road property image area, and the time interval between them is calculated. Finally, the correlation strength value is calculated based on the time interval. For example, a linear or exponential function can be used to calculate the correlation strength value; the larger the time interval, the smaller the correlation strength value.
[0049] Step S235: Filter the sensor data units whose correlation strength values meet the preset strength threshold to obtain preliminary anchoring sensor data. The preliminary anchoring sensor data carries the corresponding road property image area identifier.
[0050] The preset strength threshold is a pre-defined threshold for association strength, used to filter out sensor data units with high association strength. The initial anchored sensor data is the filtered data, carrying corresponding road property image area identifiers, indicating a strong temporal and spatial association between these sensor data and the corresponding road property image areas. When filtering sensor data units, the association strength value of each sensor data unit is compared with the preset strength threshold. If the association strength value is greater than or equal to the preset strength threshold, the sensor data unit is considered to meet the filtering criteria and is retained.
[0051] Step S236: Perform spatiotemporal label fusion on the preliminary anchoring sensing data, embedding the road property image area identifier, spatial correlation result and temporal correlation strength value into the sensing data unit to obtain spatiotemporal anchoring sensing data.
[0052] Spatiotemporal tag fusion is the process of fusing road property image region identifiers, spatial correlation results, and temporal correlation strength values with sensor data units. The aim is to enable the sensor data units to contain more comprehensive spatiotemporal information. Spatiotemporally anchored sensor data, obtained after spatiotemporal tag fusion, carries spatial location identifiers and temporal correlation identifiers corresponding to the road property image regions, enabling better association and matching with road property surface image data.
[0053] When performing spatiotemporal label fusion, the road property image region identifier, spatial correlation result, and temporal correlation strength value can be added as new fields to the sensor data unit. For example, in the database table storing the sensor data units, add the road property image region identifier field, spatial correlation result field, and temporal correlation strength value field, and fill in the corresponding information into these fields to complete the spatiotemporal label fusion and obtain the spatiotemporally anchored sensor data.
[0054] Step S240: Input the enhanced texture region and spatiotemporal anchoring sensing data into the cross-modal mapping network. Extract local texture features from the enhanced texture region through the region feature extraction layer of the cross-modal mapping network to obtain the region texture feature vector. Extract dynamic stress features from the spatiotemporal anchoring sensing data through the temporal feature extraction layer of the cross-modal mapping network to obtain the temporal stress feature vector.
[0055] Cross-modal mapping networks are network models used to map and fuse data from different modalities (road property surface image data and road property structure sensing data). They include a region feature extraction layer and a temporal feature extraction layer. The region feature extraction layer extracts local texture features from enhanced texture regions, aiming to extract representative local texture features. The region texture feature vector, obtained after extraction by the region feature extraction layer, contains local texture feature information of the enhanced texture region. The temporal feature extraction layer extracts dynamic stress features from spatiotemporally anchored sensing data, uncovering dynamic stress features within the data. The temporal stress feature vector, obtained after extraction by the temporal feature extraction layer, contains dynamic stress feature information of the spatiotemporally anchored sensing data.
[0056] When inputting enhanced texture regions and spatiotemporally anchored sensing data into a cross-modal mapping network, the data first needs to be preprocessed to meet the network's input requirements. For example, the image data of the enhanced texture regions is normalized, and the spatiotemporally anchored sensing data is standardized. Then, the processed data is input into the cross-modal mapping network. The region feature extraction layer can be implemented using a convolutional neural network (CNN), which extracts features from the enhanced texture regions through convolutional and pooling layers. The temporal feature extraction layer can be implemented using a recurrent neural network (RNN) or a long short-term memory network (LSTM) to extract temporal features from the spatiotemporally anchored sensing data.
[0057] As one implementation method, step S240 can be specifically implemented as the following steps S241~S245:
[0058] Step S241: Input the enhanced texture region into the first feature extraction unit of the region feature extraction layer. Calculate the local neighborhood feature response of the enhanced texture region through the texture-aware convolution kernel of the first feature extraction unit to obtain the first-level texture feature map. The number of channels in the first-level texture feature map corresponds to the number of texture-aware convolution kernels.
[0059] The first-layer feature extraction unit is the first processing unit in the region feature extraction layer, and it contains a texture-aware convolutional kernel. A texture-aware convolutional kernel is a type of convolutional kernel used to extract texture features; it can calculate the feature response of the local neighborhood of the enhanced texture region. The calculation of the local neighborhood feature response involves sliding the texture-aware convolutional kernel across the enhanced texture region, calculating the feature response value for each local neighborhood. The first-layer texture feature map is the feature map obtained after calculating the local neighborhood feature response. Its number of channels corresponds to the number of texture-aware convolutional kernels, with each channel corresponding to the feature information extracted by one texture-aware convolutional kernel. During the calculation of the local neighborhood feature response, the texture-aware convolutional kernel slides across the enhanced texture region. Each time it slides to a local neighborhood, the weights of the convolutional kernel are convolved with the pixel values in that neighborhood to obtain a feature response value. Combining the feature response values of all local neighborhoods yields the first-layer texture feature map.
[0060] Step S242: Input the first-level texture feature map into the first feature aggregation unit of the region feature extraction layer, and perform feature aggregation on the spatial dimension of the first-level texture feature map through the first feature aggregation unit to obtain the first aggregated feature map. The spatial size of the first aggregated feature map is smaller than the spatial size of the first-level texture feature map.
[0061] The first feature aggregation unit is used to aggregate features from the first-level texture feature map. By manipulating the spatial dimension of the first-level texture feature map, it reduces the spatial size of the feature map while retaining important feature information. Feature aggregation is a process of merging and compressing multiple feature values to reduce data redundancy and improve the expressive power of features. The first aggregated feature map is the feature map obtained after feature aggregation, and its spatial size is smaller than that of the first-level texture feature map. Pooling operations, such as max pooling and average pooling, can be used during feature aggregation. Taking max pooling as an example, the first-level texture feature map is divided into multiple non-overlapping regions, and the maximum value within each region is taken as the feature value of that region. Combining the feature values of all regions yields the first aggregated feature map.
[0062] Step S243: Input the first aggregated feature map into the second layer feature extraction unit of the region feature extraction layer, and perform multi-receptive field feature extraction on the first aggregated feature map through the multi-scale convolution kernel of the second layer feature extraction unit to obtain the second-level texture feature map. The number of channels of the second-level texture feature map is the total number of multi-scale convolution kernels.
[0063] The second-layer feature extraction unit is the second processing unit in the region feature extraction layer, containing multi-scale convolutional kernels. These multi-scale convolutional kernels are a set of kernels at different scales, capable of performing multi-receptive field feature extraction on the first aggregated feature map, uncovering texture feature information at different scales. Multi-receptive field feature extraction involves using convolutional kernels of different scales to perform convolution operations on the first aggregated feature map, extracting feature information from different receptive fields. The second-layer texture feature map is the feature map obtained after multi-receptive field feature extraction; its number of channels is equal to the total number of multi-scale convolutional kernels, with each channel corresponding to the feature information extracted by one multi-scale convolutional kernel.
[0064] When performing multi-receptive field feature extraction, multi-scale convolutional kernels are slid convolutionally applied to the first aggregated feature map, with each kernel extracting feature information at different scales. For example, by using two multi-scale convolutional kernels of different scales to extract multi-receptive field features from the first aggregated feature map, the resulting second-level texture feature map has two channels, with each channel representing the texture feature information extracted by one convolutional kernel.
[0065] Step S244: Input the second-level texture feature map into the second feature aggregation unit of the region feature extraction layer. The second feature aggregation unit performs feature aggregation on the channel dimension of the second-level texture feature map to obtain the second aggregated feature map. The number of channels in the second aggregated feature map is less than the number of channels in the second-level texture feature map.
[0066] The second feature aggregation unit is used to aggregate features from the second-level texture feature map. It operates on the channel dimension of the second-level texture feature map to reduce the number of channels while retaining important feature information. Feature aggregation is a process of merging and compressing feature information from multiple channels to reduce data redundancy and improve feature expressiveness. The second aggregated feature map is the feature map obtained after feature aggregation, and its number of channels is less than that of the second-level texture feature map.
[0067] When performing feature aggregation, methods such as global average pooling or global max pooling can be used. Taking global average pooling as an example, a global averaging operation is performed on each channel of the second-level texture feature map, averaging the values of all pixels in each channel to obtain a scalar value. Combining the scalar values of all channels yields the second aggregated feature map.
[0068] Step S245: Flatten the second aggregated feature map in terms of spatial dimensions to obtain a flattened texture vector. Perform nonlinear feature transformation on the flattened texture vector to obtain a region texture feature vector.
[0069] Spatial dimension flattening is the process of converting the two-dimensional spatial structure of the second aggregated feature map into a one-dimensional vector. Its purpose is to serialize the feature information of the feature map for easier subsequent processing. The flattened texture vector is the vector obtained after spatial dimension flattening and contains all the feature information of the second aggregated feature map. Nonlinear feature transformation is the process of performing a nonlinear transformation on the flattened texture vector, aiming to increase the expressive power of the features and uncover more complex feature information. The region texture feature vector is the vector obtained after nonlinear feature transformation; it contains local texture feature information of the enhanced texture region.
[0070] During spatial flattening, the pixel values of each row of the second aggregated feature map are concatenated sequentially to form a one-dimensional vector. Activation functions such as ReLU and Sigmoid can be used for non-linear feature transformation. Taking ReLU as an example, each element of the flattened texture vector is input into the ReLU function, resulting in a new vector, which is the region texture feature vector.
[0071] As one implementation method, in step S245, a nonlinear feature transformation is performed on the flattened texture vector to obtain a region texture feature vector, which can be specifically implemented as the following steps S2451~S2456:
[0072] Step S2451: Input the flattened texture vector into the nonlinear transformation layer. The nonlinear transformation layer includes a first nonlinear activation unit and a second nonlinear activation unit. The first nonlinear activation unit is used to nonlinearly enhance the positive feature components in the flattened texture vector, and the second nonlinear activation unit is used to nonlinearly suppress the negative feature components in the flattened texture vector.
[0073] The nonlinear transformation layer is used to perform nonlinear transformations on the flattened texture vector, and it contains a first nonlinear activation unit and a second nonlinear activation unit. The first and second nonlinear activation units are two different nonlinear activation functions, used to process the positive and negative feature components in the flattened texture vector differently, respectively. Positive feature components are the components with positive values in the flattened texture vector, and negative feature components are the components with negative values. When the flattened texture vector is input to the nonlinear transformation layer, each element of the flattened texture vector is input to the first and second nonlinear activation units respectively. The first nonlinear activation unit can use the ReLU function to nonlinearly enhance the positive feature components, that is, elements greater than 0 remain unchanged, and elements less than or equal to 0 are set to 0. The second nonlinear activation unit can use the Leaky ReLU function to nonlinearly suppress the negative feature components, that is, elements less than 0 are multiplied by a small coefficient, and elements greater than or equal to 0 remain unchanged.
[0074] Step S2452: The positive feature components in the flattened texture vector are processed by the first nonlinear activation unit to obtain the first transformation component.
[0075] The positive feature components in the flattened texture vector are processed by a first nonlinear activation unit. These positive feature components are input into the first nonlinear activation unit, and calculations are performed according to the function rules of the first nonlinear activation unit to obtain the first transformed component. For example, if the first nonlinear activation unit is a ReLU function, the positive feature components in the flattened texture vector are input into the ReLU function. The function will keep elements greater than 0 unchanged and set elements less than or equal to 0 to 0. The output result is the first transformed component.
[0076] Step S2453: The negative feature components in the flattened texture vector are processed by the second nonlinear activation unit to obtain the second transformation component.
[0077] The negative feature components in the flattened texture vector are processed by a second nonlinear activation unit. These negative feature components are input into the second nonlinear activation unit, and calculations are performed according to the function rules of the unit to obtain the second transformation component. For example, if the second nonlinear activation unit is a Leaky ReLU function, the negative feature components in the flattened texture vector are input into the Leaky ReLU function. The function multiplies elements less than 0 by a small coefficient, while elements greater than or equal to 0 remain unchanged. The output is the second transformation component.
[0078] Step S2454: Recombinate the first transformation component and the second transformation component to obtain the recombined feature vector.
[0079] Feature recombination is the process of merging and combining the first and second transformation components. Its purpose is to recombine the positive and negative feature components, which have undergone different processing, into a single vector. The recombined feature vector is the vector obtained after feature recombination, containing feature information after nonlinear enhancement and suppression processing.
[0080] When performing feature reconstruction, the first and second transformation components can be combined in the original order of the flattened texture vector to obtain the reconstructed feature vector. For example, arranging the first and second transformation components sequentially yields a new vector, which is the reconstructed feature vector.
[0081] Step S2455: Perform feature regularization on the recombined feature vector to reduce the influence of redundant feature components in the recombined feature vector, and obtain a regularized feature vector.
[0082] Feature regularization is the process of processing recombined feature vectors to reduce the influence of redundant feature components and improve the quality and expressive power of the features. Redundant feature components are those in the recombined feature vector that contribute little to the overall feature expression or contain repetitive information. Regularized feature vectors are the vectors obtained after feature regularization, where the influence of redundant feature components is reduced.
[0083] When performing feature regularization, methods such as L1 regularization or L2 regularization can be used. Taking L2 regularization as an example, the L2 norm of the reconstructed feature vector is calculated, and then each element of the reconstructed feature vector is divided by the L2 norm to obtain the regularized feature vector. This ensures that the elements of the reconstructed feature vector have the same scale, reducing the influence of redundant feature components.
[0084] Step S2456: Perform feature dimensionality reduction on the regularized feature vector, retain the feature components with higher contribution in the regularized feature vector, and obtain the region texture feature vector.
[0085] Feature dimensionality reduction is the process of processing regularized feature vectors to reduce their dimensionality while retaining the feature components with higher contribution. These feature components are those that play a crucial role in the overall feature representation. The region texture feature vector is the vector obtained after feature dimensionality reduction, containing only the feature components with higher contribution from the regularized feature vector. Methods such as Principal Component Analysis (PCA) can be used for feature dimensionality reduction. Taking PCA as an example, the regularized feature vectors are standardized to have a uniform scale and range. Then, the covariance matrix of the data is calculated. By solving for the eigenvalues and eigenvectors of the covariance matrix, the principal components of the data are obtained. The top N principal components with the largest eigenvalues are selected as the principal features. The regularized feature vector is projected onto these principal components to obtain the dimensionality-reduced feature vector, which is the region texture feature vector.
[0086] Step S250: Perform modal feature matching on the regional texture feature vector and the temporal stress feature vector, and establish a nonlinear correspondence between texture features and stress features through a dynamic mapping mechanism to obtain the matched texture feature vector and the matched stress feature vector.
[0087] Modal feature matching is the process of matching regional texture feature vectors with temporal stress feature vectors to find the intrinsic relationship between them and establish a correspondence between texture features and stress features. Dynamic mapping is a mechanism that can adaptively adjust the mapping relationship, establishing a nonlinear correspondence between texture features and stress features based on the characteristics and changes of the data. Matched texture feature vectors and matched stress feature vectors are vectors obtained after modal feature matching; they are the results of matching regional texture feature vectors and temporal stress feature vectors, respectively, reflecting the correspondence between texture features and stress features.
[0088] In modal feature matching, a modal feature mapping network is first constructed, which includes a texture-to-stress mapping branch and a stress-to-texture mapping branch. Then, the region texture feature vector is input to the texture-to-stress mapping branch, and its dynamic weight layer learns the mapping weights of the texture feature dimension to the stress feature dimension, resulting in a texture-mapped stress vector. The temporal stress feature vector is input to the stress-to-texture mapping branch, and its dynamic bias layer learns the mapping biases of the stress feature dimension to the texture feature dimension, resulting in a stress-mapped texture vector. Next, the first feature distance between the texture-mapped stress vector and the temporal stress feature vector, and the second feature distance between the stress-mapped texture vector and the region texture feature vector are calculated. The mapping parameters of the modal feature mapping network are adjusted based on these two feature distance values to ensure that both feature distance values meet a preset distance threshold. Finally, after the mapping parameters are adjusted, the region texture feature vector is passed through the adjusted stress-to-texture mapping branch to obtain a matching texture feature vector, and the temporal stress feature vector is passed through the adjusted texture-to-stress mapping branch to obtain a matching stress feature vector.
[0089] As one implementation method, step S250 can be specifically implemented as the following steps S251~S256:
[0090] Step S251: Construct a modal feature mapping network. The modal feature mapping network includes a texture-to-stress mapping branch and a stress-to-texture mapping branch. The texture-to-stress mapping branch is used to map texture features to stress feature space, and the stress-to-texture mapping branch is used to map stress features to texture feature space.
[0091] Modal Feature Mapping (MMR) networks are network models used for modal feature matching, comprising texture-to-stress mapping branches and stress-to-texture mapping branches. The texture-to-stress mapping branch is a submodule of the network whose function is to map texture features to the stress feature space, i.e., to find the correspondence between texture features and stress features, converting texture features into a stress feature representation. The stress-to-texture mapping branch is another submodule of the network whose function is to map stress features to the texture feature space, converting stress features into a texture feature representation.
[0092] When constructing modal feature mapping networks, neural network structures such as multilayer perceptrons (MLPs) can be used. The texture-to-stress mapping branch and the stress-to-texture mapping branch can each consist of multiple fully connected layers, each containing multiple neurons. Feature mapping is achieved through the connection weights and biases between neurons. For example, the texture-to-stress mapping branch can consist of three fully connected layers: the input layer receives the region's texture feature vector, the intermediate layers perform feature transformation, and the output layer outputs the texture-mapped stress vector.
[0093] Step S252: Input the region texture feature vector into the texture-to-stress mapping branch, and learn the mapping weight of the texture feature dimension to the stress feature dimension through the dynamic weight layer of the texture-to-stress mapping branch to obtain the texture-mapped stress vector.
[0094] The dynamic weight layer is a key layer in the texture-to-stress mapping branch. It learns the mapping weights between texture feature dimensions and stress feature dimensions. These mapping weights are numerical values representing the degree of influence of texture feature dimensions on stress feature dimensions. By learning these weights, the mapping relationship between texture features and stress features can be found. The texture-mapped stress vector is the vector obtained after processing by the texture-to-stress mapping branch; it is the mapping representation of the region texture feature vector in the stress feature space.
[0095] When learning the mapping weights, the dynamic weight layer adjusts based on the input region texture feature vector and the output target stress feature vector. Backpropagation can be used to update the weight parameters of the dynamic weight layer, making the texture-mapped stress vector as close as possible to the target stress feature vector. For example, the region texture feature vector is input into the dynamic weight layer of the texture-to-stress mapping branch. The dynamic weight layer calculates the error between the output texture-mapped stress vector and the target stress feature vector based on a preset loss function, and then updates the weight parameters through backpropagation, continuously adjusting the mapping weights until the error meets certain conditions.
[0096] Step S253: Input the temporal stress feature vector into the stress-to-texture mapping branch, and learn the mapping bias of the stress feature dimension to the texture feature dimension through the dynamic bias layer of the stress-to-texture mapping branch to obtain the stress-mapped texture vector.
[0097] The dynamic bias layer is a key layer in the stress-to-texture mapping branch, capable of learning the mapping bias between the stress feature dimension and the texture feature dimension. The mapping bias is a numerical value representing the degree of offset between the stress feature dimension and the texture feature dimension. By learning these biases, the mapping relationship between stress features and texture features can be found. The stress-mapped texture vector is the vector obtained after processing by the stress-to-texture mapping branch, and it is the mapping representation of the temporal stress feature vector in the texture feature space.
[0098] When learning the mapping bias, the dynamic bias layer adjusts based on the input temporal stress feature vector and the output target texture feature vector. Similarly, the backpropagation algorithm can be used to update the bias parameters of the dynamic bias layer, making the stress-mapped texture vector as close as possible to the target texture feature vector. For example, the temporal stress feature vector is input into the dynamic bias layer of the stress-to-texture mapping branch. The dynamic bias layer calculates the error between the output stress-mapped texture vector and the target texture feature vector based on a preset loss function, and then updates the bias parameters through backpropagation, continuously adjusting the mapping bias until the error meets certain conditions.
[0099] Step S254: Calculate the first feature distance between the texture mapping stress vector and the temporal stress feature vector, and calculate the second feature distance between the stress mapping texture vector and the region texture feature vector.
[0100] The first feature distance value represents the degree of difference between the texture-mapped stress vector and the temporal stress feature vector, reflecting how closely the texture features, after being mapped to the stress feature space, approximate the actual stress features. The second feature distance value represents the degree of difference between the stress-mapped texture vector and the region texture feature vector, reflecting how closely the stress features, after being mapped to the texture feature space, approximate the actual texture features.
[0101] When calculating feature distance values, methods such as cosine similarity can be used. Cosine similarity is a method for calculating the cosine of the angle between two vectors. The closer the cosine value is to 1, the more similar the two vectors are, and the smaller the feature distance value; the closer the cosine value is to -1, the less similar the two vectors are, and the larger the feature distance value is. By calculating the cosine similarity and converting it into feature distance values, we can obtain the first feature distance value and the second feature distance value.
[0102] As one implementation method, step S254 can be specifically implemented as the following steps S2541~S2544:
[0103] Step S2541: Unify the vector dimensions of the texture mapping stress vector and the temporal stress feature vector to make the number of dimensions of the two vectors consistent.
[0104] Vector dimension unification is the process of adjusting the number of dimensions of the texture mapping stress vector and the temporal stress feature vector to be consistent. When calculating feature distance, the two vectors must have the same number of dimensions; otherwise, an effective comparison cannot be performed.
[0105] When unifying vector dimensions, methods such as padding or truncation can be used. If the dimension of the texture mapping stress vector is less than the dimension of the temporal stress feature vector, zero vectors can be padded to the end of the texture mapping stress vector to make its dimension consistent with that of the temporal stress feature vector. If the dimension of the texture mapping stress vector is greater than the dimension of the temporal stress feature vector, the end of the texture mapping stress vector can be truncated to make its dimension consistent with that of the temporal stress feature vector.
[0106] Step S2542: Calculate the cosine similarity value between the texture mapping stress vector and the temporal stress feature vector after unifying the dimensions, and convert the cosine similarity value into the first feature distance value. The first feature distance value decreases as the cosine similarity value increases.
[0107] The cosine similarity value is a numerical value that represents the degree of similarity between two vectors. It is obtained by calculating the dot product of the two vectors and dividing it by the product of their magnitudes. Converting the cosine similarity value to the first feature distance value is the process of converting the similarity value to a distance value. The first feature distance value decreases as the cosine similarity value increases; that is, the more similar the two vectors are, the smaller the first feature distance value.
[0108] Step S2543: Unify the vector dimensions of the stress mapping texture vector and the region texture feature vector so that the number of dimensions of the two vectors is consistent.
[0109] Similar to step S2541, vector dimension unification is a process of adjusting the number of dimensions of the stress mapping texture vector and the region texture feature vector to be consistent in order to perform feature distance calculation.
[0110] When unifying vector dimensions, methods such as padding or truncation can also be used. The appropriate method for dimension adjustment is chosen based on the dimensionality relationship between the stress-mapped texture vector and the region texture feature vector. For example, if the number of dimensions of the stress-mapped texture vector is less than the number of dimensions of the region texture feature vector, zero vectors can be padded to the end of the stress-mapped texture vector to make its number of dimensions consistent with that of the region texture feature vector.
[0111] Step S2544: Calculate the cosine similarity value between the stress-mapped texture vector and the region texture feature vector after unifying the dimensions, and convert the cosine similarity value into the second feature distance value.
[0112] Similar to step S2542, the cosine similarity value between the stress-mapped texture vector and the region texture feature vector after unifying the dimensions is calculated, and then converted into the second feature distance value.
[0113] When calculating the cosine similarity value, the cosine similarity formula is used. The second feature distance value is obtained by subtracting 1 from the calculated cosine similarity value and taking the absolute value.
[0114] Step S255: Adjust the mapping parameters of the modal feature mapping network based on the first feature distance value and the second feature distance value so that both feature distance values meet the preset distance threshold.
[0115] The preset distance threshold is a pre-defined threshold value for feature distance, used to determine whether the mapping effect of the modal feature mapping network meets the requirements. Adjusting the mapping parameters of the modal feature mapping network based on the first and second feature distance values involves continuously adjusting the network's weights and bias parameters using an optimization algorithm, ensuring that both the first and second feature distance values are less than the preset distance threshold.
[0116] When adjusting the mapping parameters, optimization algorithms such as gradient descent can be used. A loss function is calculated based on the first and second feature distance values; the loss function can be a weighted sum of the two feature distance values. Then, the gradient descent algorithm is used to calculate the gradient of the loss function with respect to the mapping parameters. The mapping parameters are updated based on the gradient, and this process is iterated until both feature distance values meet a preset distance threshold.
[0117] Step S256: After the mapping parameters are adjusted, the region texture feature vector is passed through the adjusted stress-to-texture mapping branch to obtain the matching texture feature vector, and the temporal stress feature vector is passed through the adjusted texture-to-stress mapping branch to obtain the matching stress feature vector.
[0118] Once the mapping parameters of the modal feature mapping network are adjusted, it indicates that the network has learned a good mapping relationship between texture features and stress features. The region texture feature vector is then passed through the adjusted stress-to-texture mapping branch. This branch maps the region texture feature vector to a more suitable position in the texture feature space based on the learned mapping relationship, resulting in a matched texture feature vector. Similarly, the temporal stress feature vector is passed through the adjusted texture-to-stress mapping branch. This branch maps the temporal stress feature vector to a more suitable position in the stress feature space based on the learned mapping relationship, resulting in a matched stress feature vector.
[0119] The texture feature space is a multidimensional space composed of texture features extracted from road surface image data, while the stress feature space is a multidimensional space composed of stress features extracted from road structure sensing data. Each feature vector represents a point in its corresponding feature space, and there is an inherent physical relationship between texture features and stress features; for example, surface crack textures often correspond to stress concentrations within the road structure. When the modal feature mapping network learns the mapping relationship between texture features and stress features, it uses quantitative metrics such as cosine similarity to measure the degree of correspondence between the two.
[0120] Initially, the position of a region's texture feature vector in the texture feature space may result in a low cosine similarity to its corresponding stress feature vector. After adjustment via the stress-to-texture mapping branch, a "more suitable position in the texture feature space" refers to the region's texture feature vector moving to a new position where the cosine similarity between the texture feature vector and its corresponding stress feature vector significantly increases. This means they are closer in the feature space direction, representing a stronger correlation between the texture and stress features. From a physical perspective, different texture features on the road surface are external manifestations of the road's internal stress state. A more suitable position in the texture feature space makes the texture represented by the feature vector more physically aligned with the actual corresponding stress state. For example, a feature vector that originally represented a seemingly normal texture but actually concealing potential stress problems will, after mapping, be closer in the texture feature space to a texture feature vector representing stress anomalies, better reflecting its correspondence with potential stress problems.
[0121] Similarly, initially, the temporal stress feature vector's position in the stress feature space might result in a low cosine similarity to the corresponding texture feature vector. After the texture-to-stress mapping branch is adjusted, a "more suitable position in the stress feature space" refers to the temporal stress feature vector moving to a new position, significantly increasing its cosine similarity to the corresponding texture feature vector. This indicates that the stress feature vector and texture feature vector are more aligned in direction in the feature space, demonstrating a stronger correspondence.
[0122] In road infrastructure structures, stress state is the cause of surface texture changes. A more suitable location in the stress feature space allows the stress feature vector to better align with the physical causal relationship between the stress state and the actual texture changes it causes. For example, a stress feature vector representing a high stress growth trend, after mapping, will be closer in the stress feature space to the region corresponding to the texture feature vector showing cracks, accurately reflecting the causal relationship between stress and texture changes. The modal feature mapping network continuously adjusts the mapping parameters through dynamic weight layers and dynamic bias layers, enabling texture and stress feature vectors to find the aforementioned more suitable locations in their respective feature spaces. This adjusted location not only reflects a better correspondence in quantitative indicators but also more accurately reflects the intrinsic relationship between road surface texture and internal stress, providing a more reliable foundation for cross-modal feature fusion and road infrastructure state inference.
[0123] Step S260: Perform modal interaction fusion on the matched texture feature vector and the matched stress feature vector to obtain a cross-modal fusion feature vector.
[0124] Modal interaction fusion is the process of interacting and fusing matching texture feature vectors and matching stress feature vectors. Its purpose is to integrate feature information from different modalities and extract more valuable comprehensive information. The cross-modal fusion feature vector is the vector obtained after modal interaction fusion, containing comprehensive information from both matching texture and matching stress feature vectors.
[0125] When performing modal interaction fusion, splicing can be used to connect the matching texture feature vector and the matching stress feature vector in sequence to form a new vector, which is the cross-modal fusion feature vector.
[0126] Step S270: Arrange the cross-modal fusion feature vectors sequentially according to the sample order in the road production data set to generate a cross-modal correlation feature matrix.
[0127] The cross-modal correlation feature matrix is obtained by arranging the cross-modal fused feature vectors in the order of the samples in the road asset dataset. Each sample in the road asset dataset corresponds to a cross-modal fused feature vector. Arranging these vectors in the order of the samples forms the cross-modal correlation feature matrix.
[0128] When generating the cross-modal association feature matrix, the sample order in the road production dataset is first determined. Then, the cross-modal fusion feature vector corresponding to each sample is added to the rows of the matrix in this order, finally obtaining the cross-modal association feature matrix.
[0129] Step S300: Construct a feature evolution dependency network based on the cross-modal correlation feature matrix. The feature evolution dependency network is used to characterize the dynamic dependency relationship of cross-modal feature dimensions over time and the coupling mechanism between features.
[0130] Feature evolution dependency networks (FDCs) are network models used to represent the dynamic dependencies between cross-modal feature dimensions over time and the coupling mechanisms between features. Cross-modal feature dimensions are the feature dimensions represented by the column vectors in the cross-modal associated feature matrix; these feature dimensions may have interdependent and coupled relationships at different points in time. FDCs represent these dependencies and coupling relationships through network structure and the connections between nodes.
[0131] When constructing a feature evolution dependency network, it is first necessary to analyze and process the cross-modal correlation feature matrix to uncover the dependencies and coupling mechanisms between feature dimensions. Then, based on these relationships, a network model is constructed, and the network nodes and edges are determined. For example, network nodes can be cross-modal feature dimensions, and edges can represent the dependencies and coupling strengths between feature dimensions.
[0132] As one implementation method, step S300 can be specifically implemented as the following steps S310~S360:
[0133] Step S310: Perform feature dimension analysis on the cross-modal correlation feature matrix to obtain the image source feature sub-matrix and the sensor source feature sub-matrix. The image source feature sub-matrix is composed of the feature dimension column vectors from the road surface image data in the cross-modal correlation feature matrix, and the sensor source feature sub-matrix is composed of the feature dimension column vectors from the road structure sensor data in the cross-modal correlation feature matrix.
[0134] Feature dimension parsing is the process of analyzing and decomposing the cross-modal correlation feature matrix, aiming to divide the feature dimensions in the matrix according to the data source. The image source feature submatrix is a submatrix composed of column vectors of feature dimensions from road surface image data in the cross-modal correlation feature matrix, containing the feature information of the road surface image data. The sensor source feature submatrix is a submatrix composed of column vectors of feature dimensions from road structure sensor data in the cross-modal correlation feature matrix, containing the feature information of the road structure sensor data.
[0135] When performing feature dimension analysis, it is necessary to divide the features based on their source information. During the generation of the cross-modal correlation feature matrix, the source of each feature dimension (road property surface image data or road property structure sensor data) is recorded. Based on these records, the feature dimension column vectors from the road property surface image data are extracted to form the image source feature sub-matrix; the feature dimension column vectors from the road property structure sensor data are extracted to form the sensor source feature sub-matrix.
[0136] Step S320: Construct a dual-path collaborative dependency modeling network, which includes an image-guided dependency path and a sensor-guided dependency path. The image-guided dependency path is used to learn the dependencies of cross-modal features based on image source features, and the sensor-guided dependency path is used to learn the dependencies of cross-modal features based on sensor source features.
[0137] The dual-path collaborative dependency modeling network is a network model used to learn cross-modal feature dependencies, comprising an image-guided dependency path and a sensor-guided dependency path. The image-guided dependency path learns cross-modal feature dependencies based on image source features, uncovering dependencies by analyzing the impact of image source features on other feature dimensions. The sensor-guided dependency path learns cross-modal feature dependencies based on sensor source features, uncovering dependencies by analyzing the impact of sensor source features on other feature dimensions.
[0138] When constructing a dual-path collaborative dependency modeling network, a neural network architecture can be used. The image-guided dependency path and the sensor-guided dependency path can each consist of multiple fully connected layers and attention mechanisms. For example, the image-guided dependency path can include a feature attention mechanism layer to calculate the influence weights of the image source feature dimension on other feature dimensions; the sensor-guided dependency path can include a feature gating mechanism layer to calculate the control coefficients of the sensor source feature dimension on other feature dimensions.
[0139] Step S330: Input the image source feature sub-matrix into the image guidance dependency path, calculate the influence weight of each image source feature dimension on other feature dimensions through the feature attention mechanism of the image guidance dependency path, obtain the image guidance weight matrix, and perform weighted modulation on the cross-modal correlation feature matrix based on the image guidance weight matrix to obtain the image modulation feature matrix.
[0140] The feature attention mechanism is a key component in image-guided dependency paths, capable of calculating the influence weights of each image source feature dimension on other feature dimensions. Influence weights are numerical values representing the degree of influence of an image source feature dimension on other feature dimensions. By calculating these weights, the dependencies between image source features and other feature dimensions can be identified. The image-guided weight matrix is a matrix obtained after calculation by the feature attention mechanism, with each row representing the influence weight of one image source feature dimension on other feature dimensions. Weighted modulation of the cross-modal correlation feature matrix based on the image-guided weight matrix involves element-wise multiplication of the image-guided weight matrix and the cross-modal correlation feature matrix to obtain the image-modulated feature matrix. The image-modulated feature matrix, obtained after weighted modulation, highlights the influence of image source feature dimensions on other feature dimensions.
[0141] When calculating the influence weights, the feature attention mechanism first extracts the set of image source feature dimensions from the image source feature submatrix. Then, for each image source feature dimension, it constructs a correlation evaluation sequence between that feature dimension and all other feature dimensions in the cross-modal correlation feature matrix. Next, the correlation evaluation sequence is normalized, and the normalized sequence is input into the weight generation unit of the feature attention mechanism. This unit performs a non-linear transformation on the sequence to obtain the influence weight vector corresponding to that image source feature dimension. Finally, the influence weight vectors corresponding to all image source feature dimensions are arranged in order to generate the image guidance weight matrix.
[0142] As one implementation method, step S330 can be specifically implemented as the following steps S331-S335:
[0143] Step S331: Extract the image source feature dimension set from the image source feature submatrix. The image source feature dimension set contains the feature dimension identifiers corresponding to all column vectors of the image source feature submatrix.
[0144] The image source feature dimension set is a collection of feature dimension identifiers corresponding to all column vectors of the image source feature submatrix. Each feature dimension identifier is information used to uniquely identify each feature dimension; it can be a number, a name, or other distinctive identifier.
[0145] When extracting the feature dimension set of the image source, iterate through each column of the image source feature submatrix and extract the feature dimension identifier corresponding to each column to form a set.
[0146] Step S332: For each image source feature dimension in the image source feature dimension set, construct a correlation evaluation sequence between the image source feature dimension and all other feature dimensions in the cross-modal correlation feature matrix. Each element of the correlation evaluation sequence is the correlation value between the column vector of the image source feature dimension and the column vector of the corresponding other feature dimensions.
[0147] The correlation evaluation sequence is used to assess the degree of correlation between a feature dimension of an image source and other feature dimensions in a cross-modal correlation feature matrix. The correlation value is a numerical value representing the degree of correlation between two feature dimension column vectors, which can be obtained by calculating indicators such as correlation and similarity between the two column vectors.
[0148] When constructing the correlation evaluation sequence, for each image source feature dimension in the image source feature dimension set, the correlation degree needs to be calculated with all other feature dimensions in the cross-modal correlation feature matrix except itself. Specifically, first, an image source feature dimension is selected from the image source feature dimension set as the current evaluation dimension, and the current feature column vector corresponding to the current evaluation dimension is extracted from the cross-modal correlation feature matrix. Then, all feature dimensions in the cross-modal correlation feature matrix are traversed, and each feature dimension except the current evaluation dimension is used as a comparison evaluation dimension, and the comparison feature column vector corresponding to the comparison evaluation dimension is extracted. The correlation degree value between the current feature column vector and the comparison feature column vector is calculated, and the calculated correlation degree values are arranged in the order of traversal of the comparison evaluation dimensions to obtain the correlation evaluation sequence corresponding to the current evaluation dimension.
[0149] As one implementation method, step S332 can be specifically implemented as the following steps S3321-S3327:
[0150] Step S3321: Extract the spatial coordinate range of each node from the road property spatial association map to obtain the node coordinate range set.
[0151] In constructing the correlation assessment sequence, relevant basic information is first obtained from the road asset spatial correlation map. The road asset spatial correlation map stores the spatial information of road asset entity segments, and the set of node coordinate ranges is a set composed of the spatial coordinate ranges corresponding to each node in the map. The spatial coordinate range can be represented by a coordinate interval, for example, it can be determined by the coordinates of the lower left and upper right corners of a rectangular area on a two-dimensional plane.
[0152] When extracting the set of node coordinate ranges, this can be achieved by querying the data structure of the road property spatial association map. If the road property spatial association map is stored in a graph database, each node has a corresponding attribute field recording its spatial coordinate range. A query statement can be written to extract the spatial coordinate ranges of all nodes and form a set. For example, for a road property spatial association map with multiple road segment nodes, by querying the attributes of these nodes, the spatial coordinate range corresponding to each node can be obtained. For example, the coordinate range of node 1 is [(x1, y1), (x2, y2)]. The coordinate ranges of all nodes can then be summarized into a set of node coordinate ranges.
[0153] Step S3322: For each sensor data acquisition location in the original acquisition information, extract the coordinate information of that sensor data acquisition location.
[0154] The coordinates of the sensor data acquisition locations are crucial for determining their spatial location, and can be represented using latitude and longitude or Cartesian coordinates. When constructing the correlation assessment sequence, it is necessary to obtain the coordinates of each sensor data acquisition location for subsequent processing. Coordinate information can be directly read from the raw acquisition data. The raw acquisition data records detailed information about each acquisition location during the sensor data acquisition process, including coordinate information.
[0155] Step S3323: Traverse each node coordinate range in the set of node coordinate ranges and determine whether the coordinate information falls within the coordinate range of that node.
[0156] Traversing the set of node coordinate ranges involves checking the coordinate range of each node in the set sequentially. Determining whether coordinate information falls within a node's coordinate range involves comparing the coordinates of the sensor data acquisition location with the node's coordinate range. If the coordinate information is within the range defined by the node's coordinate range, it is considered to fall within that range.
[0157] When making a judgment, the coordinate information on the two-dimensional plane and the coordinate range of the rectangular node can be determined by comparing the horizontal and vertical coordinate values. For example, it can be determined whether the horizontal coordinate of the sensor data acquisition position is within the horizontal coordinate interval of the node coordinate range, and whether the vertical coordinate is within the vertical coordinate interval. If both conditions are met, the coordinate information is determined to fall within the coordinate range of that node.
[0158] Step S3324: When the coordinate information falls within the coordinate range of a certain node, determine the road property image area corresponding to the coordinate range of that node as the road property image area to which the sensor data acquisition location belongs.
[0159] When the coordinates of the sensor data acquisition location fall within the coordinate range of a certain node, it indicates that the sensor data acquisition location and the road property entity segment corresponding to the node are spatially related. Therefore, it can be determined that the road property image area corresponding to the coordinate range of the node is the road property image area to which the sensor data acquisition location belongs.
[0160] Step S3325: Record the correspondence between the sensor data acquisition location and the corresponding road property image area as a spatial association unit.
[0161] The spatial association unit is a basic unit used to record the correspondence between the sensor data acquisition location and the corresponding road property image area. It can be represented by a set data structure, such as a tuple (sensor data acquisition location, road property image area).
[0162] When recording spatial association units, the coordinate information of the sensor data acquisition location and the identification information of the road property image area are combined together. For example, in a database table that specifically stores road property data, two fields are created: one to record the coordinates of the sensor data acquisition location and the other to record the identification of the road property image area. These are combined into a single record, which is a spatial association unit.
[0163] Step S3326: Summarize the spatial association units corresponding to all sensor data acquisition locations to obtain the spatial association results.
[0164] The spatial correlation result is a summary of the spatial correlation units corresponding to all sensor data acquisition locations, which can be presented in the form of a list or set.
[0165] When summarizing spatial correlation units, the spatial correlation units corresponding to each sensor data acquisition location are added sequentially to a list or set. For example, for multiple sensor data acquisition locations in a road asset data set, the spatial correlation units corresponding to each location, such as (location 1 coordinates, image region 1 identifier), (location 2 coordinates, image region 2 identifier), etc., are added sequentially to a list, and this list is the final spatial correlation result.
[0166] Step S3327: Update the next image source feature dimension in the image source feature dimension set to the current evaluation dimension, and return the step of extracting the current feature column vector corresponding to the current evaluation dimension from the cross-modal correlation feature matrix until all image source feature dimensions in the image source feature dimension set have been traversed.
[0167] When constructing the correlation evaluation sequence, it is necessary to process each image source feature dimension in the image source feature dimension set. By continuously updating the current evaluation dimension, the correlation degree is calculated for each dimension in turn, until all dimensions have been traversed.
[0168] Specifically, after constructing the correlation evaluation sequence for one image source feature dimension, the next dimension in the image source feature dimension set is used as the new current evaluation dimension. Then, the process returns to the step of extracting the current feature column vector corresponding to the current evaluation dimension from the cross-modal correlation feature matrix, and repeats the subsequent calculation process.
[0169] Step S333: Normalize the correlation evaluation sequence to obtain a normalized correlation sequence. The sum of all elements in the normalized correlation sequence is the preset normalization benchmark value.
[0170] Normalization is an operation that processes a correlation evaluation sequence to give its elements a uniform scale and range. A normalized correlation sequence is a sequence obtained after normalization, where the sum of all its elements is a preset normalization baseline value, typically set to 1 for ease of subsequent calculations and comparisons. Various methods can be used for normalization, such as linear normalization. Linear normalization first calculates the sum of all elements in the correlation evaluation sequence, then divides each element in the sequence by this sum to obtain the normalized element value.
[0171] Step S334: Input the normalized correlation sequence into the weight generation unit of the feature attention mechanism, and perform a nonlinear transformation on the normalized correlation sequence through the weight generation unit to obtain the influence weight vector corresponding to the feature dimension of the image source.
[0172] The weight generation unit is a crucial component of the feature attention mechanism. It performs a non-linear transformation on the normalized correlation sequence, converting it into an influence weight vector. The influence weight vector represents the degree of influence of one image source feature dimension on other feature dimensions.
[0173] When performing nonlinear transformations, the weight generation unit can use activation functions such as the Sigmoid function and the ReLU function. Taking the Sigmoid function as an example, each element in the normalized correlation sequence is input into the Sigmoid function, which maps the input value to the interval (0, 1), obtaining the transformed element value. These transformed element values are then combined into a vector, which is the influence weight vector corresponding to the feature dimension of the image source.
[0174] Step S335: Arrange the influence weight vectors corresponding to all image source feature dimensions in the image source feature dimension set in order to generate the image guided weight matrix.
[0175] The image guidance weight matrix is a matrix composed of the influence weight vectors corresponding to all image source feature dimensions arranged in order. Each influence weight vector corresponds to a row of the image guidance weight matrix, and by arranging these vectors in sequence, a complete matrix is formed.
[0176] Step S340: Input the sensor source feature submatrix into the sensor guidance dependency path, calculate the control coefficients of each sensor source feature dimension on other feature dimensions through the feature gating mechanism of the sensor guidance dependency path, obtain the sensor control coefficient matrix, and perform gating screening on the cross-modal correlation feature matrix based on the sensor control coefficient matrix to obtain the sensor screening feature matrix.
[0177] The sensor-guided dependency path is one path in a dual-path collaborative dependency modeling network, used to learn cross-modal feature dependencies based on sensor source features. The feature gating mechanism is a key component of the sensor-guided dependency path, capable of calculating the control coefficients of each sensor source feature dimension on other feature dimensions. These control coefficients represent the degree of control a sensor source feature dimension has over other feature dimensions; they can be used to filter out sensor source features that have a significant impact on other feature dimensions.
[0178] The sensor control coefficient matrix is a matrix obtained after calculation using a feature gating mechanism. Each row represents the control coefficient of one sensor source feature dimension on other feature dimensions. Gating and filtering the cross-modal correlation feature matrix based on the sensor control coefficient matrix involves element-wise multiplication or other filtering operations between the sensor control coefficient matrix and the cross-modal correlation feature matrix to obtain the sensor filtering feature matrix. The sensor filtering feature matrix, obtained after filtering, highlights the control effect of the sensor source feature dimension on other feature dimensions.
[0179] When calculating the control coefficients, the feature gating mechanism first extracts the sequence of sensor source feature dimensions from the sensor source feature submatrix. Then, for each sensor source feature dimension, it calculates the feature fluctuation value of that feature dimension in the cross-modal correlation feature matrix. Next, the feature fluctuation value is input into the reset gate unit and update gate unit of the feature gating mechanism to obtain the reset gate value and update gate value, respectively. Finally, based on the reset gate value and update gate value, the comprehensive control coefficient of that sensor source feature dimension with respect to other feature dimensions is calculated, and the control coefficient vectors corresponding to all sensor source feature dimensions are arranged in order to generate the sensor control coefficient matrix.
[0180] As one implementation method, step S340 can be specifically implemented as the following steps S341-S346:
[0181] Step S341: Extract the sensor source feature dimension sequence from the sensor source feature submatrix. The sensor source feature dimension sequence is the sequence of feature dimensions corresponding to the column vectors of the sensor source feature submatrix arranged in a preset logical order.
[0182] The sensor source feature dimension sequence is a sequence composed of the feature dimensions corresponding to the column vectors of the sensor source feature submatrix arranged in a preset logical order. The preset logical order can be arranged according to the feature dimension number, importance, etc. When extracting the sensor source feature dimension sequence, each column of the sensor source feature submatrix is traversed, the feature dimension identifier corresponding to each column is extracted, and arranged in the preset logical order.
[0183] Step S342: For each sensor source feature dimension in the sensor source feature dimension sequence, calculate the feature fluctuation value of that sensor source feature dimension in the cross-modal correlation feature matrix. The feature fluctuation value is used to characterize the discreteness of the column vector element values of that sensor source feature dimension.
[0184] Feature fluctuation is an indicator used to measure the dispersion of the column vector elements of a sensor source feature dimension. It reflects the variation of that feature dimension in the cross-modal correlation feature matrix. When calculating feature fluctuation, statistics such as variance and standard deviation can be used. Taking variance as an example, for a sensor source feature dimension in the sequence of sensor source feature dimensions, its corresponding column vector in the cross-modal correlation feature matrix is calculated as follows: first, the mean of the column vector is calculated; then, the square of the difference between each element and the mean is calculated; finally, the average of these squared values is taken. The result is the feature fluctuation value of that sensor source feature dimension.
[0185] Step S343: Input the feature fluctuation value into the reset gate unit of the feature gating mechanism. The reset gate unit performs a gating threshold judgment on the feature fluctuation value to obtain the reset gate value. The reset gate value is used to indicate whether to retain the control effect of the feature dimension of the sensing source on other feature dimensions.
[0186] The reset gate unit is a component of the feature gating mechanism, capable of determining a gating threshold for feature fluctuation values. The gating threshold is a pre-set value used to determine whether to retain the control effect of one sensor source feature dimension on other feature dimensions. During the gating threshold determination, the feature fluctuation value is compared to the gating threshold. If the feature fluctuation value is greater than the gating threshold, it indicates that the change in that sensor source feature dimension is significant and may have a substantial impact on other feature dimensions; the reset gate value is set to 1, indicating that the control effect of that sensor source feature dimension on other feature dimensions is retained. If the feature fluctuation value is less than or equal to the gating threshold, it indicates that the change in that sensor source feature dimension is small and its impact on other feature dimensions may be minor; the reset gate value is set to 0, indicating that the control effect of that sensor source feature dimension on other feature dimensions is not retained.
[0187] Step S344: Input the feature fluctuation value into the update gate unit of the feature gating mechanism, and dynamically adjust the range of the feature fluctuation value through the update gate unit to obtain the update gate value. The update gate value is used to indicate the control strength of the feature dimension of the sensing source on other feature dimensions.
[0188] The update gate unit is another component in the feature gating mechanism, capable of dynamically adjusting the range of feature fluctuation values. Dynamic range adjustment maps feature fluctuation values to a suitable range to obtain the update gate value, which represents the control strength of this sensor feature dimension over other feature dimensions.
[0189] When adjusting the dynamic range, activation functions such as the Sigmoid function and the Tanh function can be used. Taking the Sigmoid function as an example, the feature fluctuation value is input into the Sigmoid function, which maps the input value to the (0, 1) interval. The output value is the updated gate value.
[0190] Step S345: Calculate the comprehensive control coefficient of the sensor source feature dimension on other feature dimensions based on the reset gate value and the update gate value, and obtain the control coefficient vector.
[0191] The comprehensive control coefficient is a coefficient calculated by combining the reset gate value and the update gate value, which comprehensively reflects the degree of control of one sensor source feature dimension over other feature dimensions.
[0192] When calculating the overall control coefficient, the reset gate value and the update gate value can be multiplied together to obtain the overall control coefficient. The control coefficient vector is formed by combining the overall control coefficients of each sensor feature dimension with respect to other feature dimensions.
[0193] Step S346: Arrange the control coefficient vectors corresponding to all sensor source feature dimensions in the sensor source feature dimension sequence in order to generate a sensor control coefficient matrix.
[0194] The sensor control coefficient matrix is a matrix composed of control coefficient vectors corresponding to all feature dimensions of the sensor sources, arranged in order. Each control coefficient vector corresponds to a row in the sensor control coefficient matrix, and these vectors are arranged sequentially to form a complete matrix.
[0195] Step S350: Perform feature interaction enhancement on the image modulation feature matrix and the sensor screening feature matrix to obtain the cooperative dependency feature matrix.
[0196] Feature interaction enhancement is the process of interacting and enhancing the image modulation feature matrix and the sensor screening feature matrix. Its purpose is to uncover the relationships between features in the two matrices and further improve the expressive power of the features. The co-dependency feature matrix is the matrix obtained after feature interaction enhancement, containing comprehensive information from both the image modulation feature matrix and the sensor screening feature matrix. During feature interaction enhancement, the structures of the two matrices must first be unified to ensure they have the same number of rows and columns. Then, element-wise interaction calculations are performed on the unified matrix to obtain a preliminary interaction matrix. Next, feature sparsification is applied to the preliminary interaction matrix, adjusting the element values that meet preset sparsity conditions to preset baseline values, resulting in a sparse interaction matrix. Finally, row-direction feature normalization is performed on the sparse interaction matrix to ensure that the energy values of each row's elements meet preset energy distribution conditions, resulting in the co-dependency feature matrix.
[0197] As one implementation method, step S350 can be specifically implemented as the following steps S351-S356:
[0198] Step S351: Obtain the matrix structure parameters of the image modulation feature matrix and the sensor screening feature matrix. The matrix structure parameters include the number of matrix rows and the number of matrix columns.
[0199] Matrix structure parameters describe the basic structure of a matrix, including the number of rows and columns. When performing feature interaction enhancement, it is necessary to obtain the matrix structure parameters of the image modulation feature matrix and the sensor screening feature matrix.
[0200] When obtaining matrix structure parameters, they can be read directly from the matrix's data structure. For example, in programming languages, matrices are stored as two-dimensional arrays, and the number of rows and columns can be obtained by accessing the array's length attribute. For instance, the image modulation feature matrix is a two-dimensional array image_modulated_matrix, and the number of rows can be obtained through image_modulated_matrix.length, and the number of columns can be obtained through image_modulated_matrix[0].length.
[0201] Step S352: When there is a difference between the number of rows in the image modulation feature matrix and the number of rows in the sensor screening feature matrix, the matrix with fewer rows is expanded in row dimension by supplementing the row vectors through feature interpolation so that the number of rows in the expanded matrix is the same as the number of rows in the other matrix.
[0202] Row dimension expansion is an operation performed on the matrix with fewer rows when the number of rows in the image modulation feature matrix and the sensor screening feature matrix are inconsistent. Feature interpolation is a method used to supplement row vectors, estimating and generating new row vectors using known row vector information.
[0203] When performing row dimension expansion, if the number of rows in the image modulation feature matrix is less than the number of rows in the sensor screening feature matrix, the image modulation feature matrix is expanded in row dimension. A linear interpolation method can be used to generate new row vectors based on information from adjacent row vectors.
[0204] Step S353: When there is a difference between the number of columns in the image modulation feature matrix and the number of columns in the sensor screening feature matrix, the matrix with fewer columns is expanded in column dimension by supplementing column vectors through feature mapping so that the number of columns in the expanded matrix is the same as the number of columns in the other matrix.
[0205] Column dimension expansion is an operation performed on the matrix with fewer columns when the number of columns in the image modulation feature matrix and the sensor screening feature matrix are inconsistent. Feature mapping is a method used to supplement column vectors, generating new column vectors using known column vector information and mapping relationships.
[0206] When expanding the column dimension, if the number of columns in the image modulation feature matrix is less than the number of columns in the sensor screening feature matrix, the image modulation feature matrix is expanded in terms of column dimension. Linear or nonlinear mapping methods can be used to generate new column vectors based on known column vector information and preset mapping relationships.
[0207] Step S354: Perform element-wise interactive calculations on the unified image modulation feature matrix and the sensor screening feature matrix to obtain a preliminary interactive matrix. Each element of the preliminary interactive matrix is the result of the interactive operation between the corresponding elements of the two matrices.
[0208] Element-wise interactive computation is the process of performing interactive operations on corresponding elements of the image modulation feature matrix and the sensor screening feature matrix after structural unification. Interactive operations can include addition, multiplication, subtraction, etc.
[0209] During element-wise interactive computation, each corresponding element of the two matrices is traversed and interactively operated on. For example, for the image modulation feature matrix and the sensor screening feature matrix after structural unification, their corresponding elements are a and b, respectively. If multiplication is chosen for the interactive operation, then the corresponding elements of the initial interactive matrix are a*b. Combining the results of the interactive operations on all corresponding elements yields the initial interactive matrix.
[0210] Step S355: Perform feature sparsification on the preliminary interaction matrix, and adjust the element values that meet the preset sparsity conditions to preset baseline values to obtain a sparse interaction matrix.
[0211] Feature sparsity is a process of processing the initial interaction matrix. Its purpose is to adjust the values of elements that satisfy a preset sparsity condition to a preset baseline value, making the matrix more sparse and highlighting important feature information. The preset sparsity condition can be that the element value is less than a certain threshold; the preset baseline value is usually set to 0. During feature sparsity, each element of the initial interaction matrix is traversed, and its value is compared with the preset threshold. If the element value is less than the preset threshold, it is adjusted to the preset baseline value. After adjusting all the element values that satisfy the sparsity condition, a sparse interaction matrix is obtained.
[0212] Step S356: Perform row direction feature normalization on the sparse interaction matrix so that the energy value of each row element satisfies the preset energy distribution condition, and obtain the cooperative dependency feature matrix.
[0213] Row direction feature normalization is the process of normalizing each row element of a sparse interaction matrix to ensure that the energy value of each row element meets a preset energy distribution condition. This preset condition could be that the sum of squares of each row element equals 1. During row direction feature normalization, each row element of the sparse interaction matrix is traversed, and the sum of squares of each row element is calculated. Then, each row element is divided by the square root of the sum of squares of that row element, ensuring that the sum of squares of each row element equals 1.
[0214] Step S360: Input the collaborative dependency feature matrix into the dynamic dependency learning layer of the dual-path collaborative dependency modeling network. The dynamic dependency learning layer learns the dependency strength and direction between feature dimensions as the sample sequence changes, and generates a feature evolution dependency network.
[0215] Dynamic dependency learning layers can learn the strength and direction of dependencies between feature dimensions as they change with the sample sequence. Dependency strength represents the degree of mutual influence between feature dimensions, and dependency direction represents the direction of influence between feature dimensions.
[0216] When learning the strength and direction of dependencies, dynamic dependency learning layers can use models such as Recurrent Neural Networks (RNNs) and Long Short-Term Memory Networks (LSTMs). Taking LSTM as an example, the co-dependency feature matrix is input into the LSTM network according to the sample sequence. The LSTM network learns the dependencies between feature dimensions based on the input feature sequence. Through network training and learning, the dependency strength and direction information between each feature dimension and other feature dimensions are obtained. Representing this information in the form of a graph generates the feature evolution dependency network.
[0217] Step S400: Output the road state reasoning result through the feature evolution dependency network.
[0218] The Feature Evolution Dependency Network (FEDN) is a network constructed and learned through the preceding steps. It is capable of inferring the state of road assets based on the input feature information. The road asset state inference result is the output of the FEDN, which includes information on the current structural integrity category of the road assets and the potential risk evolution direction.
[0219] The current structural integrity of road assets can be categorized into intact, slightly damaged, moderately damaged, and severely damaged categories to describe their current structural state. Information on potential risk evolution indicates the possible future direction of risk development for the road assets, such as risk escalation or mitigation. When performing comprehensive reasoning about the road asset state, the dependencies and strengths between feature dimensions learned by the feature evolution dependency network are applied to the judgment of the road asset state. By analyzing the changing trends and dependencies of features, the current structural integrity category of the road assets is determined. Simultaneously, based on the dynamic changes of features and the development of dependencies, the evolution direction of potential risks to the road assets is predicted.
[0220] Specifically, the Feature Evolution Dependency Network (FEDN) integrates cross-modal feature information from road surface image data and road structure sensor data based on a cross-modal correlation feature matrix. This allows it to characterize the dynamic dependencies between cross-modal feature dimensions over time and the coupling mechanisms between features. During comprehensive road condition reasoning, the feature information contained in the current cross-modal correlation feature matrix is first input into the FEDN. This feature information includes various aspects such as the texture features of the road surface and the internal stress-strain features. Because the FEDN learns the complex dependencies between image source features and sensor source features during its construction, it can comprehensively consider the texture features of the road surface and the internal stress-strain features. Based on the previously learned dependencies between feature dimensions, the FEDN maps and transforms the input features, fusing feature information from different modalities. Next, the FEDN analyzes the state of the input features, comparing the current feature vector with pre-defined feature templates for different structural integrity categories. These feature templates are learned during network training based on a large number of road data samples with known structural integrity categories. For example, for road properties classified as "intact," the feature template might show a uniform surface texture and internal stress and strain within a normal range; while for road properties classified as "slightly damaged," the surface might show a few fine cracks, and internal stress and strain might exhibit slight abnormal fluctuations. Based on the results of feature state analysis, feature evolution relies on the network using a classifier (such as a Softmax classifier) to make a classification decision on the current structural integrity of the road property. The classifier calculates a score based on the similarity between the feature vector and each feature template, and the category with the highest score is the current structural integrity category of the road property. For example, if the current feature vector has the highest similarity to the feature template of the "slightly damaged" category, then the network will determine that the current structural integrity category of the road property is slightly damaged.
[0221] In reasoning about the evolutionary direction of potential risks, the dynamic dependencies of feature dimensions over time and the coupling mechanisms between features learned during the construction process by the Feature Evolution Dependency Network (FEDN) play a crucial role. The FEDN analyzes the changing trends of each feature dimension in the input feature vector over time, which is directly related to the content learned by the dynamic dependency learning layer during network construction. The FEDN not only focuses on the feature state at the current moment but also analyzes the dynamic changes of features over time. By observing the strength and direction of the dependencies between learned feature dimensions as the sample sequence changes, it observes the changing trends of each feature dimension, such as the propagation rate of surface cracks and textures, and the increasing or decreasing trends of internal stress and strain. Based on the dynamic changes of features, a risk trend model is established. This model considers the mutual influence and coupling effects between features and predicts the development direction of features over a future period. For example, if the propagation rate of surface cracks and textures on the road property accelerates, and the internal stress and strain also show a continuous increasing trend, and there is a strong positive correlation between these two features, then the model judges that the risk is likely to intensify. Based on the prediction results of the risk trend model, the evolution direction of potential risks to road assets is determined. The evolution direction of potential risks can be divided into several scenarios, such as risk aggravation, risk mitigation, and risk stabilization. For example, if the model predicts that in the future, all characteristic indicators of road assets will develop in a direction unfavorable to structural stability, such as continued expansion of cracks and continuous increase in stress, then the evolution direction of potential risks can be determined as risk aggravation. Conversely, if the characteristic indicators show an improving trend, such as a slowdown in crack propagation and a decrease in stress, then the evolution direction of potential risks is risk mitigation.
[0222] It is understood that the various algorithms involved in the above descriptions of the embodiments of the present invention, such as histogram equalization algorithms, filtering algorithms, cosine distance algorithms, etc., can all be obtained from relevant content in the prior art. To save space, they will not be elaborated on in the embodiments of the present invention. In addition, those skilled in the art can supplement the details based on common knowledge in the art when implementing the solutions of the present invention. For example, they can use normalization to eliminate dimensional conflicts before feature fusion, use interpolation to eliminate dimensional differences, reasonably set thresholds based on historical data, experience or business scenario requirements, train the model based on a general model training method, set the number of layers in the model structure based on actual needs, select activation functions, etc. The present invention will not provide redundant descriptions of overly detailed implementation processes here.
[0223] Please see Figure 2 , Figure 2This is a schematic diagram of a computer system provided in an embodiment of the present invention. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 can be connected via a bus or other means. The processor 101 (or Central Processing Unit, CPU) is the computing and control core of the computer system, capable of parsing various instructions and processing various data within the computer system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for data transmission and interaction within the computer system. The memory 103 is a storage device in the computer system used to store programs and data. It is understood that the memory 103 here can include the computer system's built-in memory, or it can include extended memory supported by the computer system. The memory 103 provides storage space, which stores the computer system's operating system; this invention does not limit this storage space.
[0224] In one embodiment, the processor 101 executes the road property status monitoring data analysis method using deep learning provided above in the embodiments of the present invention by running a computer program in the memory 103.
Claims
1. A method for analyzing road asset condition monitoring data using deep learning, characterized in that, include: Acquire a road asset data set, which includes road asset surface image data and road asset structure sensing data; Performing cross-modal feature association mapping on the road asset dataset to generate a cross-modal association feature matrix includes: dividing the road asset surface image data in the road asset dataset into regions to obtain a road asset image region set, where each road asset image region in the road asset image region set contains representative morphological texture information of the road asset surface; enhancing the texture of each road asset image region in the road asset image region set to obtain enhanced texture regions; and performing spatiotemporal association anchoring on the road asset structure sensing data in the road asset dataset to obtain spatiotemporally anchored sensing data, including: from the road asset... The dataset extracts raw acquisition information from road asset structure sensor data and imaging region information from road asset surface image data. The raw acquisition information includes the sensor data acquisition location and acquisition time, and the imaging region information includes the road asset entity location corresponding to the image region and the imaging time. A road asset spatial association map is constructed, with road asset entity segments as nodes and topological connections between nodes as edges, storing the spatial coordinate range and road asset type attributes corresponding to each node. Based on the road asset spatial association map, the sensor data acquisition locations in the raw acquisition information are spatially matched with the road asset entity locations in the imaging region information to determine the road asset image region to which each sensor data acquisition location belongs, obtaining a spatial association result. Based on the spatial association result, the acquisition time in the raw acquisition information is temporally associated with the imaging time of the corresponding road asset image region, calculating the association strength value between the acquisition time and the imaging time. The association strength value decreases as the time interval increases. Sensor data units whose association strength values meet a preset strength threshold are selected to obtain preliminary anchored sensor data, which carries a corresponding road asset image region identifier. The initial anchoring sensing data is fused with spatiotemporal labels. The road property image region identifier, spatial association result, and temporal association strength value are embedded into the sensing data unit to obtain the spatiotemporal anchoring sensing data. The spatiotemporal anchoring sensing data carries a spatial location identifier and a temporal association identifier corresponding to the road property image region. The row vectors of the cross-modal association feature matrix correspond to a single heterogeneous data sample, and the column vectors correspond to the fused cross-modal feature dimensions. The single heterogeneous data sample is obtained by associating road property surface image data and road property structure sensing data at the same time period and spatial location. A feature evolution dependency network is constructed based on the cross-modal correlation feature matrix. The feature evolution dependency network is used to characterize the dynamic dependency relationship of cross-modal feature dimensions over time and the coupling effect between features. The feature evolution depends on the network outputting the road property state inference results.
2. The method according to claim 1, characterized in that, The step of performing cross-modal feature association mapping on the road asset data set to generate a cross-modal association feature matrix includes: The enhanced texture region and the spatiotemporal anchoring sensing data are input into a cross-modal mapping network. The region feature extraction layer of the cross-modal mapping network is used to extract local texture features from the enhanced texture region to obtain a region texture feature vector. The temporal feature extraction layer of the cross-modal mapping network is used to extract dynamic stress features from the spatiotemporal anchoring sensing data to obtain a temporal stress feature vector. Modal feature matching is performed on the regional texture feature vector and the temporal stress feature vector. A nonlinear correspondence between texture features and stress features is established through a dynamic mapping mechanism to obtain the matched texture feature vector and the matched stress feature vector. Modal interaction fusion is performed on the matched texture feature vector and the matched stress feature vector to obtain a cross-modal fusion feature vector; The cross-modal fusion feature vectors are arranged sequentially according to the sample order in the road production data set to generate the cross-modal association feature matrix.
3. The method according to claim 2, characterized in that, The step of extracting local texture features from the enhanced texture region through the region feature extraction layer of the cross-modal mapping network to obtain a region texture feature vector includes: The enhanced texture region is input into the first feature extraction unit of the region feature extraction layer. The texture-aware convolution kernel of the first feature extraction unit performs local neighborhood feature response calculation on the enhanced texture region to obtain a first-level texture feature map. The number of channels in the first-level texture feature map corresponds to the number of texture-aware convolution kernels. The first-level texture feature map is input to the first feature aggregation unit of the region feature extraction layer. The first feature aggregation unit performs feature aggregation on the spatial dimension of the first-level texture feature map to obtain a first aggregated feature map. The spatial size of the first aggregated feature map is smaller than the spatial size of the first-level texture feature map. The first aggregated feature map is input into the second layer feature extraction unit of the region feature extraction layer. The first aggregated feature map is subjected to multi-receptive field feature extraction through the multi-scale convolution kernel of the second layer feature extraction unit to obtain the second-level texture feature map. The number of channels of the second-level texture feature map is the total number of the multi-scale convolution kernels. The second-level texture feature map is input into the second feature aggregation unit of the region feature extraction layer. The second feature aggregation unit performs feature aggregation on the channel dimension of the second-level texture feature map to obtain a second aggregated feature map. The number of channels in the second aggregated feature map is less than the number of channels in the second-level texture feature map. The second aggregated feature map is flattened in spatial dimension to obtain a flattened texture vector; The flattened texture vector is subjected to nonlinear feature transformation to obtain the region texture feature vector.
4. The method according to claim 2, characterized in that, The modal feature matching of the region texture feature vector and the temporal stress feature vector, and the establishment of a nonlinear correspondence between texture features and stress features through a dynamic mapping mechanism, to obtain the matched texture feature vector and the matched stress feature vector, includes: A modal feature mapping network is constructed, which includes a texture-to-stress mapping branch and a stress-to-texture mapping branch. The texture-to-stress mapping branch is used to map texture features to a stress feature space, and the stress-to-texture mapping branch is used to map stress features to a texture feature space. The region texture feature vector is input into the texture-to-stress mapping branch, and the mapping weight of the texture feature dimension to the stress feature dimension is learned through the dynamic weight layer of the texture-to-stress mapping branch to obtain the texture-mapped stress vector. The temporal stress feature vector is input into the stress-to-texture mapping branch, and the dynamic bias layer of the stress-to-texture mapping branch is used to learn the mapping bias of the stress feature dimension to the texture feature dimension to obtain the stress-mapped texture vector. Calculate the first feature distance value between the texture-mapped stress vector and the temporal stress feature vector, and calculate the second feature distance value between the stress-mapped texture vector and the regional texture feature vector; The mapping parameters of the modal feature mapping network are adjusted based on the first feature distance value and the second feature distance value so that both feature distance values meet the preset distance threshold. After the mapping parameters are adjusted, the region texture feature vector is passed through the adjusted stress-to-texture mapping branch to obtain the matching texture feature vector, and the temporal stress feature vector is passed through the adjusted texture-to-stress mapping branch to obtain the matching stress feature vector.
5. The method according to claim 1, characterized in that, The construction of the feature evolution dependency network based on the cross-modal correlation feature matrix includes: The cross-modal correlation feature matrix is analyzed for feature dimensions to obtain an image source feature sub-matrix and a sensor source feature sub-matrix. The image source feature sub-matrix is composed of feature dimension column vectors from road surface image data in the cross-modal correlation feature matrix, and the sensor source feature sub-matrix is composed of feature dimension column vectors from road structure sensor data in the cross-modal correlation feature matrix. A dual-path collaborative dependency modeling network is constructed, comprising an image-guided dependency path and a sensor-guided dependency path. The image-guided dependency path is used to learn the dependencies of cross-modal features based on image source features, and the sensor-guided dependency path is used to learn the dependencies of cross-modal features based on sensor source features. The image source feature submatrix is input into the image guidance dependency path. The influence weight of each image source feature dimension on other feature dimensions is calculated through the feature attention mechanism of the image guidance dependency path to obtain the image guidance weight matrix. The cross-modal correlation feature matrix is weighted and modulated based on the image guidance weight matrix to obtain the image modulation feature matrix. The sensor source feature submatrix is input into the sensor guidance dependency path. The control coefficients of each sensor source feature dimension on other feature dimensions are calculated through the feature gating mechanism of the sensor guidance dependency path to obtain the sensor control coefficient matrix. The cross-modal correlation feature matrix is gating and filtering based on the sensor control coefficient matrix to obtain the sensor filtering feature matrix. The image modulation feature matrix and the sensing screening feature matrix are subjected to feature interaction enhancement to obtain a collaborative dependency feature matrix; The collaborative dependency feature matrix is input into the dynamic dependency learning layer of the dual-path collaborative dependency modeling network. The dynamic dependency learning layer learns the dependency strength and direction between feature dimensions as the sample sequence changes, thereby generating the feature evolution dependency network.
6. The method according to claim 5, characterized in that, The image guidance weight matrix is obtained by calculating the influence weights of each image source feature dimension on other feature dimensions through the feature attention mechanism of the image guidance dependency path, including: Extract an image source feature dimension set from the image source feature submatrix, wherein the image source feature dimension set contains feature dimension identifiers corresponding to all column vectors of the image source feature submatrix; For each image source feature dimension in the set of image source feature dimensions, a correlation evaluation sequence is constructed between the image source feature dimension and all other feature dimensions in the cross-modal correlation feature matrix. Each element of the correlation evaluation sequence is the correlation value between the column vector of the image source feature dimension and the column vector of the corresponding other feature dimensions. The correlation evaluation sequence is normalized to obtain a normalized correlation sequence, and the sum of all elements of the normalized correlation sequence is a preset normalization benchmark value. The normalized correlation sequence is input into the weight generation unit of the feature attention mechanism. The weight generation unit performs a nonlinear transformation on the normalized correlation sequence to obtain the influence weight vector corresponding to the feature dimension of the image source. The influence weight vectors corresponding to all image source feature dimensions in the image source feature dimension set are arranged in order to generate the image guiding weight matrix.
7. The method according to claim 5, characterized in that, The feature gating mechanism of the sensing guidance dependent path is used to calculate the control coefficients of each sensor source feature dimension on other feature dimensions, resulting in a sensing control coefficient matrix, including: Extract the sensor source feature dimension sequence from the sensor source feature submatrix. The sensor source feature dimension sequence is a sequence of feature dimensions corresponding to the column vectors of the sensor source feature submatrix arranged in a preset logical order. For each sensor source feature dimension in the sensor source feature dimension sequence, calculate the feature fluctuation value of that sensor source feature dimension in the cross-modal correlation feature matrix. The feature fluctuation value is used to characterize the discreteness of the column vector element values of that sensor source feature dimension. The feature fluctuation value is input to the reset gate unit of the feature gating mechanism. The reset gate unit performs a gating threshold judgment on the feature fluctuation value to obtain a reset gate value. The reset gate value is used to indicate whether to retain the control effect of the feature dimension of the sensor source on other feature dimensions. The feature fluctuation value is input to the update gate unit of the feature gating mechanism. The update gate unit dynamically adjusts the range of the feature fluctuation value to obtain the update gate value. The update gate value is used to indicate the control strength of the feature dimension of the sensing source on other feature dimensions. Based on the reset gate value and the updated gate value, calculate the comprehensive control coefficient of the sensor source feature dimension on other feature dimensions to obtain the control coefficient vector; The control coefficient vectors corresponding to all sensor source feature dimensions in the sensor source feature dimension sequence are arranged in order to generate the sensor control coefficient matrix.
8. The method according to claim 5, characterized in that, The step of performing feature interaction enhancement on the image modulation feature matrix and the sensing screening feature matrix to obtain a collaboratively dependent feature matrix includes: Obtain the matrix structure parameters of the image modulation feature matrix and the sensor screening feature matrix, wherein the matrix structure parameters include the number of matrix rows and the number of matrix columns; When there is a difference between the number of rows in the image modulation feature matrix and the number of rows in the sensor screening feature matrix, the matrix with fewer rows is expanded in row dimension by supplementing the row vectors through feature interpolation, so that the number of rows in the expanded matrix is the same as the number of rows in the other matrix. When there is a difference between the number of columns in the image modulation feature matrix and the number of columns in the sensor screening feature matrix, the matrix with fewer columns is expanded in column dimension by supplementing column vectors through feature mapping, so that the number of columns in the expanded matrix is the same as the number of columns in the other matrix. The image modulation feature matrix and the sensor screening feature matrix after structural unification are subjected to element-wise interactive calculation to obtain a preliminary interactive matrix. Each element of the preliminary interactive matrix is the result of the interactive operation of the corresponding elements of the two matrices. The preliminary interaction matrix is subjected to feature sparsification, and the element values that satisfy the preset sparsity conditions are adjusted to preset reference values to obtain a sparse interaction matrix. The row direction features of the sparse interaction matrix are normalized so that the energy value of each row element satisfies the preset energy distribution condition, thus obtaining the cooperative dependency feature matrix.
9. A computer system, characterized in that, include: A memory, wherein a computer program is stored; A processor is configured to load the computer program to implement the road property status monitoring data analysis method using deep learning as described in any one of claims 1-8.
Citation Information
Patent Citations
High-speed wireless sensing network node for monitoring active structure health
CN101262378A
Data management method and system for distributed data architecture
CN120336332A