Road network disease image recognition method and system based on deep learning

By integrating and processing road surface image sequences and inertial navigation data, a cross-domain feature matrix is ​​generated and combined with road network topology information for defect identification. This solves the problem of insufficient accuracy and reliability of defect identification in existing technologies, and achieves a more comprehensive and reliable defect identification effect.

CN122020351BActive Publication Date: 2026-07-21GUIZHOU HUILIANTONG ELECTRONIC COMMERCE SERVICE CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUIZHOU HUILIANTONG ELECTRONIC COMMERCE SERVICE CO LTD
Filing Date
2025-12-08
Publication Date
2026-07-21

Smart Images

  • Figure CN122020351B_ABST
    Figure CN122020351B_ABST
Patent Text Reader

Abstract

The application provides a kind of road network disease image recognition method and system based on deep learning, continuous image sequence of road surface and corresponding period's vehicle-mounted inertial navigation data are collected, both are integrated into joint collection data set by time stamp synchronous processing, the image frame in it is handled by spectrum-geometry dual-domain decoupling, the spectrum feature component and geometric structure component of each image frame are obtained, the correlation mapping processing is carried out to the two, and a cross-domain feature matrix is generated;It is input into the pre-trained topological perception graph neural network, and the graph structure modeling is carried out in combination with the connectivity information of road network node, and the topological embedding vector is generated by iteratively updating the node association strength;Extract the vibration spectrum data in the navigation data record corresponding to the topological embedding vector in the joint collection data set, and perform cross-modal alignment processing on the two to generate an aligned enhanced feature vector;Its recognition generates road network disease recognition result.The application can improve the accuracy of road network disease recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and more specifically, to a method and system for recognizing road network defects based on deep learning. Background Technology

[0002] With the increasing demand for road infrastructure maintenance, the analysis and processing of road surface image data can enable automated identification of the types and degrees of defects such as cracks, potholes, and settlement. Currently, computer vision and deep learning methods are widely combined to improve identification efficiency. However, existing technologies lack sufficient fine-grained decomposition and structured association of image and geometric features during the feature processing stage. This results in limited utilization of mutual information across multiple dimensions in feature representation. Furthermore, the insufficient integration of prior knowledge such as road network topology during feature modeling makes it difficult to accurately capture the correlation patterns of defect features within the road network space, thus hindering the improvement of the accuracy and reliability of defect identification. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for road network defect image recognition based on deep learning. The technical solution of this invention is implemented as follows: In a first aspect, the present invention provides a method for road network defect image recognition based on deep learning. The method includes: acquiring a continuous image sequence of the road surface and corresponding vehicle inertial navigation data for the corresponding time period; integrating the continuous image sequence and the vehicle inertial navigation data into a joint acquisition data set through timestamp synchronization processing; each data unit in the joint acquisition data set contains an image frame at the same acquisition time and a corresponding navigation data record; performing spectral-geometric dual-domain decoupling processing on the image frames in the joint acquisition data set to obtain the spectral feature components and geometric structure components of each image frame; and combining the spectral feature components and the geometric structure components... The components are correlated and mapped to generate a cross-domain feature matrix. The cross-domain feature matrix is ​​then input into a pre-trained topology-aware graph neural network. Combined with road network node connectivity information, the cross-domain feature matrix is ​​used for graph structure modeling. A topology embedding vector is generated by iteratively updating the node association strength. Vibration spectrum data is extracted from the navigation data records corresponding to the topology embedding vector in the jointly acquired data set. The topology embedding vector and the vibration spectrum data are then aligned across modes to generate an alignment-enhanced feature vector. Finally, the alignment-enhanced feature vector is subjected to joint identification processing for disease type and severity to generate a road network disease identification result.

[0004] In a second aspect, the present invention provides a computer system including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement the steps in the method described above.

[0005] This invention integrates continuous image sequences of road surfaces and corresponding time-segment vehicle inertial navigation data into a joint acquisition dataset by time-stamping and synchronizing the data. This creates a unified data unit that spatiotemporally binds visual and physical sensing data. The image frames in the joint acquisition dataset undergo spectral-geometric dual-domain decoupling processing to obtain spectral feature components and geometric structure components, which are then correlated and mapped to generate a cross-domain feature matrix. This achieves refined decomposition and structured association of spectral and geometric features, improving the richness and accuracy of feature representation. The cross-domain feature matrix is ​​input into a topology-aware graph neural network. Combined with road network node connectivity information, graph structure modeling is performed, and topology embedding vectors are generated by iteratively updating node association strength. This allows the graph network to integrate prior knowledge of road network topology and adaptively learn the propagation and association patterns of defect features within the road network. This addresses the shortcomings of traditional feature extraction methods that neglect the physical connections of the road network, enhancing the dynamic correlation between features. Corresponding vibration spectrum data is extracted and cross-modal aligned with topological embedding vectors to generate aligned and enhanced feature vectors. This establishes a spatiotemporal correlation between visual topological features and physical vibration features, resolving feature misalignment issues caused by modal differences in multimodal data and achieving deep complementary enhancement of intermodal information. The aligned and enhanced feature vectors are then subjected to joint identification processing for disease type and severity to generate road network disease identification results. This ensures the inherent consistency of disease type and severity identification results and improves the comprehensiveness and reliability of the identification results. Through the above technical solution, this method can effectively integrate multi-source data, refine image feature processing, dynamically learn road network topological correlations, deeply fuse cross-modal information, and collaboratively identify disease attributes, thereby improving the accuracy and practicality of road network disease identification and providing more comprehensive and reliable technical support for road network disease assessment. Attached Figure Description

[0006] Figure 1 This is a schematic diagram illustrating an application scenario of the deep learning-based road network defect image recognition method provided in this embodiment of the invention.

[0007] Figure 2 This is a flowchart of a road network defect image recognition method based on deep learning provided in an embodiment of the present invention.

[0008] Figure 3 This is a schematic diagram of the composition of a computer system provided in an embodiment of the present invention. Detailed Implementation

[0009] The deep learning-based road network defect image recognition method provided in this invention can be executed by a computer system, which can be various types of terminals such as laptops, tablets, and desktop computers, or it can be implemented as a server. The server can be a standalone physical server, a server cluster consisting of multiple physical servers, or a distributed system.

[0010] Figure 1 This is a schematic diagram illustrating an application scenario of the deep learning-based road network defect image recognition method provided in this embodiment of the invention. It includes multiple data acquisition devices 100, a network 200, and a computer system 300. The multiple data acquisition devices 100 and the computer system 300 are connected via the network 200. The computer system 300 is used to execute the method provided in this embodiment of the invention. The data acquisition devices 100 can be image acquisition devices on vehicles, inertial navigation devices, etc.

[0011] Specifically, embodiments of the present invention provide a method for road network defect image recognition based on deep learning, such as... Figure 2 As shown, the method includes: Step S100: Collect a continuous image sequence of the road surface and the vehicle inertial navigation data of the corresponding time period. Through timestamp synchronization processing, integrate the continuous image sequence and the vehicle inertial navigation data into a joint acquisition data set. Each data unit in the joint acquisition data set contains an image frame at the same acquisition time and the corresponding navigation data record.

[0012] A continuous image sequence of a road surface is a series of image frames captured at set time intervals as a vehicle travels along the road, using image acquisition equipment mounted on the vehicle. These image frames are arranged sequentially according to the time of capture, reflecting the actual condition of the road surface at different moments. The image acquisition equipment typically needs to have high resolution and a suitable frame rate to ensure that the captured images are clear and can capture subtle features of the road surface. For example, a high-definition camera can be selected, which can accurately record information such as texture and color of the road surface under different lighting conditions.

[0013] Vehicle inertial navigation data is collected in real time by an inertial navigation system installed on the vehicle. This system mainly consists of sensors such as accelerometers and gyroscopes, capable of measuring physical quantities such as vehicle acceleration and angular velocity. By integrating these physical quantities, information such as the vehicle's position, velocity, and attitude can be obtained. During vehicle operation, the inertial navigation system continuously records this data, yielding vehicle inertial navigation data for the corresponding time period.

[0014] Timestamp synchronization is used to ensure temporal consistency between continuous image sequences and vehicle inertial navigation data. During image and navigation data acquisition, each image frame and each navigation data record is marked with a precise timestamp. Timestamps can be provided by a high-precision clock to ensure accuracy. After data acquisition, the continuous image sequences and vehicle inertial navigation data are matched based on these timestamps. Specifically, the image frames and navigation data records are traversed to find those with the same timestamp and combined to obtain a joint acquired data set.

[0015] Step S200: Perform spectral-geometric dual-domain decoupling processing on the image frames in the joint acquisition data set to obtain the spectral feature components and geometric structure components of each image frame. Perform correlation mapping processing on the spectral feature components and geometric structure components to generate a cross-domain feature matrix.

[0016] The purpose of spectral-geometric dual-domain decoupling processing is to effectively separate the spectral information and geometric structure information in an image frame. The spectral information in the image frame reflects the absorption and reflection characteristics of different materials on the road surface to different wavelengths of light, while the geometric structure information reflects the shape, edges, and contours of the road surface. By decoupling these two types of information, a more in-depth analysis of the different features in the image frame can be achieved.

[0017] Spectral feature components are spectral-related features in an image frame, containing information about different colors and spectral responses. This information can be used to identify different road surface materials and determine the presence of defects. For example, different types of defects may exhibit different spectral characteristics; analyzing spectral feature components can lead to more accurate defect detection.

[0018] Geometric structure components are features in an image frame that relate to geometric structures, such as the edges and contours of a road surface. Association mapping is the operation that maps and associates spectral feature components with geometric structure components. Since spectral feature components and geometric structure components describe the features of an image frame from different perspectives, association mapping can integrate them to generate a cross-domain feature matrix.

[0019] For example, step S200 may specifically include the following steps S210 to S250: Step S210: Perform color space conversion processing on the image frame, convert the image frame from the original color space to the target color space, separate multiple spectral channel components, perform spatial noise suppression processing on each spectral channel component, and generate a denoised spectral channel set.

[0020] Color space conversion is the process of transforming an image frame from its original color space to a target color space. Different color spaces represent colors differently, and choosing a suitable target color space makes it easier to analyze and process the spectral information of an image. For example, the original color space might be RGB, while the target color space could be HSV. In RGB, colors are represented by three channels: red, green, and blue. In HSV, colors are represented by three channels: hue, saturation, and lightness. By converting the image frame from RGB to HSV, the color characteristics of the image can be analyzed more intuitively. After color space conversion, the image frame is separated to obtain multiple spectral channel components. Each spectral channel component corresponds to a channel in the target color space; for example, in HSV, the hue, saturation, and lightness channels are separated.

[0021] For example, step S210 may specifically include the following steps S211 to S215: Step S211: Perform local variance calculation for each spectral channel component. Use a preset sliding window to traverse each pixel of the spectral channel component, calculate the variance value of the pixel value within the window, and generate a variance distribution map.

[0022] Local variance calculation is a method used to analyze pixel value variations in local regions of an image. For each spectral channel component, a sliding window of a preset size is used to traverse the spectral channel component. The size of the sliding window can be adjusted according to the actual situation, generally taking into account factors such as noise characteristics and image resolution. During the traversal, for each pixel within the window, the variance of its surrounding pixel values ​​is calculated. The variance reflects the dispersion of pixel values ​​within the window; the larger the variance, the more drastic the pixel value variation in that area, potentially indicating the presence of noise or edge features.

[0023] After calculating the variance of the pixel values ​​within each window, these variance values ​​are arranged according to the pixel positions to generate a variance distribution map. The variance distribution map is an image with the same size as the spectral channel components, where the value of each pixel represents the variance of the pixel values ​​within the corresponding window.

[0024] Step S212: Determine the segmentation threshold between noise and non-noise regions based on the variance distribution map. Mark regions with variance values ​​greater than the segmentation threshold as noise candidate regions and regions with variance values ​​less than or equal to the segmentation threshold as non-noise regions.

[0025] The segmentation threshold is determined based on the variance distribution plot. By analyzing the variance distribution plot, a suitable threshold can be found to divide the spectral channel components into noise and non-noise regions. There are several methods for determining the segmentation threshold. For example, an adaptive thresholding method can be used to automatically determine the threshold based on the statistical characteristics of the variance distribution plot; alternatively, a manual thresholding method can be used, adjusting it based on experience and actual conditions.

[0026] After determining the segmentation threshold, each pixel in the spectral channel components is evaluated. If the variance value corresponding to the pixel is greater than the segmentation threshold, the region is marked as a noise candidate region; if the variance value is less than or equal to the segmentation threshold, the region is marked as a non-noise region. Noise candidate regions typically contain more noise and require further processing; while non-noise regions are relatively stable with less noise and can be smoothed.

[0027] Step S213: Perform adaptive median filtering on the noise candidate region, adjust the size of the filtering window according to the distribution of neighboring pixel values ​​of the pixels in the noise candidate region, sort the pixel values ​​in the window, and replace the original pixel values ​​with the median value.

[0028] Within a noise candidate region, for each pixel, an initial filtering window size is first determined. Then, based on the distribution of neighboring pixel values, it is determined whether the filtering window size needs to be adjusted. If the changes in neighboring pixel values ​​are large, it indicates that the noise in that region is more complex, and the filtering window size may need to be increased; if the changes in neighboring pixel values ​​are small, it indicates that the noise in that region is relatively simple, and the filtering window size can be maintained or decreased.

[0029] After determining the size of the filtering window, the pixel values ​​within the window are sorted. The purpose of sorting is to find the median of the pixel values ​​within the window. The median is the pixel value located in the middle after arranging the pixel values ​​in ascending order. The median is chosen as the new value for that pixel because it has good noise resistance, effectively removing noise while preserving image edge and detail information.

[0030] Step S214: Perform guided filtering on the non-noise area. Using the original spectral channel components as the guide map, perform edge-preserving smoothing on the pixel values ​​of the non-noise area to retain edge contour information while suppressing high-frequency noise.

[0031] In non-noise regions, the original spectral channel components are used as the guide map. The guide map contains the original information of the image, especially important information such as edges and contours. Through guided filtering, the pixel values ​​in non-noise regions can be smoothed while preserving the edge and contour information of the image. During guided filtering, the pixel values ​​in non-noise regions are adjusted based on the information in the guide map. Specifically, for each pixel in a non-noise region, a weighted average is calculated as the new value of that pixel, considering the information of its neighboring pixels and the information of the corresponding position in the guide map. The calculation of the weighted average is adjusted according to the edge information in the guide map, so that the pixel value changes less at the edges, thus preserving the edge contour information; while in non-edge regions, the pixel values ​​are smoothed to suppress high-frequency noise. In this way, noise can be removed while maintaining the geometric structure features of the image.

[0032] Step S215: Merge the noisy candidate region processed by adaptive median filtering with the non-noisy region processed by guided filtering to generate denoised spectral channel components, and combine all spectral channel components into a denoised spectral channel set.

[0033] Region merging is an operation that recombines noisy candidate regions and non-noise regions that have undergone different filtering processes. After adaptive median filtering and guided filtering, the pixel values ​​of both the noisy candidate regions and non-noise regions are adjusted accordingly. By merging these two regions in their original positions, a complete, denoised spectral channel component is obtained. After obtaining the denoising result for each spectral channel component, all spectral channel components are combined to obtain the denoised spectral channel set. The denoised spectral channel set removes noise from the spectral channel components, improves the quality of spectral information, and provides more accurate data for subsequent spectral feature extraction.

[0034] Step S220: Extract the edge contour of the image frame, identify the structural edge lines of the road surface in the image frame, calculate the rate of curvature change and directional gradient distribution of the structural edge lines, and generate a geometric structure description vector.

[0035] Edge contour extraction is the process of identifying the structural edge lines of a road surface from an image frame. These edge lines represent the boundaries between different areas of the road surface, such as road boundaries and the edges of cracks. Edge contour extraction separates these edge lines from the image frame, providing a basis for subsequent geometric analysis. Edge contour extraction methods include gradient-based methods, which calculate the gradient values ​​of pixels in the image frame and identify regions with larger gradient values ​​as edge lines; and threshold-based methods, which set a threshold and mark regions where pixel value changes exceeding that threshold as edge lines.

[0036] After identifying the structural edge lines of the road surface, the rate of change of curvature and the directional gradient distribution of these edge lines are calculated. The rate of change of curvature reflects the degree of change in the curvature of the edge line, indicating its bending at different locations. Calculating the rate of change of curvature reveals the complexity and variability of the road surface structure. The directional gradient distribution provides information on the gradient of the edge line in different directions, reflecting its directional characteristics. Analyzing the directional gradient distribution allows for further determination of the edge line's orientation and shape.

[0037] The calculated rate of curvature change and directional gradient distribution are combined to generate a geometric structure description vector. This vector contains the geometric features of the road surface's edge lines, with each element corresponding to a geometric feature parameter of that edge line. This geometric structure description vector allows the road surface's geometric features to be represented by a single vector, facilitating subsequent feature analysis and processing.

[0038] For example, step S220 may specifically include the following steps S221 to S225: Step S221: Extract sampling points from the identified road surface structure edge lines. Collect the coordinates of sampling points at equal intervals along the edge lines to generate a sampling point sequence. The sampling point sequence contains a set of coordinate points continuously distributed along the edge lines.

[0039] Sampling point extraction is performed to discretize the structural edge lines, facilitating subsequent calculations and analysis. For the identified road surface structural edge lines, sampling point coordinates are collected along the edge lines at equal intervals. The choice of the equal interval distance needs to be adjusted based on the complexity of the edge lines and the accuracy requirements of subsequent calculations. If the equal interval distance is too small, the computational load increases; if the equal interval distance is too large, some important edge information may be lost. When collecting sampling point coordinates, starting from one end of the edge line, coordinate points are collected sequentially at equal intervals. These coordinate points constitute a sampling point sequence. The coordinate points in the sampling point sequence are continuously distributed and can reflect the approximate shape of the edge line. By generating the sampling point sequence, the continuous edge line can be transformed into a discrete set of coordinate points, providing a data foundation for subsequent curvature and gradient calculations.

[0040] Step S222: Calculate the tangent direction vector of adjacent sampling points in the sampling point sequence. Take the derivative of the tangent direction vector of each sampling point to obtain the curvature vector. The magnitude of the curvature vector represents the magnitude of the curvature at that sampling point, and the direction represents the direction of curvature change.

[0041] The tangent direction vector is the vector that is tangent to the edge line at a sampling point, reflecting the direction of the edge line at that point. For adjacent sampling points in a sampling point sequence, the tangent direction vector between them is calculated. Specifically, by calculating the coordinate difference between adjacent sampling points, a vector can be obtained. Normalizing this vector yields the tangent direction vector.

[0042] After obtaining the tangent direction vector at each sampling point, the derivative of the tangent direction vector is calculated. The purpose of differentiation is to obtain the curvature vector. The magnitude of the curvature vector represents the magnitude of curvature at that sampling point, reflecting the degree of bending of the edge line at that point; the direction of the curvature vector represents the direction of curvature change, that is, the direction of change of the degree of bending of the edge line. By calculating the curvature vector, a deeper understanding of the geometric characteristics of the edge line can be obtained.

[0043] Step S223: Calculate the rate of change of the curvature vector. Perform a difference operation on the curvature magnitude of continuous sampling points to obtain the rate of change of curvature. The rate of change of curvature is used to characterize the degree of change in the curvature of the edge line.

[0044] The rate of change of the magnitude of a curvature vector is how the magnitude of the curvature vector changes between adjacent sampling points. To calculate the rate of change of the magnitude of the curvature vector, a difference operation is performed on the curvature magnitudes of consecutive sampling points. The difference operation calculates the difference in curvature magnitudes between adjacent sampling points. Through the difference operation, the rate of change of curvature at each sampling point can be obtained.

[0045] The rate of change of curvature is used to characterize the degree of change in the curvature of an edge line. A large rate of change of curvature indicates that the curvature of the edge line changes rapidly in that area, potentially indicating a complex geometric structure; a small rate of change of curvature indicates that the curvature of the edge line changes slowly in that area, suggesting a relatively simple geometric structure. By analyzing the rate of change of curvature, the structural characteristics of the road surface can be determined more accurately.

[0046] Step S224: Perform directional gradient calculation on the structural edge lines. Perform horizontal and vertical gradient calculation on the grayscale image of the image frame to generate a directional gradient map. The directional gradient map contains the gradient magnitude and gradient direction of each pixel.

[0047] Oriented gradient calculation is used to obtain gradient information of structural edge lines in different directions. First, the image frame is converted into a grayscale image. Grayscale images contain only brightness information, making gradient calculation easier. Then, horizontal and vertical gradients are calculated on the grayscale image. Horizontal gradient calculation is achieved by calculating the brightness difference between adjacent pixels in the horizontal direction; vertical gradient calculation is achieved by calculating the brightness difference between adjacent pixels in the vertical direction. After completing the horizontal and vertical gradient calculations, the gradient magnitude and gradient direction of each pixel are calculated based on the gradient values ​​in these two directions. The gradient magnitude represents the magnitude of the gradient at that pixel, reflecting the degree of brightness change at that point; the gradient direction represents the direction of the gradient at that pixel, reflecting the direction of brightness change at that point. Combining the gradient magnitude and gradient direction of each pixel generates the oriented gradient map. The oriented gradient map contains gradient information for each pixel in the image frame, especially the gradient information at structural edge lines.

[0048] Step S225: Extract the gradient magnitude and gradient direction of the structural edge line from the gradient orientation map, construct the gradient orientation distribution matrix, and concatenate the rate of curvature change with the gradient orientation distribution matrix as row vectors to generate a geometric structure description vector. Each element of the geometric structure description vector corresponds to a geometric feature parameter of the edge line.

[0049] Extracting the gradient magnitude and direction of the structural edge line from the gradient orientation map is to obtain the directional gradient information of the structural edge line. This is done by locating the position corresponding to the structural edge line in the gradient orientation map, extracting the gradient magnitude and direction at that position, and constructing a gradient orientation distribution matrix. The gradient orientation distribution matrix is ​​a matrix containing the directional gradient information of the structural edge line, where each element corresponds to the gradient magnitude and direction of the edge line at a specific location.

[0050] Concatenating the rate of change of curvature and the directional gradient distribution matrix as row vectors combines the curvature and directional gradient information of the structural edge line. This concatenation operation creates a new vector, the geometric structure description vector. Each element of the geometric structure description vector corresponds to a geometric feature parameter of the edge line, such as the rate of change of curvature, gradient magnitude, and gradient direction. Using this geometric structure description vector, the geometric features of the road surface structure edge line can be comprehensively represented by a single vector, providing a unified representation for subsequent feature analysis and processing.

[0051] Step S230: Input the denoised spectral channel set into the spectral feature encoder, and perform nonlinear transformation processing on each spectral channel component through a multilayer perceptron to generate spectral feature components. The spectral feature components contain the spectral response intensity distribution of each channel.

[0052] A spectral feature encoder is a model used to extract spectral features. Inputting the denoised spectral channel set into the spectral feature encoder is to further mine the spectral information within the channel set. A multilayer perceptron is the core component of the spectral feature encoder, consisting of multiple neuron layers connected by nonlinear activation functions.

[0053] In a multilayer perceptron, a nonlinear transformation is performed on each spectral channel component. This nonlinear transformation involves the neurons of the multilayer perceptron calculating and transforming the input spectral channel components. Specifically, the input spectral channel components first undergo a linear transformation by the neurons in the first layer, and then a nonlinear mapping is performed through a nonlinear activation function to obtain a new feature representation. This new feature representation is then used as the input to the next layer of neurons for further linear transformations and nonlinear mappings. After multiple layers of processing, the final spectral feature components are obtained. These spectral feature components contain the spectral response intensity distribution of each channel. The spectral response intensity distribution reflects the image's response in different spectral channels; different substances will have different response intensities in different spectral channels.

[0054] Step S240: Perform structural feature encoding on the geometric structure description vector, extract spatial dimension features from the geometric structure description vector based on a convolutional neural network, and generate geometric structure components. The geometric structure components include the spatial topological relationships of the edge lines.

[0055] Structural feature encoding further processes the geometric structure description vector to extract its spatial dimensional features. Convolutional neural networks are specifically designed for processing data with grid structures and are highly capable of extracting image and geometric features.

[0056] After the geometric structure description vector is input into a convolutional neural network (CNN), the CNN extracts its spatial dimension features. A CNN consists of multiple convolutional layers, pooling layers, and fully connected layers. In the convolutional layers, a sliding convolution operation is performed on the geometric structure description vector using a convolution kernel. The convolution operation involves element-wise multiplying the kernel with a local region of the geometric structure description vector and summing the results to obtain a new feature value. By continuously sliding the convolution kernel, the convolutional features of the entire geometric structure description vector can be obtained.

[0057] Pooling layers reduce the dimensionality of convolutional features, decreasing the number of features while retaining important feature information. Common pooling operations include max pooling and average pooling. Max pooling selects the largest feature value as the output of each pooling window; average pooling calculates the average of the feature values ​​in each pooling window and uses that average as the output.

[0058] After processing by convolutional and pooling layers, the resulting features are input into a fully connected layer. The fully connected layer further combines and transforms these features, ultimately generating geometric structural components. These components contain the spatial topological relationships of the edge lines, such as their connection methods and intersections.

[0059] Step S250: Construct a spectral-geometric correlation mapping matrix, match and align the feature dimensions of the spectral feature components with those of the geometric structure components, and perform correlation mapping processing on the spectral feature components and geometric structure components through feature connection operations to generate a cross-domain feature matrix. The row dimension of the cross-domain feature matrix corresponds to the feature channels of the spectral feature components, and the column dimension corresponds to the spatial position of the geometric structure components.

[0060] A spectral-geometric correlation mapping matrix is ​​used to correlate spectral feature components and geometric structure components. Before constructing the spectral-geometric correlation mapping matrix, the feature dimensions of the spectral feature components and geometric structure components need to be matched and aligned. Since the feature dimensions of the spectral feature components and geometric structure components may be different, methods are needed to adjust them to the same dimension. For example, interpolation, dimensionality reduction, and other operations can be used to achieve feature dimension matching and alignment. After completing the feature dimension matching and alignment, the spectral feature components and geometric structure components are correlated and mapped through a feature concatenation operation. The feature concatenation operation combines the features of the spectral feature components and geometric structure components to obtain a new feature matrix. Specifically, the spectral feature components and geometric structure components can be concatenated row by row or column by column to obtain a larger feature matrix.

[0061] The generated cross-domain feature matrix has rows corresponding to the feature channels of the spectral feature components and columns corresponding to the spatial locations of the geometric structure components. This cross-domain feature matrix integrates information from both spectral and geometric structure components, providing a more comprehensive description of the image's features. Analysis of the cross-domain feature matrix reveals the correlation between spectral and geometric structure information.

[0062] Step S300: Input the cross-domain feature matrix into the pre-trained topology-aware graph neural network, combine the road network node connectivity information to perform graph structure modeling processing on the cross-domain feature matrix, and generate topology embedding vectors by iteratively updating the node association strength.

[0063] Topology-aware graph neural networks (NATs) are specifically designed for processing graph-structured data, capturing the topological relationships between nodes in a graph. Pre-trained NATs are networks pre-trained on large-scale datasets, learning general graph structure features through pre-training. After inputting the cross-domain feature matrix into the pre-trained NAT, graph structure modeling is performed by combining it with road network node connectivity information. Road network node connectivity information describes the connections between nodes in the road network, such as which nodes are directly connected and the strength of those connections. By combining the cross-domain feature matrix with the road network node connectivity information, the cross-domain feature matrix can be transformed into a graph structure, where each node corresponds to an eigenvector in the cross-domain feature matrix, and the edges between nodes represent the connections between them.

[0064] In the graph structure modeling process, topological embedding vectors are generated by iteratively updating node association strengths. Node association strength represents the tightness of connections between nodes and is continuously updated with each iteration. Specifically, in each iteration, a new association strength between nodes is calculated based on the current node association strength and information from the cross-domain feature matrix. Then, the node features are updated according to the new association strength. After multiple iterations, the node features gradually converge, ultimately yielding the topological embedding vector.

[0065] For example, step S300 may specifically include the following steps S310 to S360: Step S310: Extract the feature node set from the cross-domain feature matrix, and use the feature vector of each row of the cross-domain feature matrix as the initial node feature of the graph neural network to generate a node feature matrix. The row dimension of the node feature matrix corresponds to the number of nodes, and the column dimension corresponds to the feature dimension.

[0066] Extracting the feature node set from the cross-domain feature matrix is ​​the first step in transforming the cross-domain feature matrix into a graph structure. Each row of the cross-domain feature matrix contains eigenvectors representing a feature, which are used as the initial node features for the graph neural network. Each eigenvector corresponds to a node in the graph, and these nodes form the feature node set. These initial node features are combined to generate the node feature matrix. The row dimension of the node feature matrix corresponds to the number of nodes, i.e., the number of nodes in the feature node set; the column dimension corresponds to the feature dimension, i.e., the length of each node's feature vector. The node feature matrix is ​​the input data for the graph neural network, containing the feature information from the cross-domain feature matrix, providing the foundation for subsequent graph structure modeling.

[0067] Step S320: Obtain road network node connectivity information. The road network node connectivity information includes the connection relationship between each road segment in the road network and the spatial distance between road segments. Construct an initial adjacency matrix based on the spatial distance. The element values ​​of the initial adjacency matrix represent the initial connection strength between the corresponding nodes.

[0068] Road network node connectivity information can be obtained in various ways, such as from data sources like map data and traffic management systems. This information includes the connection relationships between road segments, i.e., which segments are directly connected and the spatial distance between them. An initial adjacency matrix is ​​constructed based on this spatial distance. This initial adjacency matrix is ​​a two-dimensional matrix with rows and columns equal to the number of nodes. Each element of the initial adjacency matrix represents the initial connection strength between the corresponding nodes. When constructing the initial adjacency matrix, for directly connected nodes, the initial connection strength is determined based on their spatial distance. Generally, the closer the spatial distance, the stronger the initial connection strength; the farther the spatial distance, the weaker the initial connection strength. For nodes that are not directly connected, the initial connection strength can be set to zero. By constructing the initial adjacency matrix, the road network node connectivity information can be transformed into edge information in a graph structure, providing a foundation for subsequent graph convolution operations.

[0069] Step S330: Input the node feature matrix and the initial adjacency matrix into the graph construction layer of the topology-aware graph neural network to construct a road network defect feature map. The nodes of the road network defect feature map are the set of feature nodes, and the edges are the node connection relationships determined based on the initial adjacency matrix.

[0070] The graph construction layer of a topology-aware graph neural network is a module used to construct the graph structure. After the node feature matrix and the initial adjacency matrix are input into the graph construction layer, the layer constructs a road network defect feature map based on the information from the node feature matrix and the initial adjacency matrix.

[0071] The nodes in the road network defect feature graph are a set of feature nodes, namely, feature nodes extracted from the cross-domain feature matrix. Edges represent node connections determined based on the initial adjacency matrix. The element values ​​in the initial adjacency matrix represent the connection strength between nodes; the graph construction layer uses these connection strengths to determine whether edges exist between nodes and their weights. By constructing the road network defect feature graph, the cross-domain feature matrix and road network node connectivity information can be integrated to obtain a complete graph structure, providing a foundation for subsequent graph feature aggregation processing.

[0072] Step S340: Call the graph convolutional layer of the topology-aware graph neural network to perform graph feature aggregation processing on the road network defect feature map. Through the multiplication operation of the adjacency matrix and the node feature matrix, the features of the neighboring nodes of each node are aggregated to the current node to generate an aggregated feature matrix.

[0073] For example, step S340 may specifically include the following steps S341 to S346: Step S341: Input the node feature matrix of the road network defect feature map and the updated adjacency matrix into the feature transformation module of the graph convolutional layer, perform linear transformation on the node feature matrix, and map the node features to a higher-dimensional feature space through the multiplication operation of the weight matrix and the node feature matrix to generate the transformed node feature matrix.

[0074] The feature transformation module of the graph convolutional layer performs linear transformations on node features. The node feature matrix of the road network defect feature map and the updated adjacency matrix are input into the feature transformation module. The updated adjacency matrix is ​​obtained during the iteration process based on the update of node association strength. In the feature transformation module, the node features are linearly transformed through multiplication of the weight matrix and the node feature matrix. The weight matrix is ​​a learnable parameter in the feature transformation module, which maps the node features to a higher-dimensional feature space. This linear transformation increases the expressive power of the node features, enabling them to better capture information from the graph structure. After the linear transformation, a transformed node feature matrix is ​​generated. The transformed node feature matrix has a higher feature dimension than the original node feature matrix, containing more feature information and providing a richer feature representation for subsequent feature aggregation operations.

[0075] Step S342: Perform symmetric normalization on the updated adjacency matrix, calculate the sum of the updated adjacency matrix and the identity matrix to obtain the adjacency matrix with self-loops, and perform degree matrix normalization on the adjacency matrix with self-loops to generate a symmetric normalized adjacency matrix.

[0076] Symmetric normalization is performed to improve the stability and interpretability of the adjacency matrix during feature aggregation. First, the updated adjacency matrix is ​​summed with the identity matrix to obtain an adjacency matrix with self-loops. The identity matrix is ​​a matrix where all diagonal elements are 1s and the rest are 0s. Adding self-loops to the adjacency matrix ensures that each node considers its own feature information, preventing the loss of its own features during feature aggregation.

[0077] Next, the adjacency matrix with self-loops is normalized to its degree. The degree matrix is ​​a diagonal matrix where the elements on the diagonal represent the degree of each node, i.e., the number of edges connected to that node. Degree normalization involves left-multiplying the adjacency matrix with self-loops by the inverse square root of the degree matrix and then right-multiplying it. This normalization ensures that the element values ​​of the adjacency matrix are comparable across different nodes, avoiding imbalances in feature aggregation due to differences in node degrees. Finally, a symmetric normalized adjacency matrix is ​​generated, which can more effectively transmit information between nodes during feature aggregation.

[0078] Step S343: Perform matrix multiplication on the symmetric normalized adjacency matrix and the transformed node feature matrix to complete the aggregation operation of the neighborhood node features and generate a preliminary aggregated feature matrix.

[0079] The core step in graph feature aggregation is matrix multiplication, which involves combining the symmetric normalized adjacency matrix with the transformed node feature matrix. The symmetric normalized adjacency matrix represents the connectivity and strength between nodes, while the transformed node feature matrix represents the features of each node. In the matrix multiplication, for each node, the transformed features of its neighbors are weighted and summed based on the connectivity strength between that node and its neighbors in the symmetric normalized adjacency matrix, resulting in a new feature vector that serves as the initial aggregated feature for the current node. This process is repeated for all nodes, ultimately generating the initial aggregated feature matrix. This matrix contains the initial aggregated features of each node and represents the result of aggregating the features of its neighbors to the current node. The initial aggregated feature matrix provides a more comprehensive representation of the node's feature information, laying the foundation for subsequent nonlinear activation processing.

[0080] Step S344: Perform nonlinear activation processing on the preliminary aggregated feature matrix. Based on the ReLU activation function, perform activation operation on each element of the preliminary aggregated feature matrix to enhance the nonlinear expressive power of the features and generate the activated aggregated feature matrix.

[0081] Nonlinear activation is used to enhance the nonlinear expressive power of features. In graph convolutional layers, the ReLU activation function is used to activate each element of the initially aggregated feature matrix. The ReLU activation function is a nonlinear activation function, and its expression is: when the input value is greater than 0, the output value is equal to the input value; when the input value is less than or equal to 0, the output value is 0.

[0082] The ReLU activation function allows for non-linear changes in the elements of the initial aggregated feature matrix. In neural networks, non-linear activation functions introduce non-linearity, enabling the network to learn more complex features and patterns. After non-linear activation, an activated aggregated feature matrix is ​​generated.

[0083] Step S345: Construct feature aggregation residual connections. Perform element-wise addition operations on the transformed node feature matrix and the activated aggregated feature matrix to generate a residual enhanced aggregated feature matrix. The residual enhanced aggregated feature matrix is ​​used to alleviate the gradient vanishing problem in deep networks.

[0084] In the graph convolutional layer, feature aggregation residual connections are constructed, performing element-wise addition between the transformed node feature matrix and the activated aggregated feature matrix. Element-wise addition adds corresponding elements from the transformed and activated aggregated feature matrices to obtain a new feature matrix. By constructing residual connections, the network can more easily learn feature changes during training. In deep networks, as the number of layers increases, gradients gradually decrease during backpropagation, making training difficult. Residual connections allow gradients to be directly passed from later layers to earlier layers, mitigating the vanishing gradient problem. The generated residual-enhanced aggregated feature matrix contains information from both the transformed and activated aggregated feature matrices, providing a more stable representation of node features and laying the foundation for subsequent batch normalization.

[0085] Step S346: Perform batch normalization on the residual enhancement aggregated feature matrix, calculate the mean and variance of each feature channel, standardize the feature values, and generate an aggregated feature matrix. The feature distribution of the aggregated feature matrix is ​​more stable.

[0086] Batch normalization is performed to stabilize the feature distribution of the feature matrix. In batch normalization, the mean and variance of each feature channel in the residual augmented aggregated feature matrix are calculated. The mean represents the average eigenvalue of that feature channel, and the variance represents the dispersion of the eigenvalues. Then, the eigenvalues ​​of the residual augmented aggregated feature matrix are standardized. Standardization involves subtracting the mean of that feature channel from each eigenvalue and then dividing by the square root of the variance of that feature channel. Standardization ensures that the eigenvalues ​​of each feature channel have the same mean and variance, thus stabilizing the feature distribution. In neural networks, a stable feature distribution accelerates network training and improves generalization ability. After batch normalization, the final aggregated feature matrix is ​​generated. The more stable feature distribution of the aggregated feature matrix provides more reliable feature information for subsequent updates to node association strength.

[0087] Step S350: Calculate the node association strength update value based on the aggregated feature matrix, determine the association strength adjustment coefficient by comparing the similarity of node features before and after aggregation, and perform iterative update processing on the initial adjacency matrix based on the association strength adjustment coefficient to generate the updated adjacency matrix.

[0088] Calculating the node association strength update value based on the aggregated feature matrix is ​​to dynamically adjust the association strength between nodes. In a graph structure, the association strength between nodes reflects the tightness of the connection between nodes. As graph features are aggregated and updated, the association strength between nodes also needs to be adjusted accordingly.

[0089] The association strength adjustment coefficient is determined by comparing the similarity of node features before and after aggregation. The similarity of node features before and after aggregation can be obtained by calculating the distance or similarity metric between node feature vectors. For example, Euclidean distance, cosine similarity, etc., can be used to calculate similarity. The association strength adjustment coefficient is a coefficient used to adjust the association strength between nodes, typically ranging from 0 to 1. Higher similarity results in a larger association strength adjustment coefficient, indicating that the association strength between nodes needs to be strengthened; lower similarity results in a smaller association strength adjustment coefficient, indicating that the association strength between nodes needs to be weakened. The initial adjacency matrix is ​​iteratively updated based on the association strength adjustment coefficient. In each iteration, the association strength adjustment coefficient is multiplied by the corresponding element of the initial adjacency matrix to obtain a preliminary adjusted adjacency matrix. Then, the preliminary adjusted adjacency matrix undergoes further processing, such as normalization, to finally generate the updated adjacency matrix. The updated adjacency matrix reflects the new association strength between nodes, providing a foundation for subsequent graph feature aggregation processing.

[0090] For example, step S350 may specifically include the following steps S351 to S356: Step S351: Calculate the similarity between the aggregated feature matrix and the node feature matrix before the update. For each node, calculate the similarity value between the feature vector of the node in the aggregated feature matrix and the feature vector in the node feature matrix before the update, and generate a similarity matrix.

[0091] Calculating the similarity between the aggregated feature matrix and the node feature matrix before the update is to measure the degree of change in a node before and after feature aggregation. For each node, the similarity value is calculated between the node's eigenvector in the aggregated feature matrix and the eigenvector in the node's feature matrix before the update. Similarity values ​​can be calculated using various methods, such as Euclidean distance and cosine similarity.

[0092] The similarity values ​​of each node are combined to generate a similarity matrix. The number of rows and columns in the similarity matrix is ​​equal to the number of nodes, and each element in the matrix represents the similarity value of the corresponding node before and after feature aggregation.

[0093] Step S352: Map each element value in the similarity matrix to a preset interval, perform non-linear transformation on the similarity values, and generate an association strength adjustment coefficient. The value range of the association strength adjustment coefficient is 0 to 1.

[0094] Mapping each element in the similarity matrix to a preset interval transforms the similarity values ​​into values ​​suitable for use as association strength adjustment coefficients. The preset interval is typically between 0 and 1, and the association strength adjustment coefficient also ranges from 0 to 1. Applying a non-linear transformation to the similarity values ​​ensures that the association strength adjustment coefficient better reflects the relationships between nodes. This non-linear transformation can be achieved using non-linear functions such as the Sigmoid function and the Tanh function. These non-linear functions map similarity values ​​to the interval between 0 and 1 and exhibit non-linear variation characteristics.

[0095] Through nonlinear transformation, each element value in the similarity matrix is ​​converted into an association strength adjustment coefficient. The association strength adjustment coefficient represents the degree of adjustment of the association strength between nodes. The closer the value is to 1, the stronger the association between nodes needs to be; the closer the value is to 0, the weaker the association between nodes needs to be.

[0096] Step S353: Multiply the association strength adjustment coefficient with the corresponding element of the initial adjacency matrix to obtain the preliminary adjusted adjacency matrix. The element value of the preliminary adjusted adjacency matrix is ​​the product of the initial connection strength and the association strength adjustment coefficient.

[0097] Multiplying the association strength adjustment coefficient by the corresponding element of the initial adjacency matrix is ​​to perform an initial adjustment to the initial adjacency matrix based on the association strength adjustment coefficient. The initial adjacency matrix represents the initial connection strength between nodes, and the association strength adjustment coefficient represents the degree of adjustment of the association strength between nodes.

[0098] In the multiplication operation, for each element in the initial adjacency matrix, it is multiplied by the corresponding association strength adjustment coefficient to obtain the element value at the corresponding position in the preliminary adjusted adjacency matrix. The element values ​​of the preliminary adjusted adjacency matrix are the product of the initial connection strength and the association strength adjustment coefficient. In this way, the connection strength between nodes can be dynamically adjusted according to changes in node characteristics. The preliminary adjusted adjacency matrix provides the foundation for subsequent row normalization processing.

[0099] Step S354: Perform row normalization on the preliminary adjusted adjacency matrix, calculate the sum of the elements in each row, divide the value of each element in a row by the sum of the elements in that row, so that the sum of the elements in each row is 1, and generate a probabilistic adjacency matrix.

[0100] Row normalization of the initial adjacency matrix is ​​performed to give it a probabilistic distribution property. In a graph structure, row normalization makes the adjacency relationship of each node probabilistic, meaning that the sum of the elements in each row being 1 indicates that the sum of the connection strengths between that node and all its neighboring nodes is 1.

[0101] Specifically, the sum of the elements in each row of the initially adjusted adjacency matrix is ​​calculated. Then, the value of each row element is divided by the sum of the elements in that row to obtain a new matrix, namely the probabilistic adjacency matrix. The sum of the elements in each row of the probabilistic adjacency matrix is ​​1, and each element represents the connection probability between that node and its corresponding neighboring nodes. Through row normalization, the adjacency matrix can be made to better conform to the characteristics of a probability distribution, providing a basis for subsequent regularization processing.

[0102] Step S355: Construct the association strength regularization term, calculate the regularization loss value based on the Frobenius norm of the probabilistic adjacency matrix, and add the regularization loss value to the association strength update objective function.

[0103] The association strength regularization term is constructed to prevent the elements of the adjacency matrix from deviating excessively from their normal range, thus avoiding overfitting. The regularization loss is calculated based on the Frobenius norm of the probabilistic adjacency matrix. The Frobenius norm is the matrix norm, defined as the square root of the sum of the squares of all elements in the matrix. Including the regularization loss in the association strength update objective function allows for the consideration of regularization constraints during adjacency matrix updates. The association strength update objective function is typically a function used to measure the effectiveness of adjacency matrix updates; by incorporating the regularization loss, the updates can be made more stable and reasonable. During training, an optimization algorithm minimizes the association strength update objective function to obtain the optimal adjacency matrix. Constructing the association strength regularization term improves the model's generalization ability and stability.

[0104] Step S356: Minimize the association strength update objective function using the gradient descent algorithm, iteratively adjust the element values ​​of the probabilistic adjacency matrix until the regularization loss value is less than the preset threshold, and use the probabilistic adjacency matrix at this time as the updated adjacency matrix.

[0105] Minimizing the association strength update objective function using gradient descent aims to find the adjacency matrix element values ​​that minimize the objective function. In each iteration, the gradient of the association strength update objective function with respect to the adjacency matrix elements is calculated. The gradient represents the rate of change of the objective function at the current adjacency matrix element value. Then, based on the direction and magnitude of the gradient, the adjacency matrix element values ​​are adjusted. The adjustment step size is typically controlled by the learning rate, a hyperparameter that needs to be adjusted according to specific circumstances.

[0106] The elements of the probabilistic adjacency matrix are iteratively adjusted until the regularization loss value is less than a preset threshold. The preset threshold is a pre-defined standard; when the regularization loss value is less than this threshold, it indicates that the adjacency matrix update has reached a relatively ideal state. At this point, the probabilistic adjacency matrix is ​​used as the updated adjacency matrix.

[0107] Step S360: Repeat the graph feature aggregation and adjacency matrix update process a preset number of times until the node association strength converges. Use the final node feature matrix as the topology embedding vector, which contains dynamic association information between nodes.

[0108] Repeating the graph feature aggregation and adjacency matrix update processes a predetermined number of times is to allow the node association strength to gradually converge to a stable value. In each iteration, graph feature aggregation is performed first, aggregating the features of each node's neighboring nodes to the current node and updating the node's feature representation; then, adjacency matrix update is performed, adjusting the association strength between nodes based on changes in node features.

[0109] The preset number of iterations is a pre-defined number of iterations, which usually needs to be adjusted based on the specific dataset and model. During the iteration process, the node association strength will continuously change, and as the number of iterations increases, the node association strength will gradually converge to a stable value. When the node association strength converges, it indicates that the node relationships in the graph structure have reached a relatively stable state. The final node feature matrix is ​​used as the topological embedding vector. The topological embedding vector contains dynamic association information between nodes and is the result obtained after multiple graph feature aggregations and adjacency matrix updates.

[0110] Step S400: Extract vibration spectrum data from the navigation data record corresponding to the topology embedding vector in the joint acquisition data set, perform cross-modal alignment processing on the topology embedding vector and the vibration spectrum data, and generate an alignment enhancement feature vector.

[0111] The jointly acquired dataset includes continuous image sequences of the road surface and corresponding time-period vehicle inertial navigation data. The topology embedding vector is obtained by graph structure modeling of the cross-domain feature matrix, containing the topological structure and feature information of the road network. Vibration spectrum data is extracted from the navigation data records corresponding to the topology embedding vectors in the jointly acquired dataset to correlate the vibration information in the image data and navigation data.

[0112] Vibration spectrum data is a part of navigation data recordings, reflecting the vibration of the vehicle during operation. Vibration spectrum data can be obtained by performing spectral analysis on the vibration acceleration data in the navigation data recordings. Spectral analysis converts the vibration acceleration data from a time-domain signal to a frequency-domain signal, obtaining the amplitude values ​​of different frequency components.

[0113] Cross-modal alignment of topological embedding vectors and vibration spectrum data aims to match and fuse data from different modes. This alignment aligns the topological embedding vectors and vibration spectrum data in both time and feature dimensions, allowing for better utilization of information from both datasets. Cross-modal alignment generates aligned and enhanced feature vectors.

[0114] For example, step S400 may specifically include the following steps S410 to S460: Step S410: Filter the navigation data records corresponding to the topology embedding vector from the joint acquisition data set, and match the vehicle inertial navigation data with the same time stamp according to the generation timestamp of the topology embedding vector to obtain the corresponding navigation data records.

[0115] The process of filtering navigation data records corresponding to topology embedding vectors from the jointly acquired dataset ensures that the navigation data records and topology embedding vectors correspond in time. The generation timestamp of the topology embedding vector records the specific time of its generation. Based on this timestamp, the system searches for vehicle inertial navigation data with the same timestamp in the jointly acquired dataset.

[0116] In the jointly acquired dataset, each navigation data record has a corresponding timestamp. By comparing timestamps, navigation data records with the same timestamp as the topology embedding vector generation time are selected. These navigation data records contain the vehicle's motion state and position information at the time of topology embedding vector generation, especially the vibration acceleration data, which provides the foundation for subsequent vibration spectrum data generation. By selecting the corresponding navigation data records, the temporal consistency of the data used in subsequent processing can be ensured, improving the accuracy and reliability of the data.

[0117] Step S420: Perform spectral analysis on the vibration acceleration data in the navigation data record, convert the vibration acceleration data from a time domain signal to a frequency domain signal, and generate vibration spectrum data, which contains the amplitude values ​​of different frequency components.

[0118] Spectral analysis of vibration acceleration data in navigation data records aims to convert the time-domain signal into a frequency-domain signal, thereby obtaining the distribution of vibration acceleration data at different frequencies. A time-domain signal is a signal with time as its independent variable, reflecting the signal's changes over time; a frequency-domain signal is a signal with frequency as its independent variable, reflecting the energy distribution of the signal at different frequencies.

[0119] For example, step S420 may specifically include the following steps S421 to S426: Step S421: Preprocess the vibration acceleration data in the navigation data record, remove the DC component from the vibration acceleration data, filter the vibration acceleration data, and retain the AC component.

[0120] Preprocessing vibration acceleration data in navigation data records improves the accuracy and reliability of spectrum analysis. The DC component, a constant offset in the vibration acceleration data, does not reflect actual vibration changes and therefore needs to be removed. This can be achieved by calculating the average value of the vibration acceleration data and then subtracting this average from each data point. Filtering removes noise and unwanted frequency components from the vibration acceleration data, retaining the AC component. The AC component is the time-varying part of the vibration acceleration data, reflecting the actual vibration conditions. Filtering can be implemented using various filters, such as low-pass filters, high-pass filters, and band-pass filters. Low-pass filters remove high-frequency noise, high-pass filters remove low-frequency noise, and band-pass filters retain only signals within a specified frequency range. Appropriate filters are selected based on specific requirements and the characteristics of the vibration acceleration data. Through preprocessing, the DC component and noise are removed, and the AC component is retained, making the vibration acceleration data more suitable for spectrum analysis.

[0121] Step S422: The preprocessed vibration acceleration data is segmented according to a preset time window to generate multiple continuous vibration data segments, each with the same length.

[0122] The preprocessed vibration acceleration data is segmented according to a preset time window to facilitate spectral analysis. The preset time window is a pre-defined period of time; the vibration acceleration data is divided into multiple consecutive vibration data segments according to this period. Each vibration data segment has the same length, ensuring consistent processing for each segment in subsequent spectral analysis. The purpose of segmentation is to decompose long vibration acceleration data sequences into multiple shorter data segments, allowing for independent spectral analysis of each segment. In practice, data segments can be extracted sequentially according to the preset time window length, starting from the beginning of the vibration acceleration data, until the end of the data stream.

[0123] Step S423: Window each vibration data segment. The vibration data segments are weighted based on the Hamming window function to reduce the spectral leakage effect and generate the windowed data segments.

[0124] Windowing is applied to each vibration data segment to reduce spectral leakage. Spectral leakage occurs during spectral analysis because the signal has a finite time domain, causing leakage in the frequency domain. Windowing reduces spectral leakage by weighting the vibration data segments, gradually reducing the weights at both ends of the segment.

[0125] The vibration data segment is weighted using the Hamming window function. The Hamming window function is a general-purpose window function. In windowing, the Hamming window function is multiplied element-wise with the vibration data segment to obtain the windowed data segment. The amplitude of the windowed data segment gradually decreases at both ends. This reduces spectral leakage during Fast Fourier Transform (FFT) and improves the accuracy of spectral analysis.

[0126] Step S424: Perform Fast Fourier Transform on the windowed data segment to convert the vibration acceleration signal in the time domain into a complex spectrum in the frequency domain, calculate the magnitude of the complex spectrum, and generate the amplitude spectrum.

[0127] After performing a Fast Fourier Transform (FFT), a complex spectrum in the frequency domain is obtained. This complex spectrum contains the real and imaginary parts of each frequency component, representing the amplitude and phase information of the signal at that frequency. To obtain the amplitude information, the magnitude of the complex spectrum is calculated. The magnitude of a complex number is the square root of the sum of the squares of its real and imaginary parts. By calculating the magnitude of the complex spectrum for each frequency component, the amplitude spectrum is obtained.

[0128] Step S425: Extract the frequency axis and amplitude value from the amplitude spectrum, divide the frequency axis into multiple frequency intervals, calculate the average amplitude value in each frequency interval, and generate vibration spectrum data. The horizontal axis of the vibration spectrum data is the frequency interval, and the vertical axis is the average amplitude value.

[0129] Extracting the frequency axis and amplitude values ​​from the amplitude spectrum is for further processing of the amplitude spectrum to obtain vibration spectrum data. The frequency axis represents the position of different frequencies in the amplitude spectrum, and the amplitude value represents the magnitude of the amplitude at each frequency.

[0130] Dividing the frequency axis into multiple frequency intervals allows for a more detailed analysis of the amplitude spectrum. The frequency intervals can be adjusted based on specific needs and the characteristics of the vibration signal. Within each frequency interval, the average amplitude value is calculated. This is done by summing all amplitude values ​​within the interval and then dividing by the number of amplitude values ​​in that interval, thus obtaining the average amplitude value. This method integrates detailed information from the amplitude spectrum, providing a more concise representation of the vibration signal's characteristics across different frequency ranges.

[0131] By using the frequency range as the horizontal axis and the average amplitude value as the vertical axis, vibration spectrum data is generated. This vibration spectrum data can intuitively display the energy distribution of the vibration acceleration signal in different frequency ranges, facilitating subsequent analysis and processing of vibration characteristics.

[0132] Step S426: Smooth the vibration spectrum data. The amplitude values ​​of the vibration spectrum data are averaged using a moving average filter to generate smoothed vibration spectrum data.

[0133] The purpose of smoothing vibration spectrum data is to reduce noise and fluctuations in the data, making the spectrum data smoother and easier to analyze. Moving average filters are a commonly used smoothing method, achieving a smoothing effect by averaging the data through a sliding window.

[0134] Specifically, for each amplitude value in the vibration spectrum data, a sliding window of a preset size is selected centered on that value. This sliding window includes the amplitude value and several amplitude values ​​before and after it. The average value of all amplitude values ​​within the sliding window is calculated and used as the new amplitude value at that position. This process is repeated for each amplitude value in the vibration spectrum data to complete the sliding window averaging process. The smoothing effect of the moving average filter effectively suppresses high-frequency noise and random fluctuations in the vibration spectrum data, highlighting the main characteristics of the vibration signal.

[0135] Step S430: Perform feature extraction processing on the vibration spectrum data, extract the peak frequency, spectral centroid and bandwidth parameters from the vibration spectrum data, and generate a vibration spectrum feature vector. The dimension of the vibration spectrum feature vector is consistent with the feature dimension of the topological embedding vector.

[0136] Feature extraction processing of vibration spectrum data aims to extract key parameters that represent the characteristics of the vibration signal. The peak frequency is the frequency with the largest amplitude in the vibration spectrum data, reflecting the main frequency components of the vibration signal and representing the frequency location where the vibration signal energy is most concentrated. The centroid of the spectrum is the energy center location of the vibration spectrum, calculated by weighted averaging of frequency and amplitude, taking into account the distribution of the vibration spectrum across various frequencies. The centroid of the spectrum reflects the overall frequency characteristics of the vibration signal and embodies the distribution bias of vibration energy along the frequency axis.

[0137] The bandwidth parameter describes the width of the vibration spectrum, and is typically determined by calculating the range of frequencies in the spectrum with amplitude values ​​greater than a certain threshold. The bandwidth parameter reflects the frequency distribution range of the vibration signal, embodying its frequency richness. The extracted peak frequency, spectral centroid, and bandwidth parameter are combined to generate a vibration spectrum feature vector. For effective cross-modal alignment with the topological embedding vector, the dimension of the vibration spectrum feature vector needs to match the feature dimension of the topological embedding vector. Dimensional matching can be achieved through appropriate feature selection, dimensionality reduction, or augmentation methods, ensuring that the two feature vectors can be compared and fused in the same dimensional space.

[0138] Step S440: Input the topology embedding vector and the vibration spectrum feature vector into the dynamic time warping module, calculate the feature distance matrix between the two, find the optimal alignment path through the dynamic programming algorithm, and generate the alignment path matrix.

[0139] For each time point in the topological embedding vector and the vibration spectrum feature vector, the distance between their feature vectors at that time point is calculated, for example, using a metric such as Euclidean distance. Combining the feature distance calculations for all time points yields the feature distance matrix. Next, a dynamic programming algorithm is used to find the optimal alignment path within this matrix. The dynamic programming algorithm calculates the cumulative distance step-by-step, starting from the beginning of the feature distance matrix and selecting the path with the smallest cumulative distance according to certain rules, until the end of the matrix is ​​reached. During this process, each step is recorded, ultimately forming an optimal alignment path. The matrix coordinates of the optimal alignment path are recorded, generating an alignment path coordinate sequence. This sequence is then converted into an alignment path matrix. The alignment path matrix is ​​a binary matrix; elements with a value of 1 indicate an alignment relationship between the topological embedding vector and the vibration spectrum feature vector at the corresponding time point, while elements with a value of 0 indicate no alignment relationship. This alignment path matrix will guide subsequent time-dimensional alignment processing of the topological embedding vector and the vibration spectrum feature vector.

[0140] For example, step S440 may specifically include the following steps S441 to S446: Step S441: Convert the topology embedding vector and the vibration spectrum feature vector into time series features. The time dimension of the topology embedding vector corresponds to the timestamp of image acquisition, and the time dimension of the vibration spectrum feature vector corresponds to the timestamp of vibration data acquisition.

[0141] To process the topology embedding vectors and vibration spectrum feature vectors in the dynamic time warping module, they need to be converted into time series features. The topology embedding vectors are generated based on image data, and their time dimension corresponds to the timestamp of image acquisition. This means that each element of the topology embedding vector corresponds to image information acquired at a specific time.

[0142] The vibration spectrum feature vector is generated based on vibration data, and its time dimension corresponds to the timestamp of vibration data acquisition. That is, each element of the vibration spectrum feature vector represents the vibration characteristics at a specific moment. Through this correspondence, the topological embedding vector and the vibration spectrum feature vector are transformed into time series features with a clear temporal order, making them comparable and aligned in the time dimension.

[0143] Step S442: Calculate the feature distance matrix between the topological embedding vector and the vibration spectrum feature vector. For each time point in the time series, calculate the Euclidean distance between the feature vector of the topological embedding vector at that time and the feature vector of the vibration spectrum feature vector at each time point, and generate the distance matrix.

[0144] After converting the topological embedding vector and the vibration spectrum feature vector into time series features, their feature distance matrix is ​​calculated. For each time point in the time series, the feature vectors of the topological embedding vector at that time and the feature vectors of the vibration spectrum feature vector at each time point are extracted. The Euclidean distance between these two feature vectors is calculated. Euclidean distance is a commonly used distance metric, which measures the similarity between two vectors by calculating the square root of the sum of the squares of the differences between corresponding elements. This calculation is performed sequentially for each time point of the topological embedding vector, and all the calculated Euclidean distances are combined to obtain a matrix, i.e., the distance matrix. Each element in the distance matrix represents the distance between the feature vector of the topological embedding vector at one time point and the feature vector of the vibration spectrum feature vector at another time point, reflecting their degree of difference in the feature space.

[0145] Step S443: Initialize the dynamic programming cumulative distance matrix. The element values ​​of the cumulative distance matrix represent the minimum cumulative distance from the start time to the current time. The first row and first column of the cumulative distance matrix are initialized by adding the corresponding elements of the cumulative distance matrix.

[0146] The cumulative distance matrix in dynamic programming records the minimum cumulative distance from the starting time to each time point, and is the foundation for the dynamic programming algorithm to find the optimal alignment path. The cumulative distance matrix needs to be initialized before starting the dynamic programming process.

[0147] The cumulative distance matrix has the same size as the distance matrix, and its elements are initially set to infinity, except for the first row and first column. The values ​​of the elements in the first row and first column are determined by summing the corresponding elements of the cumulative distance matrix. Specifically, the first element of the first row of the cumulative distance matrix is ​​equal to the first element of the distance matrix, and the other elements of the first row are obtained by adding the previous element of the cumulative distance matrix to the current element of the distance matrix; similarly, the elements of the first column are obtained by summing the elements of the first column of the cumulative distance matrix. This initialization method provides the starting conditions for the subsequent calculations of the dynamic programming algorithm.

[0148] Step S444: Fill the cumulative distance matrix using a dynamic programming algorithm. For each element in the cumulative distance matrix, select the minimum value as the cumulative distance value of the current element based on the sum of the cumulative distance values ​​of the elements to its left, above, and above its left.

[0149] After initializing the cumulative distance matrix, a dynamic programming algorithm is used to fill the entire matrix. For each element in the cumulative distance matrix (except the first row and first column), the cumulative distance values ​​of its three adjacent elements (left, top, and top-left) are considered. The sum of these three cumulative distance values ​​is then added to the sum of the corresponding element's value in the current distance matrix. The minimum sum is selected as the cumulative distance value for the current element. This calculation method ensures that each element in the cumulative distance matrix records the minimum cumulative distance from the initial time step to that time step, reflecting the optimal path selected during dynamic programming. This process is repeated, calculating and updating each element in the cumulative distance matrix sequentially, until the entire matrix is ​​filled. The filled cumulative distance matrix provides crucial information for subsequently finding the optimal alignment path.

[0150] Step S445: Start backtracking from the bottom right element of the cumulative distance matrix to find the optimal alignment path, record the matrix coordinate points passed by the path, and generate the alignment path coordinate sequence.

[0151] After the cumulative distance matrix is ​​filled, backtracking begins from the bottom right element to find the optimal alignment path. Since each element in the cumulative distance matrix records the minimum cumulative distance from the start time to that time, the backtracking process uses this minimum cumulative distance information to find the previous optimal choice. Starting from the bottom right element, the cumulative distance values ​​of its three adjacent elements (left, top, and top left) are compared. The adjacent element with the smallest cumulative distance value is selected as the previous point for backtracking, and its matrix coordinates are recorded. Backtracking continues from this point, repeating the comparison and selection process until the top left element of the cumulative distance matrix is ​​reached. During the backtracking process, the matrix coordinates of the points traversed by the path are recorded sequentially, resulting in the alignment path coordinate sequence.

[0152] Step S446: Convert the alignment path coordinate sequence into an alignment path matrix. The alignment path matrix is ​​a binary matrix. Elements with a value of 1 in the matrix represent the alignment relationship at the corresponding time point. The alignment path matrix is ​​used to guide the alignment of the topology embedding vector with the vibration spectrum feature vector in the time dimension.

[0153] For each coordinate point in the alignment path coordinate sequence, set the corresponding element in the alignment path matrix to 1, indicating that the topological embedding vector and the vibration spectrum feature vector have an alignment relationship at that time point; set the elements at the remaining positions to 0, indicating that there is no alignment relationship.

[0154] In this way, the alignment path coordinate sequence is transformed into an intuitive alignment path matrix. The alignment path matrix clearly shows the alignment of the topological embedding vector and the vibration spectrum feature vector in the time dimension, providing clear guidance for subsequent rearrangement and alignment operations of the topological embedding vector and the vibration spectrum feature vector based on this matrix.

[0155] Step S450: Perform time-dimensional alignment processing on the topology embedding vector and the vibration spectrum feature vector according to the alignment path matrix, and rearrange the feature sequence of the topology embedding vector and the feature sequence of the vibration spectrum feature vector according to the alignment path matrix to generate the aligned dual-modal feature matrix.

[0156] Aligning the topological embedding vector and the vibration spectrum feature vector in the time dimension according to the alignment path matrix is ​​to align the features of these two different modes in time, enabling effective feature fusion later. The feature sequences of the topological embedding vector and the vibration spectrum feature vector are rearranged according to the alignment relationship indicated by the elements with a value of 1 in the alignment path matrix. Specifically, the alignment path matrix determines which time-sequence elements in the topological embedding vector and the vibration spectrum feature vector should correspond together, and then these corresponding elements are rearranged and combined. After rearrangement, an aligned bimodal feature matrix is ​​generated. This matrix effectively aligns the topological embedding vector and the vibration spectrum feature vector in the time dimension, allowing their features to be compared and analyzed on the same time scale, providing a good data foundation for subsequent feature fusion processing.

[0157] Step S460: Perform feature fusion processing on the aligned dual-modal feature matrix. Based on the feature connection operation, fuse the aligned topological embedding vector with the vibration spectrum feature vector to generate an aligned enhanced feature vector. The aligned enhanced feature vector contains cross-modal correlation information between the image and vibration.

[0158] Feature fusion processing of the aligned bimodal feature matrix aims to integrate the information from the topological embedding vector and the vibration spectrum feature vector, thereby uncovering cross-modal correlations between them. Feature concatenation is a commonly used feature fusion method, which achieves feature fusion by concatenating the aligned topological embedding vector and the vibration spectrum feature vector in a specific manner. Specifically, the aligned topological embedding vector and the vibration spectrum feature vector can be concatenated row-wise or column-wise to obtain a new vector, namely the aligned enhanced feature vector. This vector integrates the information from the topological embedding vector generated from image data and the vibration spectrum feature vector generated from vibration data, containing cross-modal correlation information between the image and vibration.

[0159] Step S500: Perform joint identification processing of disease type and severity on the aligned and enhanced feature vector to generate road network disease identification results. The road network disease identification results include disease type labels and severity descriptors.

[0160] For example, step S500 may specifically include the following steps S510 to S560: Step S510: Input the alignment enhancement feature vector into the disease type identification sub-network. The disease type identification sub-network contains multiple fully connected layers. Perform nonlinear transformation processing on the alignment enhancement feature vector to generate type feature vector.

[0161] The road damage type identification subnetwork is a neural network module specifically designed for identifying road damage types. It consists of multiple fully connected layers, where each neuron is connected to all neurons in the layer above. After the aligned and enhanced feature vector is input into the damage type identification subnetwork, the network performs a non-linear transformation. In the fully connected layers, each neuron performs a linear combination of the input feature vector, i.e., multiplying the input vector by the weight matrix and adding a bias term. The result of this linear combination is then input into a non-linear activation function for non-linear transformation. Non-linear activation functions, such as ReLU and Sigmoid, introduce non-linearity, enabling the network to learn more complex features and patterns.

[0162] Through sequential processing by multiple fully connected layers, the alignment enhancement feature vector is continuously subjected to nonlinear transformations, ultimately generating a type feature vector. This type feature vector, after processing by the disease type identification sub-network, better reflects the characteristics of the disease type, providing a foundation for disease type classification.

[0163] Step S520: Call the type classification layer of the disease type identification sub-network to classify the type feature vector, calculate the probability distribution of different disease types, and generate disease type probability vectors. The element values ​​of the disease type probability vectors represent the confidence level of the corresponding disease type.

[0164] The type classification layer of the disease type identification subnetwork is responsible for classifying type feature vectors and determining the probability of different disease types. The type classification layer typically contains hidden layers and an output layer for further transformation and mapping of the type feature vectors. In the type classification layer, some preprocessing operations are first performed on the type feature vectors to improve classification accuracy. Then, the hidden layers perform deeper feature extraction and transformation on the type feature vectors to enhance the discriminative power of the features. In the output layer, the probability distribution of different disease types is calculated. This is usually implemented using classification algorithms such as the Softmax function. The Softmax function converts the raw scores of the output layer into probability values, such that the sum of the probability values ​​of all disease types is 1. Each probability value represents the confidence level of the corresponding disease type, i.e., the likelihood of that disease type occurring. These probability values ​​are combined to generate a disease type probability vector. Each element in the disease type probability vector corresponds to a disease type, and its value reflects the confidence level of that disease type.

[0165] For example, step S520 may specifically include the following steps S521 to S526: Step S521: Perform batch normalization on the type feature vector, calculate the mean and variance of each feature channel of the type feature vector, standardize the feature values, and generate normalized type feature vectors.

[0166] Batch normalization of categorical feature vectors aims to improve the training stability and convergence speed of neural networks. Batch normalization standardizes the feature values ​​by calculating the mean and variance of each feature channel within the categorical feature vector.

[0167] Specifically, for each feature channel of the categorical feature vector, the mean and variance of all elements in that channel are calculated. Then, the mean is subtracted from each feature value of that channel, and then divided by the square root of the variance to obtain the standardized feature value. This method ensures that the feature values ​​of each feature channel have the same mean and variance, reducing correlation and scale differences between features. The generated normalized categorical feature vectors propagate more stably in subsequent network processing, avoiding problems such as vanishing or exploding gradients, thus improving the training efficiency and performance of the network.

[0168] Step S522: Input the normalized type feature vector into the hidden layer of the type classification layer, and perform a linear transformation by multiplying the weight matrix and the normalized type feature vector to generate a linearly transformed feature vector.

[0169] After the normalized feature vector is input into the hidden layer of the classification layer, the hidden layer performs a linear transformation on it. In the hidden layer, each neuron has a set of corresponding weights, which form a weight matrix. The normalized feature vector is linearly combined by multiplying the weight matrix with the normalized feature vector. Each neuron multiplies each element of the input normalized feature vector with its corresponding weight, adds the results, and adds a bias term to obtain the neuron's output. The outputs of all neurons are combined to generate a linearly transformed feature vector. This linearly transformed feature vector is the result of the normalized feature vector undergoing a linear transformation in the hidden layer, representing a mapping and transformation in the feature space, providing input for subsequent nonlinear activation processing.

[0170] Step S523: Perform nonlinear activation processing on the linear transformation feature vector. Based on the Leaky ReLU activation function, perform activation operation on each element of the linear transformation feature vector to enhance the nonlinear discrimination ability of the features and generate activated type feature vectors.

[0171] Applying nonlinear activation to the linearly transformed feature vectors introduces nonlinearity to enhance their nonlinear discriminative power. In the type classification layer, the Leaky ReLU activation function is used to activate each element of the linearly transformed feature vector.

[0172] The Leaky ReLU activation function is an improved version of the ReLU activation function. When the input value is less than 0, it assigns a small non-zero slope, avoiding the problem of neuron death when the input is negative, which is common in traditional ReLU functions. For each element of the linearly transformed feature vector, the activated element value is calculated according to the rules of the Leaky ReLU activation function.

[0173] This nonlinear activation process makes the elements in the linearly transformed feature vector undergo nonlinear changes, which can better represent the complex feature differences between different disease types.

[0174] Step S524: Call the output layer of the type classification layer to perform classification mapping processing on the activated type feature vector, and map the activated type feature vector to the dimension of disease type quantity through the output weight matrix to generate the original score vector.

[0175] The output layer of the type classification layer is responsible for classifying and mapping the activated type feature vectors to the dimension of the number of disease types. The output layer contains an output weight matrix, the size and structure of which are determined by the number of disease types.

[0176] The activated type feature vectors are linearly mapped by multiplying the output weight matrix with the activated type feature vectors. Each disease type corresponds to a row in the output weight matrix, and the result of the multiplication is a vector with the number of elements equal to the number of disease types. This vector is the original score vector, representing the score of the activated type feature vectors on each disease type, but these scores are not yet probability values.

[0177] Step S525: Perform the Softmax function operation on the original score vector to calculate the probability value of each disease type. The probability value is obtained by dividing the exponential function value of the original score vector by the sum of all exponential function values, thus generating a disease type probability vector.

[0178] The purpose of performing the Softmax function operation on the original score vector is to convert the original score into probability values, so that the sum of the probability values ​​of all disease types is 1, which facilitates classification decision-making.

[0179] The Softmax function is calculated by taking the exponential function value for each element of the original score vector and then dividing each exponential function value by the sum of all exponential function values. Each element obtained in this way represents the probability value of the corresponding disease type. These probability values ​​indicate the likelihood that the activated feature vector belongs to each disease type. Combining the probability values ​​of all disease types generates a disease type probability vector. Each element in the disease type probability vector is between 0 and 1, and the sum of all elements is 1, clearly showing the confidence distribution of different disease types, providing an important basis for subsequent disease type selection and decision-making.

[0180] Step S526: Perform confidence screening on the disease type probability vector, retain disease types with probability values ​​greater than the preset confidence threshold, and generate the screened disease type probability vector. The screened disease type probability vector is used for subsequent type decision.

[0181] The confidence screening process for disease type probability vectors aims to remove disease types with low confidence levels, thereby improving the accuracy of disease type identification. A pre-set confidence threshold is a pre-defined standard used to determine whether the confidence level of a disease type is high enough. For each element in the disease type probability vector, if its probability value is greater than the pre-set confidence threshold, the corresponding disease type is retained; if the probability value is less than or equal to the pre-set confidence threshold, the corresponding disease type is excluded. The probability values ​​of the retained disease types are then recombine to generate the filtered disease type probability vector.

[0182] Step S530: Input the alignment enhancement feature vector into the severity assessment sub-network. The severity assessment sub-network uses a convolutional neural network to extract spatial features from the alignment enhancement feature vector to generate a severity feature vector.

[0183] The severity assessment subnetwork is a neural network module used to evaluate the severity of road defects. It employs a convolutional neural network (CNN) to extract spatial features from the aligned and enhanced feature vectors. CNNs have powerful spatial feature extraction capabilities, automatically learning local features and spatial structure within the data.

[0184] After the alignment-enhanced feature vector is input into the severity assessment subnetwork, the convolutional layers in the network use convolutional kernels to perform convolution operations on the input feature vector. The convolutional kernels slide across the input data, extracting and combining features from local regions, resulting in a series of feature maps. These feature maps reflect the spatial characteristics of the alignment-enhanced feature vector at different scales and locations. Pooling layers downsample the feature maps output by the convolutional layers, reducing their size while retaining important feature information. Through multiple convolution and pooling operations, features are continuously extracted and compressed, ultimately generating a severity feature vector. This severity feature vector contains spatial feature information related to the severity of road damage, providing a foundation for subsequent severity assessment.

[0185] Step S540: Perform regression processing on the severity feature vector, and map the severity feature vector to the severity assessment space through a fully connected layer to generate a severity assessment value, which reflects the degree of damage to the disease.

[0186] The purpose of regression processing on the severity feature vector is to map it to a severity assessment space, obtaining an assessment value that reflects the degree of disease damage. Fully connected layers are the key component for achieving this mapping. Each neuron in a fully connected layer is connected to all elements of the severity feature vector, and the severity feature vector is linearly combined through multiplication of the weight matrix with the severity feature vector. Then, a bias term is added to obtain a preliminary prediction value. To improve the accuracy and stability of the prediction, this preliminary prediction value can be further processed. For example, nonlinear activation processing might be performed to introduce nonlinear factors and enhance the model's expressive power. Finally, the processed result is mapped to the severity assessment space to obtain the severity assessment value.

[0187] For example, step S540 may specifically include the following steps S541 to S546: Step S541: Perform feature dimensionality reduction on the severity feature vector. Based on the principal component analysis algorithm, reduce the dimensionality of the severity feature vector, retain the main feature components, and generate a dimensionality-reduced severity feature vector.

[0188] Feature dimensionality reduction of severity feature vectors aims to reduce feature dimensionality, remove redundant information, and retain key feature components. The specific steps of principal component analysis (PCA) are as follows: First, calculate the covariance matrix of the severity feature vectors. The covariance matrix reflects the correlation between the features in the feature vector. Then, solve for the eigenvalues ​​and eigenvectors of the covariance matrix. The eigenvalues ​​represent the importance of the principal components, and the eigenvectors represent the orientation of the principal components. Select the top few principal components with larger eigenvalues ​​and project the severity feature vectors onto these principal components to obtain the dimensionality-reduced feature vectors. This dimensionality-reduced severity feature vector retains the most important information in the severity feature vector while reducing the number of features, lowering the complexity of subsequent processing, and improving the training efficiency and generalization ability of the model.

[0189] Step S542: Input the dimensionality reduction severity feature vector into the first fully connected layer of the severity evaluation sub-network, and perform a linear transformation by multiplying the weight matrix with the dimensionality reduction severity feature vector to generate the first transformed feature vector.

[0190] After the dimensionality-reduced severity feature vector is input into the first fully connected layer of the severity evaluation sub-network, the first fully connected layer performs a linear transformation on it. Each neuron in the first fully connected layer has a corresponding set of weights, which together form a weight matrix.

[0191] The dimensionality-reduced severity feature vector is linearly combined by multiplying the weight matrix with the dimensionality-reduced severity feature vector. Each neuron multiplies each element of the input dimensionality-reduced severity feature vector with its corresponding weight, adds the results together, and adds a bias term to obtain the neuron's output.

[0192] The outputs of all neurons are combined to generate the first transformed feature vector. This first transformed feature vector is the result of the dimensionality reduction severity feature vector being linearly transformed by the first fully connected layer. It undergoes certain mappings and transformations in the feature space, providing input for subsequent nonlinear activation processing.

[0193] Step S543: Perform nonlinear activation processing on the first transformed feature vector, and perform activation operation on the first transformed feature vector based on the ELU activation function to generate the first activated feature vector.

[0194] The nonlinear activation of the first transformed feature vector is applied to introduce nonlinearity and enhance the expressive power of the features. In the severity evaluation subnetwork, the first transformed feature vector is activated using the ELU activation function.

[0195] The ELU activation function is an activation function with a smooth negative half-axis. When the input value is less than 0, it assigns a negative exponential function value, allowing the neuron to still have a certain response even with negative inputs, avoiding the problem of neuron death when the traditional ReLU function has negative inputs. For each element of the first transformed feature vector, the activation value is calculated according to the rules of the ELU activation function. Through this non-linear activation processing, the elements in the first transformed feature vector have non-linear changes, which can better represent the complex features related to the severity of road damage.

[0196] Step S544: Input the first activation feature vector into the second fully connected layer and perform a second linear transformation to generate a second transformed feature vector. The second transformed feature vector has a dimension of 1 and corresponds to the preliminary prediction of the severity assessment value.

[0197] After the first activation feature vector is input into the second fully connected layer, the second fully connected layer performs a second linear transformation on it. The structure and weights of the second fully connected layer are determined according to the severity assessment target, and its output dimension is 1.

[0198] The first activation feature vector is linearly mapped by multiplying its weight matrix with the first activation feature vector. Each neuron multiplies each element of the input first activation feature vector with its corresponding weight, adds the results, and adds a bias term to obtain the neuron's output. Since the second fully connected layer has only one neuron, the final output is a scalar value, which is the second transformation feature vector, corresponding to the preliminary prediction of the severity assessment value.

[0199] Step S545: Construct a severity assessment loss function, calculate the error between the second transformed feature vector and the true severity label based on the mean squared error loss function, and adjust the weight parameters of the second fully connected layer through the backpropagation algorithm.

[0200] The purpose of constructing a severity assessment loss function is to measure the difference between the second-transformation feature vector and the true severity label, so as to improve the accuracy of severity assessment by adjusting the model parameters through optimization algorithms. The mean squared error loss function is a general loss function that measures the magnitude of the error by calculating the average of the squared errors between the predicted value (second-transformation feature vector) and the true value (true severity label).

[0201] Specifically, for each sample, the squared difference between the second transformed feature vector and the true severity label is calculated. Then, the squared errors of all samples are summed, and the result is divided by the number of samples to obtain the mean squared error loss value. This loss value reflects the overall error of the model in severity assessment.

[0202] The backpropagation algorithm adjusts the weight parameters of the second fully connected layer based on the gradient information of the loss function. Starting with the loss function, backpropagation calculates the gradient layer by layer, transmitting the gradient information back to each layer of the network, and updating the weight parameters according to the direction and magnitude of the gradient. Through continuous iteration of this process, the value of the loss function gradually decreases, and the model's prediction results gradually approach the true value, improving the accuracy of severity assessment.

[0203] Step S546: Perform numerical normalization on the second transformed feature vector, map the feature values ​​to a preset severity assessment range, and generate a severity assessment value. The range of the severity assessment value corresponds to the actual severity level of the disease.

[0204] The purpose of numerically normalizing the second transformation feature vector is to map its feature values ​​to a preset severity assessment range, so that the severity assessment value has actual physical meaning and can correspond to the actual severity level of the disease.

[0205] The preset severity assessment range is determined based on the actual severity classification criteria for diseases. For example, it can be divided into numerical ranges corresponding to different levels such as mild, moderate, and severe. The numerical normalization method can be linear mapping, which calculates the minimum and maximum values ​​of the second transformation eigenvector and linearly transforms its eigenvalues ​​to the preset severity assessment range.

[0206] After numerical normalization, a severity assessment value is generated. This severity assessment value takes a value within a preset severity assessment range, and its magnitude directly corresponds to the severity level of the actual road damage. The severity assessment value provides a clear understanding of the extent of road damage, offering a precise reference for road maintenance and management.

[0207] Step S550: Construct a type-severity correlation matrix, model the correlation between the disease type probability vector and the severity assessment value, and fuse the feature information of the two through feature connection operation to generate a joint identification feature vector.

[0208] The purpose of constructing the type-severity correlation matrix is ​​to associate and integrate the disease type identification results and severity assessment results, and to uncover the potential relationship between the two. The disease type probability vector represents the confidence level of different disease types, while the severity assessment value reflects the degree of damage caused by the disease.

[0209] By using feature concatenation operations, the disease type probability vector and severity assessment value are combined. The severity assessment value can be used as a new feature dimension and concatenated with the disease type probability vector row by row or column to obtain a new vector, namely the joint identification feature vector.

[0210] The joint identification feature vector integrates information on both the type and severity of road defects, encompassing not only the probability of different defect types but also information on their severity. Through this correlation modeling and feature fusion, the characteristics of road defects can be described more comprehensively, providing richer evidence for subsequent decision-making.

[0211] Step S560: Perform decision output processing on the joint identification feature vector, extract the disease type with the highest confidence from the disease type probability vector as the identification disease type, and generate the road network disease identification result by combining the corresponding severity assessment value. The road network disease identification result includes disease type label and severity descriptor.

[0212] The final step in the entire road network defect image recognition method is to process the joint identification feature vector for decision output. The aim is to determine the final road network defect recognition result based on the joint identification feature vector. First, the defect type with the highest confidence level is extracted from the defect type probability vector. Each element in the defect type probability vector represents the confidence level of the corresponding defect type. By comparing the values ​​of these elements, the defect type corresponding to the maximum value is identified and used as the recognized defect type. Then, the road network defect recognition result is generated by combining the severity assessment value corresponding to this identified defect type. Based on the preset severity assessment interval where the severity assessment value falls, the corresponding severity descriptor, such as mild, moderate, or severe, is determined.

[0213] It should be noted that the various algorithms involved in the above descriptions of the embodiments of the present invention can all be obtained from relevant content in the prior art. To save space, they will not be elaborated on in the embodiments of the present invention. In addition, those skilled in the art can supplement the details based on common knowledge in the art when implementing the solutions of the present invention. For example, they can use normalization to eliminate dimensional conflicts before feature fusion, use interpolation to eliminate dimensional differences, reasonably set thresholds based on historical data, experience or business scenario requirements, train the model based on a general model training method, set the number of layers in the model structure based on actual needs, select activation functions, etc. The present invention will not provide redundant descriptions of overly detailed implementation processes.

[0214] This invention also provides a computer system, including a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement some or all of the steps in the above-described method.

[0215] Please refer to Figure 3 , Figure 3 This is a hardware entity diagram of a computer system provided in an embodiment of the present invention, such as... Figure 3 As shown, the hardware entities of the computer system 300 include: a processor 310, a communication interface 320, and a memory 330. The processor 310 typically controls the overall operation of the computer system 300. The communication interface 320 enables the computer system to communicate with other terminals or servers via a network. The memory 330 is configured to store instructions and applications executable by the processor 310, and can also cache data to be processed or already processed by the processor 310 and various modules in the computer system 300. It can be implemented using flash memory (FLASH) or random access memory (RAM). Data transmission between the processor 310, the communication interface 320, and the memory 330 can be performed via a bus 340. It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of the present invention, please refer to the description of the method embodiments of the present invention for understanding.

Claims

1. A method for road network defect image recognition based on deep learning, characterized in that, The method includes: A continuous image sequence of the road surface and vehicle inertial navigation data for the corresponding time period are collected. The continuous image sequence and the vehicle inertial navigation data are integrated into a joint acquisition data set through timestamp synchronization processing. Each data unit in the joint acquisition data set contains an image frame at the same acquisition time and the corresponding navigation data record. The image frames in the jointly acquired data set are subjected to spectral-geometric dual-domain decoupling processing to obtain the spectral feature components and geometric structure components of each image frame. The spectral feature components and the geometric structure components are then correlated and mapped to generate a cross-domain feature matrix. The cross-domain feature matrix is ​​input into a pre-trained topology-aware graph neural network, and graph structure modeling is performed on the cross-domain feature matrix in combination with road network node connectivity information. Topology embedding vectors are generated by iteratively updating node association strength. Vibration spectrum data is extracted from navigation data records corresponding to the topology embedding vector in the jointly acquired data set. Cross-modal alignment processing is performed on the topology embedding vector and the vibration spectrum data to generate an alignment enhancement feature vector. The aligned and enhanced feature vectors are subjected to joint identification processing of disease type and severity to generate road network disease identification results.

2. The method according to claim 1, characterized in that, The process involves performing spectral-geometric dual-domain decoupling processing on the image frames in the jointly acquired data set to obtain the spectral feature components and geometric structure components of each image frame. Then, the spectral feature components and geometric structure components are correlated and mapped to generate a cross-domain feature matrix, including: The image frame is subjected to color space conversion processing to convert the image frame from the original color space to the target color space, and multiple spectral channel components are separated. Spatial noise suppression processing is performed on each spectral channel component to generate a noise-reduced spectral channel set. The image frame is processed to extract edge contours, identify the structural edge lines of the road surface in the image frame, calculate the rate of curvature change and directional gradient distribution of the structural edge lines, and generate a geometric structure description vector. The denoised spectral channel set is input into a spectral feature encoder, and a multilayer perceptron is used to perform nonlinear transformation on each spectral channel component to generate spectral feature components, which contain the spectral response intensity distribution of each channel. The geometric structure description vector is subjected to structural feature encoding processing, and spatial dimension features are extracted from the geometric structure description vector based on a convolutional neural network to generate geometric structure components, wherein the geometric structure components include the spatial topological relationship of the edge lines; A spectral-geometric correlation mapping matrix is ​​constructed, and the feature dimensions of the spectral feature components and the geometric structure components are matched and aligned. The spectral feature components and the geometric structure components are correlated and mapped through feature connection operations to generate a cross-domain feature matrix. The row dimension of the cross-domain feature matrix corresponds to the feature channels of the spectral feature components, and the column dimension corresponds to the spatial position of the geometric structure components.

3. The method according to claim 2, characterized in that, The process of performing spatial noise suppression on each spectral channel component to generate a denoised spectral channel set includes: For each spectral channel component, local variance is calculated. A preset sliding window is used to traverse each pixel of the spectral channel component, and the variance value of the pixel value within the window is calculated to generate a variance distribution map. Based on the variance distribution map, a segmentation threshold for noise and non-noise regions is determined. Regions with variance values ​​greater than the segmentation threshold are marked as noise candidate regions, and regions with variance values ​​less than or equal to the segmentation threshold are marked as non-noise regions. Adaptive median filtering is performed on the noise candidate region. The size of the filtering window is adjusted according to the distribution of neighboring pixel values ​​of the pixels in the noise candidate region. The pixel values ​​in the window are sorted and the median value is used to replace the original pixel value. The non-noise region is subjected to guided filtering, and the original spectral channel components are used as the guide map to perform edge-preserving smoothing on the pixel values ​​of the non-noise region. The noisy candidate regions processed by adaptive median filtering are merged with the non-noisy regions processed by guided filtering to generate denoised spectral channel components. All spectral channel components are then combined into a denoised spectral channel set.

4. The method according to claim 2, characterized in that, The calculation of the curvature change rate and directional gradient distribution of the edge lines of the structure to generate a geometric structure description vector includes: The identified road surface structure edge line is processed by sampling point extraction. The coordinates of the sampling points are collected at equal intervals along the edge line to generate a sampling point sequence, which contains a set of coordinate points continuously distributed on the edge line. Calculate the tangent direction vectors of adjacent sampling points in the sampling point sequence, and differentiate the tangent direction vector of each sampling point to obtain the curvature vector. The magnitude of the curvature vector represents the magnitude of the curvature at that sampling point, and the direction represents the direction of curvature change. The rate of change of the magnitude of the curvature vector is calculated by performing a difference operation on the curvature magnitude of continuous sampling points to obtain the rate of change of curvature. The rate of change of curvature is used to characterize the degree of change in the curvature of the edge line. The directional gradient is calculated on the edge line of the structure, and the gradients in the horizontal and vertical directions are calculated on the grayscale image of the image frame to generate a directional gradient map, which includes the gradient magnitude and gradient direction of each pixel. The gradient magnitude and gradient direction at the location of the structural edge line are extracted from the directional gradient map, and a directional gradient distribution matrix is ​​constructed. The curvature change rate and the directional gradient distribution matrix are concatenated as row vectors to generate a geometric structure description vector. Each element of the geometric structure description vector corresponds to a geometric feature parameter of the edge line.

5. The method according to claim 1, characterized in that, The step of inputting the cross-domain feature matrix into a pre-trained topology-aware graph neural network, combining road network node connectivity information to perform graph structure modeling processing on the cross-domain feature matrix, and generating a topology embedding vector by iteratively updating node association strength includes: Extract a set of feature nodes from the cross-domain feature matrix, and use the feature vector of each row of the cross-domain feature matrix as the initial node feature of the graph neural network to generate a node feature matrix. The row dimension of the node feature matrix corresponds to the number of nodes, and the column dimension corresponds to the feature dimension. Obtain road network node connectivity information, which includes the connection relationship between each road segment in the road network and the spatial distance between road segments. Construct an initial adjacency matrix based on the spatial distance, where the element values ​​of the initial adjacency matrix represent the initial connection strength between the corresponding nodes. The node feature matrix and the initial adjacency matrix are input into the graph construction layer of the topology-aware graph neural network to construct a road network defect feature map. The nodes of the road network defect feature map are a set of feature nodes, and the edges are the node connection relationships determined based on the initial adjacency matrix. The graph convolutional layer of the topology-aware graph neural network is invoked to perform graph feature aggregation processing on the road network defect feature map. By multiplying the adjacency matrix and the node feature matrix, the features of the neighboring nodes of each node are aggregated to the current node to generate an aggregated feature matrix. The node association strength update value is calculated based on the aggregated feature matrix. The association strength adjustment coefficient is determined by comparing the similarity of node features before and after aggregation. The initial adjacency matrix is ​​iteratively updated based on the association strength adjustment coefficient to generate the updated adjacency matrix. Repeat the graph feature aggregation and adjacency matrix update process a preset number of times until the node association strength converges. Use the final node feature matrix as the topology embedding vector, which contains dynamic association information between nodes.

6. The method according to claim 5, characterized in that, The process of calculating node association strength update values ​​based on the aggregated feature matrix, determining association strength adjustment coefficients by comparing the similarity of node features before and after aggregation, and iteratively updating the initial adjacency matrix based on the association strength adjustment coefficients to generate an updated adjacency matrix includes: The similarity between the aggregated feature matrix and the node feature matrix before the update is calculated. For each node, the similarity value between the feature vector of the node in the aggregated feature matrix and the feature vector in the node feature matrix before the update is calculated, and a similarity matrix is ​​generated. Each element value in the similarity matrix is ​​mapped to a preset interval, and the similarity values ​​are subjected to nonlinear transformation to generate an association strength adjustment coefficient. The correlation strength adjustment coefficient is multiplied by the corresponding element of the initial adjacency matrix to obtain the preliminary adjusted adjacency matrix. The element value of the preliminary adjusted adjacency matrix is ​​the product of the initial connection strength and the correlation strength adjustment coefficient. The preliminary adjusted adjacency matrix is ​​subjected to row normalization processing. The sum of the elements in each row is calculated, and the value of each element in a row is divided by the sum of the elements in that row so that the sum of the elements in each row is 1, thereby generating a probabilistic adjacency matrix. Construct an association strength regularization term, calculate the regularization loss value based on the Frobenius norm of the probabilistic adjacency matrix, and add the regularization loss value to the association strength update objective function; The association strength update objective function is minimized by using the gradient descent algorithm, and the element values ​​of the probabilistic adjacency matrix are iteratively adjusted until the regularization loss value is less than a preset threshold. The probabilistic adjacency matrix at this point is then used as the updated adjacency matrix.

7. The method according to claim 5, characterized in that, The process of invoking the graph convolutional layer of the topology-aware graph neural network to perform graph feature aggregation processing on the road network defect feature map involves aggregating the features of neighboring nodes of each node to the current node through multiplication of the adjacency matrix and the node feature matrix, generating an aggregated feature matrix, including: The node feature matrix of the road network defect feature map and the updated adjacency matrix are input into the feature transformation module of the graph convolution layer. The node feature matrix is ​​linearly transformed and mapped to the feature space through the multiplication operation of the weight matrix and the node feature matrix to generate the transformed node feature matrix. The updated adjacency matrix is ​​symmetrically normalized, and the sum of the updated adjacency matrix and the identity matrix is ​​calculated to obtain an adjacency matrix with self-loops. The adjacency matrix with self-loops is then normalized to generate a symmetrically normalized adjacency matrix. The symmetric normalized adjacency matrix and the transformed node feature matrix are multiplied together to complete the aggregation operation of the neighborhood node features and generate a preliminary aggregated feature matrix. The preliminary aggregated feature matrix is ​​subjected to nonlinear activation processing. Based on the activation function, each element of the preliminary aggregated feature matrix is ​​activated to enhance the nonlinear expressive power of the features and generate the activated aggregated feature matrix. Construct a feature aggregation residual connection, and perform element-wise addition operation between the transformed node feature matrix and the activated aggregated feature matrix to generate a residual enhanced aggregated feature matrix. The residual enhanced aggregated feature matrix is ​​used to alleviate the gradient vanishing problem in deep networks. The residual enhanced aggregated feature matrix is ​​batch normalized, the mean and variance of each feature channel are calculated, the feature values ​​are standardized, and an aggregated feature matrix is ​​generated. The feature distribution of the aggregated feature matrix is ​​more stable.

8. The method according to claim 1, characterized in that, The step of extracting vibration spectrum data from navigation data records corresponding to the topology embedding vector in the jointly acquired data set, and performing cross-modal alignment processing on the topology embedding vector and the vibration spectrum data to generate an alignment-enhanced feature vector includes: From the jointly acquired data set, navigation data records corresponding to the topology embedding vector are selected. Based on the generation timestamp of the topology embedding vector, vehicle inertial navigation data with the same timestamp are matched to obtain the corresponding navigation data records. The vibration acceleration data in the navigation data record is subjected to spectral analysis to convert the vibration acceleration data from a time domain signal to a frequency domain signal, generating vibration spectrum data, which contains amplitude values ​​of different frequency components. The vibration spectrum data is subjected to feature extraction processing to extract the peak frequency, spectral centroid and bandwidth parameters from the vibration spectrum data, and a vibration spectrum feature vector is generated. The dimension of the vibration spectrum feature vector is consistent with the feature dimension of the topological embedding vector. The topological embedding vector and the vibration spectrum feature vector are input into the dynamic time warping module to calculate the feature distance matrix between them. The optimal alignment path is found through a dynamic programming algorithm to generate the alignment path matrix. The topology embedding vector and the vibration spectrum feature vector are aligned in the time dimension according to the alignment path matrix. The feature sequence of the topology embedding vector and the feature sequence of the vibration spectrum feature vector are rearranged according to the alignment path matrix to generate an aligned dual-modal feature matrix. The aligned dual-modal feature matrix is ​​subjected to feature fusion processing. Based on feature connection operation, the aligned topological embedding vector and the vibration spectrum feature vector are fused to generate an aligned enhanced feature vector. The aligned enhanced feature vector contains cross-modal correlation information between the image and the vibration.

9. The method according to claim 8, characterized in that, The step of performing spectral analysis on the vibration acceleration data in the navigation data record, converting the vibration acceleration data from a time-domain signal to a frequency-domain signal, and generating vibration spectrum data includes: The vibration acceleration data in the navigation data record is preprocessed to remove the DC component from the vibration acceleration data and filtered to retain the AC component. The preprocessed vibration acceleration data is segmented according to a preset time window to generate multiple continuous vibration data segments, each of which has the same length. Windowing is applied to each vibration data segment. The vibration data segments are weighted based on the Hamming window function to reduce spectral leakage and generate windowed data segments. The windowed data segment is processed by Fast Fourier Transform to convert the vibration acceleration signal in the time domain into a complex spectrum in the frequency domain. The magnitude of the complex spectrum is calculated to generate the amplitude spectrum. Extract the frequency axis and amplitude value from the amplitude spectrum, divide the frequency axis into multiple frequency intervals, calculate the average amplitude value in each frequency interval, and generate vibration spectrum data. The horizontal axis of the vibration spectrum data is the frequency interval, and the vertical axis is the average amplitude value. The vibration spectrum data is smoothed by performing a sliding window averaging process on the amplitude values ​​of the vibration spectrum data based on a moving average filter to generate smoothed vibration spectrum data.

10. A computer system comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 9.