Adversarial Generation Data Augmentation Method and System for Highway Patrol and Maintenance Images

By constructing multimodal generation networks, the spatial and temporal dislocation of equipment vibration and visual images is solved, and high-quality highway patrol maintenance images are generated, which improves the accuracy of highway condition assessment and the reliability of maintenance decisions.

CN120031733BActive Publication Date: 2025-07-25JIANGSU ZHICHENG HUINING TRANSPORTATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510512537.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-25
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing technology cannot effectively simulate the spatial and temporal misalignment of equipment vibrations and lidar point clouds and visual images in highway patrol and maintenance images, resulting in poor image quality, affecting the accurate evaluation of highway conditions and the scientific nature of maintenance decisions.

Method used

By building a multimodal generation network with space-time alignment, integrating inertial navigation data, laser point clouds and visual images, collecting six-axis vibration signals and point cloud data of the maintenance equipment in real time, generating vibration frequency domain fingerprint maps and point cloud space-time encoding, and generating enhanced images using the adversarial generation network.

Benefits of technology

It improves the quality and authenticity of highway patrol and maintenance images, can more accurately identify road surface diseases and facility damage, provide a more reliable basis for highway maintenance decisions, and improves the efficiency and quality of maintenance work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031733B_ABST
    Figure CN120031733B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image enhancement, and discloses an adversarial generation data enhancement method and system for highway inspection and maintenance images. The method constructs a vibration frequency domain fingerprint map by collecting six-axis vibration signals of maintenance equipment in real time, collects point cloud data and visual images and processes them to construct a multi-modal fusion feature vector, and then constructs and trains an adversarial generation network to finally generate enhanced highway inspection and maintenance images. The present invention effectively solves the problem that traditional image enhancement only relies on visual modal data. Through multi-modal collaborative enhancement technology, it can more realistically simulate the spatio-temporal misalignment problems of equipment vibration, point cloud and visual images in maintenance operations, improve the quality and analysis accuracy of highway inspection and maintenance images, and provide a reliable basis for highway maintenance decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image enhancement, and more specifically, to an adversarial generation data enhancement method and system for highway inspection and maintenance images. Background Art

[0002] In highway inspection and maintenance work, the quality of image data plays a crucial role in accurately evaluating the highway condition. With the development of technology, image enhancement technology has been widely used in the field of highway inspection and maintenance, but traditional image enhancement methods have many limitations.

[0003] The Chinese patent with the authorization announcement number CN107153928B discloses a visual highway maintenance decision-making system, which improves the maintenance efficiency by collecting highway maintenance information and using index evaluation and visualization means to assist in maintenance decision-making. However, this system mainly focuses on the visualization of maintenance decisions and index analysis, and does not involve image enhancement technology, and cannot solve the problem of poor image quality caused by equipment vibration, spatio-temporal misalignment between lidar point cloud and visual image, etc. in highway inspection and maintenance images. In actual highway inspections, equipment vibration will cause the captured images to be blurred, and the spatio-temporal misalignment between lidar point cloud and visual image will lead to inaccurate information matching. These problems will seriously affect the judgment of highway diseases and facility conditions, and this patent does not provide solutions to these problems.

[0004] The patent application with the publication number CN118333608A discloses a highway inspection and maintenance system, which mainly focuses on the processing of maintenance status data, model establishment, and optimization of sprinkler operation to improve the sprinkler uniformity and utilization rate. However, this prior art also does not consider the impact of equipment vibration, spatio-temporal misalignment between lidar point cloud and visual image on the image quality in highway inspection and maintenance images. Due to the lack of effective processing of image data, it is difficult to make accurate judgments based on high-quality images in analyzing highway pavement conditions, road facility integrity, etc., which is not conducive to timely discovery and handling of highway diseases and facility damage problems.

[0005] The prior art cannot effectively simulate the problems of equipment vibration, spatio-temporal misalignment between lidar point cloud and visual image in the processing of highway inspection and maintenance images, and it is difficult to provide high-quality highway inspection and maintenance images, thus affecting the accurate evaluation of highway conditions and the scientific nature of maintenance decisions. Summary of the Invention

[0006] To overcome the above-mentioned defects of the prior art, the present invention provides an adversarial generation data enhancement method and system for highway inspection and maintenance images, aiming to innovatively integrate inertial navigation data, lidar point clouds, and visual images. By constructing a spatio-temporal alignment multi-modal generation network, the problems that traditional methods cannot simulate equipment vibration and the spatio-temporal misalignment between point clouds and visual images are effectively solved. Through multi-modal collaborative enhancement technology, a spatio-temporal alignment multi-modal generation network is constructed to improve the quality of highway inspection and maintenance images, providing a more reliable basis for highway maintenance decision-making.

[0007] The present invention is mainly applied to the work scenario of highway inspection and maintenance. During daily highway inspections, maintenance personnel use vehicles equipped with relevant equipment to conduct inspections on the highway. During the inspection process, the equipment will collect a large amount of highway image data. However, due to factors such as the vibration of the vehicle during driving and the differences in the acquisition time and space between lidar point clouds and visual images, the collected images often have problems such as blurring and inaccurate information. The method and system of the present invention can process these images, enhance the image quality, help maintenance personnel observe pavement diseases, damage to road facilities, etc. more clearly, and thus formulate maintenance plans more efficiently to ensure the safe and unobstructed operation of the highway.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] An adversarial generation data enhancement method for highway inspection and maintenance images, comprising:

[0010] Real-time collect the six-axis vibration signals of the maintenance equipment and construct a vibration frequency domain fingerprint map; perform frequency domain decomposition on the collected six-axis vibration signals based on a sliding time window, extract the characteristic frequency bands related to the movement speed of the maintenance equipment, and generate a vibration-speed mapping relationship matrix; dynamically correct the vibration-speed mapping relationship matrix and output a calibrated vibration frequency domain fingerprint map;

[0011] Collect the point cloud data of the highway inspection area, perform entropy value weighted encoding on the point cloud data to generate a point cloud spatio-temporal code; obtain the visual image of the highway inspection area, perform cross-modal correlation analysis on the six-axis vibration signals and the visual image to obtain a weighted visual feature map; construct a multi-modal fusion feature vector according to the calibrated vibration frequency domain fingerprint map, the point cloud spatio-temporal code, and the weighted visual feature map;

[0012] Construct and train an adversarial generation network including a generator and a discriminator according to the multi-modal fusion feature vector; obtain a new multi-modal fusion feature vector, and generate enhanced highway inspection and maintenance images according to the new multi-modal fusion feature vector and the trained adversarial generation network.

[0013] Further, the six-axis vibration signals collected by the real-time collection and maintenance equipment include: installing an inertial measurement unit on the highway inspection and maintenance equipment, and according to the installed inertial measurement unit, collecting the six-axis vibration signals of the maintenance equipment during driving in real time; the six-axis vibration signals include acceleration signals a x (t), a y (t), a z (t) in the three directions of the X, Y, and Z axes and angular velocity signals ω x (t), ω y (t), ω z (t).

[0014] Further, the construction of the vibration frequency domain fingerprint spectrum includes:

[0015] Performing frequency domain conversion on the collected six-axis vibration signals to obtain the main frequency band energy distribution and harmonic distortion rate; analyzing the mutual relationship between the acceleration signals in the three directions of the X, Y, and Z axes and the angular velocity signals in the three directions of the X, Y, and Z axes to obtain the inter-axis coupling coefficient; integrating the main frequency band energy distribution, harmonic distortion rate, and inter-axis coupling coefficient to construct the vibration frequency domain fingerprint spectrum.

[0016] Further, the obtaining of the inter-axis coupling coefficient includes:

[0017] Performing timestamp alignment on the acceleration signals a x (t), a y (t), a z (t) in the three directions of the X, Y, and Z axes and the angular velocity signals ω x (t), ω y (t), ω z (t);

[0018] According to the aligned a x (t), a y (t), a z (t), ω x (t), ω y (t), ω z (t), constructing a multi-signal space matrix;

[0019] According to the multi-signal space matrix, obtaining the correlation coefficient matrix R and the non-linear correlation matrix N;

[0020] Fusing the correlation coefficient matrix and the non-linear correlation matrix to obtain the comprehensive correlation matrix C, and extracting the inter-axis coupling coefficient from the comprehensive correlation matrix C.

[0021] Further, the construction of the multi-signal space matrix includes:

[0022] The aligned a x (t), a y (t), a z (t), ω x (t), ω y (t), ω z (t) are combined into a six-dimensional signal vector ;

[0023] Within a period of time collect six-dimensional signal vectors at multiple moments to construct a multi-signal space matrix , where is the six-dimensional signal vector at the nth moment.

[0024] Furthermore, the obtained correlation coefficient matrix R includes:

[0025] For any two components in the multi-signal space matrix and , calculate and the correlation coefficients between all different axes to obtain a correlation coefficient matrix , and the element in the correlation coefficient matrix represents the correlation between the signals of the '' axis and the '' axis; where is the six-dimensional signal vector at the ith moment, is the six-dimensional signal vector at the jth moment, 1 ≤ i ≤ n, 1 ≤ j ≤ n, .

[0026] Furthermore, the obtained non-linear correlation matrix N includes: calculating and the mutual information value , and constructing the non-linear correlation matrix N according to the mutual information value .

[0027] Furthermore, the generation of the vibration-velocity mapping relationship matrix includes:

[0028] Perform time-frequency representation on the six-axis vibration signals of each characteristic frequency band, and extract the amplitude statistics of each characteristic frequency band; obtain the motion speed of the maintenance equipment, associate the amplitude statistics of each characteristic frequency band with the motion speed of the maintenance equipment, and establish a vibration-velocity mapping relationship matrix M vs ; where the matrix element M vs (f i', v j' ) represents at the i'-th characteristic frequency band f i' below, the j'-th speed value v j' corresponding amplitude statistic.

[0029] Further, the output calibrated vibration frequency domain fingerprint map includes:

[0030] Dynamically correct the vibration-velocity mapping relationship matrix through a Kalman filter to obtain the corrected M vs matrix; according to the corrected M vs matrix, calibrate the vibration frequency domain fingerprint map to obtain the calibrated vibration frequency domain fingerprint map.

[0031] Further, the generation of point cloud spatio-temporal encoding includes:

[0032] Collect point cloud data in the highway inspection area, record the point cloud acquisition timestamp and spatial coordinates; based on the point cloud acquisition timestamp and spatial coordinates, calculate the spatio-temporal entropy weight coefficient of each point; extract the surface features of the point cloud data, and the surface features include reflection intensity gradient, normal vector offset, and local curvature mutation threshold; according to the spatio-temporal entropy weight coefficient of each point and the surface features of the point cloud data, perform entropy value weighted encoding on the point cloud data to generate point cloud spatio-temporal encoding.

[0033] Further, the obtaining of the weighted visual feature map includes:

[0034] Calculate the gradient direction of each pixel point of the visual image to form the visual image gradient direction; calculate the mutual information amount between each characteristic frequency band of the six-axis vibration signal and the visual image gradient direction to obtain the pixel-frequency band mutual information matrix I pf ; construct a non-linear mapping model w = F(m1, n1, I pf ) of the maintenance equipment rigidity coefficient m1, road surface roughness n1, pixel-frequency band mutual information matrix I pf and the pixel vibration sensitivity weight w; where, F is a non-linear function based on a deep neural network; determine the influence weight of m1, n1, I pf on w through the analytic hierarchy process, and substitute it into the model w = F(m1, n1, I pf ) to calculate the pixel vibration sensitivity weight of each pixel point; perform feature extraction on the visual image of the highway inspection area to obtain the visual feature map; multiply each pixel point's pixel vibration sensitivity weight with the visual feature map element by element to obtain the weighted visual feature map.

[0035] An adversarial generation data augmentation system for highway inspection and maintenance images, which is used to implement the above-mentioned adversarial generation data augmentation method for highway inspection and maintenance images. The system includes:

[0036] Spectrum construction module: It is used to collect the six-axis vibration signals of maintenance equipment in real time and construct vibration frequency-domain fingerprint spectra; perform frequency-domain decomposition on the collected six-axis vibration signals based on a sliding time window, extract characteristic frequency bands related to the movement speed of the maintenance equipment, and generate a vibration-speed mapping relationship matrix; dynamically correct the vibration-speed mapping relationship matrix and output the calibrated vibration frequency-domain fingerprint spectra;

[0037] Feature fusion module: It is used to collect point cloud data of the highway inspection area, perform entropy value weighted coding on the point cloud data to generate point cloud spatio-temporal coding; obtain the visual image of the highway inspection area, perform cross-modal correlation analysis on the six-axis vibration signal and the visual image to obtain a weighted visual feature map; construct a multi-modal fusion feature vector according to the calibrated vibration frequency-domain fingerprint spectra, point cloud spatio-temporal coding, and weighted visual feature map;

[0038] Adversarial generation module: According to the multi-modal fusion feature vector, construct and train an adversarial generation network including a generator and a discriminator; obtain a new multi-modal fusion feature vector, and generate an enhanced highway inspection and maintenance image according to the new multi-modal fusion feature vector and the trained adversarial generation network.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] The present invention innovatively fuses multi-modal data, effectively overcoming the limitations of traditional image enhancement methods. By collecting the six-axis vibration signals of maintenance equipment, point cloud data of the highway inspection area, and visual images in real time, constructing a multi-modal fusion feature vector, and then using an adversarial generation network to generate enhanced images. This not only solves the problems that traditional methods cannot simulate equipment vibration and spatio-temporal misalignment of lidar point clouds and visual images, but also can generate motion-blurred areas with physical realism, improving the authenticity and reliability of images. In the actual application of highway inspection and maintenance, the enhanced images help to more accurately identify road surface diseases, damage to road facilities, etc., provide more accurate and comprehensive basis for highway maintenance decision-making, improve the efficiency and quality of highway maintenance work, and ensure the safe and stable operation of highways. Brief description of the drawings

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0042] Figure 1 It is the principle flow chart of the adversarial generation data enhancement method for highway inspection and maintenance images in the present invention;

[0043] Figure 2 This is the method flowchart for constructing the vibration frequency-domain fingerprint spectrum in the adversarial generation data augmentation method for highway inspection and maintenance images of the present invention;

[0044] Figure 3 This is the schematic diagram of the principle of timestamp alignment of the present invention;

[0045] Figure 4 This is the method flowchart for generating the vibration-velocity mapping relationship matrix in the adversarial generation data augmentation method for highway inspection and maintenance images of the present invention;

[0046] Figure 5 This is the method flowchart for outputting the calibrated vibration frequency-domain fingerprint spectrum in the adversarial generation data augmentation method for highway inspection and maintenance images of the present invention;

[0047] Figure 6 This is the method flowchart for generating the point cloud spatio-temporal encoding in the adversarial generation data augmentation method for highway inspection and maintenance images of the present invention;

[0048] Figure 7 This is the method flowchart for constructing the multi-modal fusion feature vector in the adversarial generation data augmentation method for highway inspection and maintenance images of the present invention;

[0049] Figure 8 This is the functional module diagram of the adversarial generation data augmentation system for highway inspection and maintenance images of the present invention. Detailed implementation manners

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] Embodiment 1

[0052] Please refer to Figure 1 As shown, this embodiment provides an adversarial generation data augmentation method for highway inspection and maintenance images, including:

[0053] Step S1000, collect the six-axis vibration signals of the maintenance equipment in real time and construct the vibration frequency-domain fingerprint spectrum; perform frequency-domain decomposition on the collected six-axis vibration signals based on a sliding time window, extract the characteristic frequency bands related to the movement speed of the maintenance equipment, and generate the vibration-velocity mapping relationship matrix; dynamically correct the vibration-velocity mapping relationship matrix and output the calibrated vibration frequency-domain fingerprint spectrum;

[0054] Further, step S1000 includes:

[0055] Step S1100: Collect the six-axis vibration signals of the maintenance equipment in real time and construct a vibration frequency-domain fingerprint spectrum.

[0056] Furthermore, as Figure 2 shown, Step S1100 includes:

[0057] Step S1110: Install an inertial measurement unit on the highway inspection and maintenance equipment. According to the installed inertial measurement unit, collect the six-axis vibration signals of the maintenance equipment during driving in real time; the six-axis vibration signals include the acceleration signals a x (t), a y (t), a z (t) in the three directions of the X, Y, and Z axes and the angular velocity signals ω x (t), ω y (t), ω z (t) in the three directions of the X, Y, and Z axes.

[0058] Specifically, the purpose of Step S1110 is to obtain the six-axis vibration signals of the highway inspection and maintenance equipment during driving, providing a raw data basis for subsequent analysis of the equipment vibration state and generation of relevant data. An inertial measurement unit (IMU) is a combination of sensors that can measure the acceleration and angular velocity of an object, and its working principle is based on Newton's second law and the law of conservation of angular momentum. In this step, the inertial measurement unit is installed on the highway inspection and maintenance equipment. Through its internal acceleration sensors and angular velocity sensors, the acceleration signals a x (t), a y (t), a z (t) in the three directions of the X, Y, and Z axes and the angular velocity signals ω x (t), ω y (t), ω z (t) in the three directions of the X, Y, and Z axes of the equipment during driving are collected respectively. Usually, the driving direction of the equipment is used as an important reference to establish a coordinate system. Generally, the forward direction of the equipment is set as the positive direction of the X axis. In this way, during the driving of the equipment, the acceleration signal a x (t) in the X-axis direction mainly reflects information related to changes in the driving speed such as vehicle acceleration and deceleration. The direction perpendicular to the driving direction and in the horizontal direction of the equipment can be set as the Y axis, and the acceleration signal a y (t) in the Y-axis direction can reflect the lateral force and vibration of the vehicle during operations such as turning and avoidance. The Z axis is perpendicular to the plane determined by the X and Y axes, that is, perpendicular to the bottom surface of the equipment upward, and the acceleration signal a z (t) in the Z-axis direction can be used to monitor the vertical vibration of the vehicle caused by road bumps, uphill and downhill, etc.

[0059] Taking the inspection of a section of road with curves and straight sections as an example, when driving on a curve, the device will not only generate acceleration changes in the X and Y directions, but also may generate acceleration in the Z-axis direction due to vehicle tilt; when driving on a straight section, the acceleration change in the X-axis direction is more obvious, and the angular velocity signal will also change accordingly under different driving conditions. By collecting these signals in real time, the vibration conditions of the device under various driving conditions can be accurately recorded. The beneficial effect of this real-time collection method is that it can comprehensively and dynamically obtain the vibration information of the device. Because the driving environment of road inspection and maintenance equipment is complex and changeable, different road conditions, driving speeds and other factors will lead to different vibration states of the device. By continuously collecting six-axis vibration signals, these changes can be captured in time, providing rich data support for the subsequent accurate analysis of the vibration characteristics of the device. For example, if a section of the road surface is uneven, the device will generate frequent and complex vibrations during driving. Through the collection of six-axis vibration signals, the performance of this vibration in each axial direction can be accurately reflected, providing a basis for analyzing the operating state and potential faults of the device. From the perspective of data integrity, the six-axis vibration signal covers the acceleration and angular velocity information of the device in three spatial dimensions. Compared with only collecting single or partial axial signals, it can describe the vibration state of the device more comprehensively, avoid missing important information, and improve the accuracy and reliability of subsequent analysis results.

[0060] Step S1120: Perform frequency-domain conversion on the collected six-axis vibration signals to obtain the main frequency band energy distribution and harmonic distortion rate.

[0061] Specifically, frequency-domain conversion is the process of converting a time-domain signal (i.e., a vibration signal that changes with time) to the frequency domain. Its principle is based on the Fourier transform. Through the Fourier transform, a complex time series can be decomposed into a combination of sine and cosine waves of different frequencies, thereby revealing the energy distribution of the signal at different frequency components. In the specific calculation process, the fast Fourier transform (FFT) algorithm can be used. This is an efficient method for calculating the discrete Fourier transform (DFT), which can quickly obtain the frequency-domain representation of the vibration signal. After obtaining the frequency-domain signal, the calculation of the main frequency band energy distribution is to determine the frequency range where the vibration signal energy is mainly concentrated, that is, the main frequency band, and then calculate the proportion of the energy in this frequency band to the total energy. For example, if it is found through analysis that the vibration energy of the device is more concentrated in the 5-15 Hz frequency band, calculate the sum of the energies of all frequency components in this frequency band and divide it by the total energy of the entire frequency-domain signal to obtain the main frequency band energy distribution. The harmonic distortion rate reflects the distortion of the signal, and its calculation is based on the relationship between the fundamental wave and the harmonics. The fundamental wave is the basic frequency component of the signal, and the harmonics are the frequency components that are integer multiples of the fundamental wave frequency. When calculating the harmonic distortion rate, first determine the amplitude of the fundamental wave, then calculate the sum of the squares of the amplitudes of all harmonics, then divide the sum of the squares of the harmonic amplitudes by the square of the amplitude of the fundamental wave, and finally take the square root of the result and multiply it by 100% to obtain the harmonic distortion rate.

[0062] This step can understand the concentrated frequency range of the equipment vibration energy by obtaining the energy distribution of the main frequency band. For example, when it is found that the energy distribution of the equipment in a certain frequency band is abnormally increased, it may mean that there is a problem with the components or operating status of the equipment related to the frequency. If on a certain maintenance equipment, the main frequency band energy is mainly concentrated in 10-12Hz when the engine is operating normally, but the data collected at a certain time shows that the energy in the 15-18Hz frequency band has increased significantly, further inspection shows that a certain transmission component of the engine is worn, resulting in a change in the vibration energy distribution. The harmonic distortion rate can reflect the degree of signal distortion. If the harmonic distortion rate is too high, it means that there are more harmonic components in the signal, which may be caused by electrical faults inside the equipment, abnormal friction of mechanical parts, etc. By monitoring the harmonic distortion rate, potential fault hazards of the equipment can be discovered in time, providing a basis for preventive maintenance of the equipment. Moreover, these two features complement each other, can analyze the vibration characteristics of the equipment more comprehensively and deeply, help technicians more accurately judge the operating status of the equipment, improve the reliability and safety of the equipment, and reduce the interruption and loss of highway inspection and maintenance work caused by equipment failure.

[0063] Step S1130, analyzing the relationship between the acceleration signals in three directions and the angular velocity signals in three directions to obtain the inter-axis coupling coefficient;

[0064] The purpose of step S1130 is to deeply analyze the relationship between the acceleration signal and the angular velocity signal in the six-axis vibration signal of the highway inspection and maintenance equipment, obtain the inter-axis coupling coefficient, and provide key data support for constructing a vibration frequency domain fingerprint that can accurately reflect the vibration characteristics of the equipment. When the highway inspection and maintenance equipment is running, the acceleration signal a in the three directions of X, Y, and Z axes is x (t), a y (t), a z (t) and the angular velocity signals ω in the three directions of X, Y, and Z axes x (t),ω y (t),ω z(t) are mutually affected by factors such as the driving state of the device and the road surface conditions. The inter-axis coupling coefficient obtained by analyzing the relationships between these signals can reflect the degree of correlation between vibrations in different axis directions, which is of great significance for accurately judging the vibration state of the device. From the perspective of practical applications, if a certain highway inspection and maintenance device suddenly has an increase in the acceleration signal in the X-axis direction when passing through a rough road surface, and at this time the inter-axis coupling coefficient shows that the vibrations of the X-axis and Y-axis are closely related, then it is very likely that the vibration of the Y-axis will also be significantly affected. By obtaining the inter-axis coupling coefficient, the interaction between vibrations in different axes can be captured, providing a basis for comprehensively evaluating the vibration situation of the device in the follow-up. The beneficial effects of this step are reflected in many aspects. First of all, it helps to more precisely understand the internal mechanism of device vibration. The vibrations of different axes do not exist in isolation but are interrelated, and the inter-axis coupling coefficient can quantify this correlation, enabling technicians to have a deeper understanding of device vibration. Secondly, in terms of device fault diagnosis, changes in the inter-axis coupling coefficient can be used as an important fault indication signal. For example, when an abnormal fluctuation occurs in a certain inter-axis coupling coefficient, it may mean that there are faults in the relevant components of the device, such as loose connection components, bearing wear, etc., which helps to detect and solve potential problems in a timely manner, ensure the normal operation of the device, and improve the efficiency and reliability of highway inspection and maintenance work. Moreover, accurate inter-axis coupling coefficients provide more precise data for subsequent simulation of device motion blur, and the simulation images generated based on this can more realistically reflect the situation of the device during actual operation, thereby improving the accuracy of highway inspection and maintenance image analysis.

[0065] Further, step S1130 includes:

[0066] Step S1131, perform timestamp alignment on the acceleration signals a x (t), a y (t), a z (t) in the three directions of the X, Y, and Z axes and the angular velocity signals ω x (t), ω y (t), ω z (t) in the three directions of the X, Y, and Z axes;

[0067] Specifically, during the actual acquisition process, due to factors such as the sampling frequency of the sensors and transmission delays, these signals may have time deviations. If timestamp alignment is not performed on them, errors will occur in the subsequent analysis based on these signals, resulting in inaccurate results. The specific method of timestamp alignment is to record the exact time points of each signal sampling, and through a time synchronization algorithm, make all signals in the same reference system on the time axis. For example, as Figure 3 shown, before timestamp alignment, assume that the sampling time of the acceleration signal a x (t) is at time t1, and ω xThe sampling time of (t) is at the moment of t1 + Δt (Δt is the time deviation). Through the timestamp alignment operation, the time of ω x (t) is adjusted to the moment of t1, and the two are synchronized in time after the timestamp alignment.

[0068] The beneficial effects of this step are significant. First, it ensures the consistency and accuracy of the data. When analyzing the relationship between signals, only the data synchronized in time can truly reflect the internal connection between signals. If the signal times are not synchronized, the calculated parameters such as correlation will deviate, thus affecting the judgment of the vibration state of the equipment. Second, it lays a solid foundation for the subsequent construction of the multi-signal space matrix and the calculation of the correlation coefficient matrix and the non-linear correlation matrix. Only when the time is aligned can the calculation results of these matrices be reliable and accurately reflect the distribution of signals in time and dimension and the mutual relationship between signals. Third, in equipment condition monitoring and fault diagnosis, the time-aligned data can more accurately reflect the real-time changes during the operation of the equipment, which helps to detect the abnormal conditions of the equipment in time. For example, when a certain component of the equipment suddenly fails, the time-aligned signals can more clearly show the changes in the vibration signals of each axis at the moment of the fault occurrence, providing a more accurate basis for fault diagnosis.

[0069] Step S1132, construct a multi-signal space matrix according to the aligned a x (t), a y (t), a z (t), ω x (t), ω y (t), ω z (t);

[0070] Furthermore, step S1132 includes:

[0071] Step S11321, combine the aligned a x (t), a y (t), a z (t), ω x (t), ω y (t), ω z (t) into a six-dimensional signal vector ; represents the transpose of the vector;

[0072] Step S11322, within a period of time , collect six-dimensional signal vectors at different moments to construct a multi-signal space matrix , where is the six-dimensional signal vector at the nth moment.

[0073] Specifically, the core of step S1132 is to construct a multi-dimensional signal space matrix based on the time-stamp alignment, so as to comprehensively reflect the distribution of the six-axis vibration signals in terms of time and dimensions. When specifically implemented, first perform step S11321, and combine the aligned to form a six-dimensional signal vector . This step integrates vibration signals of different dimensions into one vector, facilitating subsequent processing. For example, at a certain moment t, the collected acceleration signals . This step is to integrate vibration signals of different dimensions into one vector for subsequent processing. For example, at a certain moment t, the collected acceleration signals , , and the angular velocity signal , then the six-dimensional signal vector at this moment is . Then perform step S11322. Within a period of time , collect six-dimensional signal vectors at different moments to construct a multi-dimensional signal space matrix , where is the six-dimensional signal vector at the th moment. For example, within 10 seconds (i.e., ), collect signals once every 0.1 second, and a total of 100 signals at different moments are collected (i.e., ). Arrange these 100 six-dimensional signal vectors in order to form a multi-dimensional signal space matrix .

[0074] The beneficial effects of this step are reflected in many aspects. From the perspective of data representation, the multi-dimensional signal space matrix organizes a large amount of vibration signal data in an orderly manner, intuitively showing the distribution of the signals in terms of time and dimensions. By observing the changes in the elements of the matrix, the changing trends of the vibration signals of each axis at different moments can be clearly seen. In terms of data analysis, this matrix form facilitates the subsequent calculation of the correlation coefficient matrix and the non-linear correlation matrix, and can efficiently analyze the relationships between signals by using matrix operation methods. For example, by operating on the matrix M, the correlations between signals of different axes can be quickly obtained, so as to discover potential vibration patterns. In the evaluation of equipment performance, the multi-dimensional signal space matrix can comprehensively reflect the vibration state of the equipment within a period of time, providing strong data support for evaluating the stability and reliability of the equipment. If there are abnormal fluctuations in a certain column of elements (representing the six-dimensional signal vector at a certain moment) in the matrix, it may mean that the equipment has abnormal conditions at this moment, and technicians can further analyze the reasons and take corresponding measures accordingly.

[0075] Step S1133: Obtain the correlation coefficient matrix R and the non-linear correlation matrix N according to the multi-dimensional signal space matrix;

[0076] Furthermore, step S1133 includes:

[0077] Step S11331: For the multi - variable signal space matrix for any two components and in it, calculate the correlation coefficients between all different axes of and to obtain a correlation coefficient matrix of . The element in the correlation coefficient matrix represents the correlation between the signals of the -th axis and the -th axis; where is the six - dimensional signal vector at the i - th moment, is the six - dimensional signal vector at the j - th moment, 1 ≤ i ≤ n, 1 ≤ j ≤ n, ;

[0078] Step S11332: Calculate the mutual information value and of , and construct the non - linear correlation matrix N according to the mutual information value .

[0079] Specifically, in step S11331, the Pearson correlation coefficient is used to measure the linear correlation between signals of different axes. In step S11332, mutual information is a concept in information theory, used to measure the degree of mutual dependence between two random variables. By calculating the mutual information values between all different axes, a 6×6 non - linear correlation matrix N is constructed, and the matrix elements represent the non - linear correlation between the corresponding axes. The beneficial effect of this step is very important. By obtaining the correlation coefficient matrix R and the non - linear correlation matrix N, the correlation between signals of different axes can be comprehensively analyzed from both linear and non - linear perspectives. In the actual operation of the device, the relationship between signals often contains both linear and non - linear components. The correlation coefficient matrix R can reveal the degree of linear correlation between signals. For example, if Close to 1 or -1 indicates a strong linear relationship between the signals of the i''-th axis and the j''-th axis; if close to 0, the linear relationship is weak. The non-linear correlation matrix N complements the non-linear information. For some complex equipment vibration situations, non-linear correlation may play a key role. For example, during the vibration of certain equipment, the vibrations of different axes may generate non-linear coupling through complex mechanical structures. In this case, relying solely on linear correlation analysis may not be able to accurately capture the relationship between signals, while the non-linear correlation matrix N can discover these hidden non-linear connections. In equipment fault diagnosis, comprehensive correlation analysis helps to more accurately determine the type and location of faults. For example, when the vibration of a certain axis is abnormal, by analyzing the correlation coefficient matrix R and the non-linear correlation matrix N, it is possible to determine which axes have a strong association with the abnormal axis, thereby narrowing the scope of fault investigation and improving the efficiency and accuracy of fault diagnosis.

[0080] Step S1134, fuse the correlation coefficient matrix and the non-linear correlation matrix to obtain a comprehensive correlation matrix C, and extract the coupling coefficients between axes from the comprehensive correlation matrix C.

[0081] Specifically, use the formula to fuse the correlation coefficient matrix R and the non-linear correlation matrix N, where α is the fusion coefficient, generally taking values between 0.5 - 0.8. The value of α is determined based on a large number of experiments and practical application experiences. When the value of α is close to 0.5, it means that the importance of the linear correlation matrix R and the non-linear correlation matrix N is relatively balanced; when α is close to 0.8, it relatively emphasizes the role of the correlation coefficient matrix R more. For example, in the vibration analysis scenario of a certain highway inspection and maintenance equipment, through multiple experiments, it is found that when α = 0.6, it can better comprehensively reflect the relationship between the vibrations of different axes of the equipment. The element C i''j'' in the fused comprehensive correlation matrix C is the coupling coefficient between any two axes i'' and j''. This coupling coefficient combines linear and non-linear correlation information and can more accurately describe the mutual influence of vibrations between axes than using only linear or non-linear correlation indicators.

[0082] The beneficial effects of this step are reflected at multiple levels. From the perspective of analyzing the vibration characteristics of the equipment, the inter-axis coupling coefficient synthesizes linear and non-linear related information, and can more comprehensively and accurately reflect the interaction mechanism between vibrations in different axis directions. For example, during the operation of the equipment, the vibrations of different axes may affect each other through complex physical processes, including both linear transmission relationships and non-linear resonance phenomena. The inter-axis coupling coefficient can quantify these complex relationships and help technicians deeply understand the essence of equipment vibration. In terms of equipment fault prediction, the change in the inter-axis coupling coefficient can be used as a sensitive indicator. When potential faults such as wear and looseness gradually occur in equipment components, the inter-axis coupling coefficient will change accordingly. By monitoring the change trend of the inter-axis coupling coefficient, abnormal conditions of the equipment can be detected in advance, providing a basis for preventive maintenance of the equipment, reducing the probability of equipment failures, and minimizing the interruption of highway inspection and maintenance work and economic losses caused by equipment failures. When simulating the motion blur of the equipment, an accurate inter-axis coupling coefficient can provide more reliable data support for generating a more realistic motion blur effect, thereby improving the accuracy and reliability of subsequent highway inspection and maintenance image analysis.

[0083] Step S1140: Integrate the main frequency band energy distribution, harmonic distortion rate, and inter-axis coupling coefficient to construct a vibration frequency domain fingerprint map.

[0084] Specifically, the main frequency band energy distribution reflects the concentration degree of the equipment vibration energy in different frequency bands. As mentioned above, it can help determine the main distribution area of the equipment vibration energy. The harmonic distortion rate reflects the distortion of the signal. The higher this value, the greater the degree of deviation of the signal from the ideal sine wave, which means that there may be abnormal interference factors during the operation of the equipment. The inter-axis coupling coefficient shows the mutual influence between vibrations in different axis directions. By calculating the correlation coefficient matrix R and the non-linear correlation matrix N and fusing them to obtain the comprehensive correlation matrix C, the element C extracted from C i''j'' is the coupling coefficient between axis i'' and axis j'', which quantifies the correlation degree between vibrations in different axes.

[0085] For example, assume that during the operation of a certain highway inspection and maintenance equipment, the vibration in the X-axis direction is relatively large. Through the inter-axis coupling coefficient, it is found that there is a strong coupling relationship between the X-axis and the Y-axis, which means that the vibration of the Y-axis may be significantly affected by the vibration of the X-axis. Integrating these three characteristics to construct a vibration frequency domain fingerprint map is like generating a unique "fingerprint" for the vibration state of the equipment. Because for different equipment in normal operation and fault states, the combination of these three characteristics will present different patterns, just like everyone's fingerprint is unique.

[0086] Its beneficial effects are reflected in multiple aspects. From the perspective of equipment status monitoring, the vibration frequency-domain fingerprint spectrum can establish an accurate vibration characteristic model for the equipment. When a fault occurs or the performance of the equipment deteriorates, its vibration frequency-domain fingerprint spectrum will change accordingly. By comparing the current spectrum with the spectrum in the normal state, technicians can quickly and accurately determine whether there are abnormalities in the equipment, as well as the possible locations and causes of the abnormalities. For example, if it is found that the harmonic distortion rate of a certain equipment suddenly increases, and at the same time, the inter-axis coupling coefficient changes significantly between some axis pairs, combined with the change in the main frequency band energy distribution, it can be inferred that there may be a problem with a component of the equipment that involves multi-axis movement and is prone to harmonic interference. In terms of simulating equipment motion blur, the vibration frequency-domain fingerprint spectrum provides key input information for the subsequent steps. Since it accurately reflects the vibration state of the equipment, the motion blur simulation generated based on this can be more in line with the blur effect generated during the actual operation of the equipment, thereby laying a foundation for generating high-quality highway inspection and maintenance simulation images, improving the accuracy of subsequent image analysis, and helping to more accurately evaluate the highway inspection and maintenance status.

[0087] Step S1200: Based on a sliding time window, perform frequency-domain decomposition on the collected six-axis vibration signals, extract the characteristic frequency bands related to the movement speed of the maintenance equipment, and generate a vibration-speed mapping relationship matrix;

[0088] Furthermore, as Figure 4 shown, step S1200 includes:

[0089] Step S1210: Based on a sliding time window, perform frequency-domain decomposition on the collected six-axis vibration signals, and extract the characteristic frequency bands related to the movement speed of the maintenance equipment;

[0090] Specifically, the sliding time window is a commonly used method in time series data processing. By sliding a window of a fixed length on the time axis, the data within the window is analyzed. Selecting an appropriate sliding time window length and sliding step size Δt' is crucial. For example, setting T windowLet \( \Delta t \) be 1 second and \( \Delta t' \) be 0.1 second. In this way, while ensuring data continuity, the vibration signals in different time periods can be analyzed in detail. When processing the six-axis vibration signals within each sliding time window, the short-time Fourier transform (STFT) in Fourier transform is used to convert the time-domain signal into a frequency-domain signal. The principle of STFT is to slide a window function over the signal and perform Fourier transform on the signal within the window, so as to obtain the distribution of the frequency components of the signal at different local times. The reason for choosing STFT is that it can display the characteristics of the signal simultaneously in both the time and frequency dimensions and is suitable for analyzing non-stationary signals, while the vibration signals of the maintenance equipment are non-stationary during driving. The frequency band of 0.5 - 20 Hz is selected as the characteristic frequency band related to the equipment movement speed, which is determined based on a large number of experiments and practical experiences. When the equipment is moving, there is a close connection between the signal components within this frequency band and the speed change. For example, when the equipment accelerates, the signal intensity of certain frequencies within this frequency band may increase; when decelerating, it will decrease. By extracting the signal components within this characteristic frequency band from the frequency-domain signal of each time window, the vibration characteristics related to the speed can be accurately obtained.

[0091] The beneficial effects of this step are significant. First of all, the method of combining the sliding time window with STFT can dynamically track the change of the frequency components of the vibration signal over time. For example, when the equipment passes through a rough road surface, the vibration signal will change suddenly, and this method can capture these changes in time and accurately extract the frequency characteristics corresponding to the current motion state, avoiding ignoring local changes due to overall signal analysis. Secondly, accurately extracting the characteristic frequency band related to the speed improves the accuracy of subsequent analysis of the relationship between the equipment vibration and the speed. By analyzing the signals in these characteristic frequency bands, the vibration law of the equipment at different speeds can be understood more deeply, providing a reliable data basis for establishing an accurate vibration-speed mapping relationship matrix. Moreover, in terms of equipment fault diagnosis, the changes in these characteristic frequency bands can be used as important diagnostic bases. If the signal in a certain characteristic frequency band appears abnormal, it may mean that there is a problem with a certain speed-related component of the equipment, which helps to quickly locate the fault source and improve the reliability and safety of the equipment.

[0092] Step S1220: Perform time-frequency representation on the six-axis vibration signals of each characteristic frequency band and extract the amplitude statistics of each characteristic frequency band;

[0093] Specifically, the amplitude statistics include the amplitude mean and the peak value. The time-frequency representation uses the short-time Fourier transform (STFT) because STFT can simultaneously display the characteristics of the signal in both the time and frequency dimensions and is suitable for analyzing non-stationary vibration signals. After performing STFT processing on the six-axis vibration signals in each characteristic frequency band, a time-frequency diagram is obtained. From the time-frequency diagram, the energy distribution of the signal at different times and frequencies can be clearly observed. When calculating the amplitude mean, first determine the time interval and frequency range corresponding to each characteristic frequency band in the time-frequency diagram, then accumulate the amplitude values of all sampling points within this range, and divide by the total number of sampling points to obtain the amplitude mean. When calculating the amplitude peak value, in the time-frequency diagram corresponding to each characteristic frequency band, find the maximum value among all the amplitude values of the sampling points, and this value is the amplitude peak value.

[0094] The beneficial effects of this step are reflected in multiple aspects. In equipment condition monitoring, the amplitude mean and peak value can intuitively reflect the intensity characteristics of the vibration signal. The amplitude mean can reflect the average intensity of the equipment's vibration over a period of time. If the amplitude mean exceeds the normal range, it may mean that the overall vibration of the equipment has increased and there are potential problems. The amplitude peak value, on the other hand, can reflect the instantaneous maximum intensity in the vibration signal. When the amplitude peak value is too high, it may indicate that the equipment has received a large impact at a certain moment, and it is necessary to pay attention to whether the relevant components are damaged. In equipment fault diagnosis, the changes in the amplitude statistics can be used as an important basis for fault diagnosis. For example, when a certain component of the equipment wears or loosens, the amplitude mean and peak value of its vibration signal may change significantly. By monitoring these changes, faults can be detected in a timely manner and corresponding measures can be taken. In addition, when simulating equipment motion blur, the amplitude statistics help to more accurately simulate the impact of equipment vibration on the image. Different amplitude statistics correspond to different degrees of vibration blur effects, providing more accurate parameter support for generating motion blur images that conform to the actual situation and improving the authenticity and reliability of image simulation.

[0095] Step S1230: Obtain the motion speed of the maintenance equipment, associate the amplitude statistics of each characteristic frequency band with the motion speed of the maintenance equipment, and establish a vibration-speed mapping relationship matrix M vs ; where the matrix element M vs (f i' , v j' ) represents the amplitude statistics corresponding to the i'-th characteristic frequency band f i' and the j'-th speed value v j' .

[0096] Specifically, the movement speed of the maintenance equipment is obtained through devices such as GPS. GPS (Global Positioning System) is a positioning technology based on satellite navigation. By receiving signals transmitted by multiple satellites, it calculates the position information of the equipment and calculates the movement speed based on the position changes at different times. Establish the vibration - speed mapping relationship matrix M vs When vs (f i' ,v j' ) represents the amplitude statistic corresponding to the i'-th characteristic frequency band f i' and the j'-th speed value v j' . For example, assume the characteristic frequency band is , and the speed value . Through the previous steps, the amplitude mean at this characteristic frequency band and speed is , and the amplitude peak is . Then in the matrix , M vs (f1,v1) can be represented by a data structure containing and , such as .

[0097] The beneficial effects of this step are very significant. From the perspective of equipment performance analysis, the vibration - speed mapping relationship matrix M vs clearly shows the vibration conditions of each characteristic frequency band of the equipment at different speeds. By analyzing the data in the matrix, the vibration characteristics of the equipment at different operating speeds can be understood, and the stability and reliability of the equipment can be evaluated. For example, if it is found that within a certain speed range, the amplitude statistic of a specific characteristic frequency band increases abnormally, it indicates that the equipment may have vibration problems at this speed and relevant components need to be further inspected. In terms of equipment fault prediction, this matrix provides a strong basis for predicting equipment faults. As the equipment operates, by continuously monitoring the changes in the vibration - speed mapping relationship, potential fault signs can be detected in a timely manner. For example, when some elements in the matrix show abnormal change trends, it may indicate that the equipment is about to fail, so maintenance can be arranged in advance to avoid work interruptions and losses caused by equipment failures. In addition, in the simulation and analysis of highway inspection and maintenance images, the vibration - speed mapping relationship matrix provides key data support for generating more realistic motion - blurred images. Based on this matrix, according to the actual movement speed and vibration conditions of the equipment, the blur effect of the image can be accurately simulated, improving the accuracy of image analysis and helping to more precisely evaluate the highway inspection and maintenance status.

[0098] Step S1200 aims to process the collected six-axis vibration signals, extract the characteristic frequency bands related to the movement speed of the maintenance equipment, and establish a vibration-speed mapping relationship matrix, so as to provide key data support for the subsequent analysis of the relationship between equipment vibration and speed, and also lay the foundation for generating a calibrated vibration frequency domain fingerprint spectrum, so that the spectrum can more accurately reflect the vibration characteristics of the equipment at different speeds. During the driving process of highway inspection and maintenance equipment, its vibration situation is complex and closely related to the movement speed. Through this series of operations, the intrinsic relationship between vibration signals and speed can be deeply explored, thereby providing a more accurate basis for the processing and analysis of highway inspection and maintenance images. Under different road conditions, such as highways, rural roads, etc., the driving speed of maintenance equipment is different, and the vibration signals generated will also be different. Through step S1200, these differences can be effectively captured, providing strong support for the subsequent accurate evaluation of the equipment operation status and the generation of high-quality simulation images. The beneficial effects of this step are reflected in many aspects. First, in the field of equipment status monitoring, the vibration-speed mapping relationship matrix can intuitively display the vibration characteristics of the equipment at different speeds. For example, when the amplitude of the vibration signal in a certain speed range is found to be abnormally increased, combined with the mapping relationship matrix, it can be judged that the equipment may have a potential failure risk at this speed, which is helpful for timely equipment maintenance and ensuring the smooth progress of highway inspection and maintenance work. Second, when simulating equipment motion blur, the accurate vibration-speed mapping relationship can provide key data for generating motion blur images that conform to the actual situation. The images generated based on these data are closer to the images taken in the real scene, which is conducive to the subsequent analysis and processing of highway inspection and maintenance images, and improves the accuracy and reliability of image analysis. Third, for optimizing highway inspection and maintenance strategies, by analyzing the vibration-speed mapping relationship, the vibration of the equipment at different speeds can be understood, and then the driving speed of the equipment can be reasonably adjusted to reduce equipment wear and extend the service life of the equipment, while improving the efficiency and quality of highway inspection and maintenance.

[0099] Step S1300, dynamically correct the vibration-velocity mapping relationship matrix through a Kalman filter, and output a calibrated vibration frequency domain fingerprint.

[0100] Furthermore, if Figure 5 As shown, step S1300 includes:

[0101] Step S1310, dynamically modify the vibration-velocity mapping relationship matrix through the Kalman filter to obtain the modified M vs matrix;

[0102] Step S1320, based on the modified M vs The vibration frequency domain fingerprint spectrum is calibrated by the matrix to obtain the calibrated vibration frequency domain fingerprint spectrum.

[0103] Specifically, the core objective of step S1300 is to dynamically correct the vibration-velocity mapping relationship matrix through a Kalman filter and calibrate the vibration frequency-domain fingerprint spectrum based on the corrected matrix, so that the calibrated spectrum can more accurately reflect the actual vibration state of the device, providing a reliable basis for generating a motion-blurred image that conforms to the actual situation in the subsequent process. During the highway inspection and maintenance process, the operating environment of the device is complex and variable, and factors such as road surface conditions and vehicle loads can cause changes in the vibration signals and motion speeds of the device, making the previously established vibration-velocity mapping relationship matrix may not accurately reflect the actual situation. Therefore, it is necessary to dynamically correct it to ensure the accuracy and reliability of subsequent analysis and applications based on this matrix.

[0104] The main task of step S1310 is to dynamically correct the vibration-velocity mapping relationship matrix using a Kalman filter to obtain the corrected matrix M vs matrix. The Kalman filter is an algorithm that uses the state equation of a linear system and, through system input and output observation data, optimally estimates the system state. Its core principle is based on two steps: prediction and update.

[0105] In the prediction stage, the Kalman filter predicts the state at the current moment based on the state estimate value at the previous moment and the state transition equation of the system. For the vibration-velocity mapping relationship matrix, assuming the matrix at the previous moment is , according to the physical model and relevant parameters of the device's motion, the state transition equation can be established, where is the state transition matrix, which describes the variation law of the matrix over time. For example, if the device's motion state is relatively stable, the elements of the matrix may represent the linear variation relationship of each element in the matrix over time; if the device's motion state is complex and variable, the elements of the matrix need to be determined according to a more complex physical model.

[0106] In the update stage, the Kalman filter combines the latest vibration signal and speed measurement data to correct the prediction result. First, calculate the Kalman gain , and its calculation formula is , where is the prediction error covariance matrix, which reflects the uncertainty of the predicted value; is the observation matrix, which is used to relate the system state to the observation data; is the observation noise covariance matrix, represents the transpose of the observation matrix. Then, update the predicted matrix according to the Kalman gain to obtain the corrected matrix , where is the observed data at the current moment, that is, the part related to the matrix after the latest vibration signal and speed measurement data are processed.

[0107] In the correction process, it is of great significance to set the adaptive noise covariance matrix. According to the changes in the equipment operating environment, such as the change in road surface flatness, vehicle load change, etc., the parameters of the noise covariance matrix are dynamically adjusted. When the road surface becomes rough, the noise of the equipment vibration signal increases and the uncertainty increases. At this time, the element values of the noise covariance matrix are appropriately increased, so that the Kalman filter pays more attention to the new observed data during the update process, thus better adapting to the uncertainty in the signal; when the road surface is relatively flat, the signal noise is relatively small, and the element values of the noise covariance matrix are reduced to improve the accuracy of the filter, making the corrected matrix closer to the real situation. For example, when driving on a rough road surface, the elements related to the vibration signal in the noise covariance matrix increase from the initial value Q1 to Q2 (Q2 > Q1). In this way, when calculating the Kalman gain, the weight of the new observed data increases, and the corrected M vs matrix can more accurately reflect the vibration - speed relationship of the equipment under this road condition.

[0108] The beneficial effects of this step are remarkable. Through the dynamic correction of the Kalman filter, the accuracy and adaptability of the vibration - speed mapping relationship matrix can be effectively improved. An accurate matrix can more precisely describe the relationship between vibration and speed of the equipment under different operating states, providing a more reliable data basis for subsequent analysis. In equipment fault diagnosis, an accurate matrix helps to more accurately judge the correlation between abnormal equipment vibration and speed change, thus more precisely locating the cause of the fault. For example, if the equipment vibrates abnormally at a certain speed, based on the corrected matrix, it can be more accurately analyzed whether it is the normal vibration fluctuation caused by the speed change or the abnormal vibration caused by a fault in the equipment itself. In addition, in the generation of simulated equipment motion blurred images, an accurate vibration - speed mapping relationship matrix can provide more realistic parameters for the simulation process, making the generated motion blurred images more truly reflect the imaging situation of the equipment under different speeds and vibration states, improving the quality and reliability of image simulation.

[0109] The goal of step S1320 is to calibrate the vibration frequency - domain fingerprint spectrum based on the corrected M vs matrix obtained in step S1310, and then obtain the calibrated vibration frequency - domain fingerprint spectrum. The vibration frequency - domain fingerprint spectrum consists of features such as the main frequency band energy distribution, harmonic distortion rate, and inter - axis coupling coefficient, etc. These features comprehensively reflect the vibration state of the equipment. And the corrected M vs matrix contains more accurate information about the relationship between equipment vibration and speed. Using this information to calibrate the fingerprint spectrum can optimize the characterization of the equipment vibration state by the spectrum.

[0110] When performing the calibration operation, each feature of the vibration frequency-domain fingerprint spectrum needs to be adjusted separately. For the main frequency band energy distribution, the element M in the M vs matrix vs (f i' , v j' ) represents the amplitude statistic corresponding to the i'-th characteristic frequency band f i′ and the j'-th speed value v j′ . By analyzing the changes in the amplitude statistics of each characteristic frequency band at different speeds, the actual distribution of the device vibration energy in different frequency bands can be inferred. For example, if at a certain speed v2, the M vs matrix shows that the average amplitude in the 0.5 - 10 Hz frequency band increases significantly, which means that the vibration energy in this frequency band increases. Then, in the main frequency band energy distribution, the energy proportion of this frequency band should be increased accordingly. When calculating specifically, assume that the original energy proportion of this frequency band is E1. By comparing and analyzing the amplitude statistics of this frequency band and speed in the M vs matrix with those of other frequency bands and speeds, according to the relationship between energy and amplitude (generally, energy is proportional to the square of the amplitude), recalculate the energy proportion of this frequency band as E2 (E2 > E1), thereby completing the calibration of the main frequency band energy distribution.

[0111] For the harmonic distortion rate, the change in the device vibration - speed relationship may affect the harmonic components of the signal. The changes in the vibration conditions of each characteristic frequency band at different speeds in the matrix can be used as a basis for judging the change in the harmonic distortion rate. For example, when the device speed changes, if the matrix shows that the vibration modes of some characteristic frequency bands change, it may lead to a change in the harmonic content. By analyzing the matrix, combined with the calculation method of the harmonic distortion rate (the harmonic distortion rate is the ratio of the harmonic content to the fundamental wave content. When calculating, first determine the fundamental wave amplitude, then calculate the sum of the squares of all harmonic amplitudes, then divide the sum of the squares of the harmonic amplitudes by the square of the fundamental wave amplitude, and finally take the square root and multiply by ), recalculate the harmonic distortion rate. Assume that the original harmonic distortion rate is . After analyzing the matrix and recalculating, a new harmonic distortion rate is obtained, thereby completing the calibration of the harmonic distortion rate.

[0112] Regarding the inter-axis coupling coefficient, the changes in the device vibration and speed may change the vibration correlation between different axes. The matrix reflects the changes in the amplitude statistics of the vibrations of each axis at different speeds, and these changes are related to the inter-axis coupling coefficient. For example, if the matrix shows that at a certain speed, axis and The change in the vibration amplitude of the shaft shows a specific correlation, which may mean that the inter-axis coupling coefficient needs to be adjusted. According to the information in the matrix, combined with the previous method of calculating the inter-axis coupling coefficient (by calculating the correlation coefficient matrix and the non-linear correlation matrix and fusing them to obtain the comprehensive correlation matrix , extract the inter-axis coupling coefficient from ), recalculate the inter-axis coupling coefficient. Assume that the original coupling coefficient between the X-axis and the Y-axis is , and after recalculation, we get , completing the calibration of the inter-axis coupling coefficient.

[0113] Through the above calibration of the main frequency band energy distribution, harmonic distortion rate, and inter-axis coupling coefficient, the calibration of the vibration frequency domain fingerprint spectrum is achieved, and the calibrated vibration frequency domain fingerprint spectrum is obtained. This step has various beneficial effects. In the field of equipment condition monitoring and fault diagnosis, the calibrated vibration frequency domain fingerprint spectrum can provide more accurate equipment vibration information, helping technicians to detect potential problems of the equipment in a timely manner. For example, by comparing the spectra before and after calibration, if it is found that the calibration result of a certain feature is significantly different from the normal range, such as the abnormal increase in the main frequency band energy distribution in a certain frequency band, combined with the equipment operation speed information, it can be inferred that the corresponding components of the equipment at this speed may have wear or potential faults, so as to arrange inspections and maintenance in a timely manner, avoid the occurrence of equipment failures, reduce the equipment maintenance cost, and improve the reliability and operation efficiency of the equipment. In the generation of simulated equipment motion blurred images, the calibrated vibration frequency domain fingerprint spectrum, as the benchmark input for the generator motion blur simulation, can significantly improve the quality of the simulated images. Since the calibrated spectrum more accurately reflects the actual vibration state of the equipment, the generated motion blurred images can more realistically simulate the imaging situation of the equipment during highway inspection and maintenance. For example, when analyzing the highway pavement condition, more realistic simulated images can help technicians observe the subtle cracks, potholes and other diseases on the pavement more clearly, improve the efficiency and accuracy of highway inspection and maintenance work, and provide a more reliable basis for highway maintenance decision-making.

[0114] Step S2000: Collect the point cloud data of the highway inspection area, perform entropy value weighted coding on the point cloud data to generate point cloud spatio-temporal coding; obtain the visual image of the highway inspection area, perform cross-modal correlation analysis on the six-axis vibration signal and the visual image to obtain the weighted visual feature map; construct a multi-modal fusion feature vector according to the calibrated vibration frequency domain fingerprint spectrum, point cloud spatio-temporal coding, and weighted visual feature map;

[0115] Furthermore, step S2000 includes:

[0116] Step S2100, collect the point cloud data of the highway patrol area, perform entropy value weighted encoding on the point cloud data, and generate point cloud spatio-temporal encoding;

[0117] Further, as Figure 6 shown, step S2100 includes:

[0118] Step S2110, collect the point cloud data of the highway patrol area, and record the point cloud acquisition timestamp and spatial coordinates;

[0119] Specifically, the point cloud data is obtained by a laser scanning device. The laser scanning device measures the distance information of the points on the object surface by using the time difference between the laser beam emission and reception, so as to obtain the three-dimensional spatial coordinates of the object. During the acquisition process, the scanning parameters are dynamically adjusted according to the complexity of the scene, in order to improve the acquisition efficiency while ensuring the data quality. For example, in curved roads or areas with complex structures, such as road intersections and near bridge structures, the scanning resolution is increased. Because the object shapes and spatial layouts in these areas are relatively complex, a higher resolution can obtain more detailed point cloud data to accurately describe the shape and position information of the objects. Suppose at a road bend, high-resolution scanning can clearly obtain the detailed point cloud data of the guardrail, roadside signs, and road surface texture at the bend. These data are of great significance for subsequent analysis of the safety conditions of the bend and the integrity of road facilities. In the open straight road area, the resolution is appropriately reduced to improve the scanning efficiency. The scene in the open straight road area is relatively simple. Reducing the resolution will not have a great impact on the acquisition of key information, and at the same time, it can reduce the data acquisition volume and processing time. For example, on a long straight highway, after reducing the scanning resolution, the laser scanning device can complete the scanning of this area faster, improving the overall patrol efficiency without losing key road information.

[0120] Recording the point cloud acquisition timestamp and spatial coordinates is for the needs of subsequent processing and analysis. The timestamp can record the specific moment when each point cloud data is acquired, which is very important for analyzing the changes of highway facilities over time. For example, by comparing the data of the same point cloud area acquired at different times, the wear and deformation of road facilities can be monitored. The spatial coordinates clarify the position of each point in the three-dimensional space, providing basic data for constructing an accurate three-dimensional model and performing spatial analysis. For example, when constructing a three-dimensional model of a highway, accurate spatial coordinates can ensure the accuracy of the model, enabling technicians to intuitively observe the terrain, slope of the highway, and the spatial layout of road facilities. The beneficial effect of this step is that it provides accurate and complete basic data for subsequent point cloud data processing, analysis, and fusion with other modality data, helps improve the quality and efficiency of highway patrol and maintenance work, and provides strong data support for the maintenance and management of highway facilities.

[0121] Step S2120, calculating the spatiotemporal entropy weight coefficient of each point based on the point cloud acquisition timestamp and spatial coordinates;

[0122] Specifically, entropy is a concept in information theory that measures the uncertainty or confusion of information. The calculation of the spatiotemporal entropy weight coefficient combines the temporal and spatial characteristics of point cloud data, aiming to highlight the information contribution of key points and provide a basis for subsequent entropy value weighted encoding. First, for each point cloud data point, its entropy value in the time dimension and space dimension is calculated according to its timestamp and spatial coordinates. Assume that in the time period Collected within 2 point cloud data points, each point cloud data point has a timestamp of First, divide the time axis into several equally spaced time intervals. , count the frequency of point cloud data points in each time interval is the number of time intervals). Here the frequency The calculation method is, The number of point cloud data points in the time interval divided by the total number of points 2. According to the calculation formula of entropy in information theory , for the entropy value of point cloud data in the time dimension , substitute the frequency of each time interval into the formula for calculation. For example, if In the time interval, the frequencies of point cloud data points appearing are , then the entropy value in the time dimension is for:

[0123]

[0124] For the calculation of entropy value in spatial dimension, it is necessary to consider the distribution of point cloud data points in space. Assume that the spatial coordinates of point cloud data points are , is the total number of point cloud data points. First, calculate the spatial distance between each point cloud data point and all other point cloud data points ; Next, determine a spatial neighborhood radius , count each point cloud data point with a radius of The number of points in the neighborhood of . Calculate the density of points in the neighborhood ,in is the volume of the neighborhood; according to the entropy calculation formula, the entropy value in the spatial dimension is .

[0125] The beneficial effects of this step are reflected in multiple aspects. In terms of data processing, by calculating the spatio-temporal entropy weight coefficient, key points carrying important information can be highlighted, and redundant information can be suppressed. For example, in a large amount of point cloud data, some key points may represent the critical parts or abnormal conditions of road facilities. By assigning higher weights to these points, more attention can be paid to these points in subsequent processing, improving the efficiency and accuracy of data processing. In terms of highway condition analysis, highlighting key points helps to more accurately identify problems such as road diseases and facility damages. For example, for a small pothole on the road surface, the corresponding point cloud data may account for a small proportion in the overall data. However, through the calculation of the spatio-temporal entropy weight coefficient, the information of these points is highlighted, making it easier for technicians to discover and analyze the problem. In addition, in multi-modal data fusion, accurate spatio-temporal entropy weight coefficients can better fuse point cloud data with other modal data (such as visual images and vibration signals), improving the fusion effect and providing more reliable support for subsequent highway inspection and maintenance image analysis and decision-making.

[0126] Step S2130, extract the surface features of the point cloud data, where the surface features include the reflection intensity gradient, the normal vector offset, and the local curvature mutation threshold;

[0127] Specifically, the purpose of step S2130 is to extract the surface features of the point cloud data. These features include the reflection intensity gradient, the normal vector offset, and the local curvature mutation threshold. These features can further describe the surface characteristics of the point cloud data, providing a basis for generating more targeted spatio-temporal encoding of the point cloud and helping to more accurately reflect the actual situation of the highway inspection area.

[0128] The reflection intensity gradient is used to measure the degree of change in the reflection intensity of each point in the point cloud data. When calculating, for each point in the point cloud data, a suitable neighborhood (such as a spherical neighborhood with a radius of 1) is selected centered on this point. Within this neighborhood, the gradient of the reflection intensity is calculated. Assuming that the reflection intensity of point is , and the reflection intensity of other points in the neighborhood is , then the calculation of the reflection intensity gradient can be achieved by taking the derivative of the change in the reflection intensity in space, generally using the method of numerical difference. For example, in the Cartesian coordinate system, calculate the reflection intensity gradient component of point in the direction. Through points in the neighborhood and points offset by in the direction (the coordinates can be expressed as , where The reflection intensity (etc., the coordinate values are determined based on the coordinate information of the points in the neighborhood)) and ; is the reflection intensity of the point at a distance of Δx from the point in the x - direction; is the reflection intensity of the point at a distance of - Δx from the point in the x - direction; The numerical difference method is used to calculate and in the same way in other directions, and finally the complete reflection intensity gradient is obtained, where is the reflection intensity gradient component of the point in the direction, is the reflection intensity gradient component of the point in the direction. In the area with a large reflection intensity gradient, it indicates that the reflection intensity changes violently, which may correspond to the boundary of the object surface, material changes, or other important features.

[0129] The normal vector offset is used to describe the difference between the normal vector of the point cloud data point and the normal vector of the reference plane. First, the normal vector of the point cloud data point needs to be determined. For each point, a local plane is formed by fitting the points in its neighborhood, and the normal vector of this plane is the normal vector of this point. Then, a reference plane (such as a horizontal plane or a certain specific reference plane) is selected, and the angle between the normal vector of the point and the normal vector of the reference plane is calculated. The normal vector offset can be calculated by , where is the normal vector of the point , and is the normal vector of the reference plane. The normal vector offset can reflect the inclination degree and direction change of the surface where the point cloud data point is located, and helps to identify the concavity, convexity, slope change, etc. of the object surface.

[0130] The local curvature mutation threshold is used to detect the area where the local curvature in the point cloud data changes significantly. The local curvature reflects the degree of bending of the point cloud data surface. When calculating the local curvature, a neighborhood is also selected with the point as the center, a quadratic surface is formed by fitting the points in the neighborhood, and the local curvature is calculated according to the parameters of the quadratic surface. Suppose the local curvature of a certain point is , in its neighborhood, a threshold is set. When ( When the local curvature of a point is significantly different from that of other points in the neighborhood (i.e., the local curvature of other points in the neighborhood), it is considered that the point is in the local curvature mutation region. The determination of the local curvature mutation threshold needs to be adjusted according to the characteristics of the actual point cloud data and application requirements, and generally, appropriate values are determined through multiple experiments and analyses.

[0131] The extraction of these surface features has various beneficial effects. In highway inspections, the reflection intensity gradient can help identify objects with obvious reflection intensity changes, such as road signs and traffic facilities, because the reflection intensity gradients of these objects are usually large. For example, the white markings on the road have different reflection intensities from the surrounding road surface, and they can be clearly distinguished through the reflection intensity gradient. The normal vector offset helps analyze the slope and undulation of the road surface, which is of great significance for judging road safety and vehicle driving stability. For instance, in the inspection of mountain roads, sections with large slopes can be detected through the normal vector offset, and corresponding safety measures can be taken in advance. The local curvature mutation threshold can detect diseases such as potholes and bumps on the road surface because these diseases will cause local curvature mutations. By accurately extracting these surface features, the characteristics of the point cloud data can be described more comprehensively and meticulously, providing key information for subsequent generation of high-quality point cloud spatio-temporal coding, thereby improving the efficiency and accuracy of highway inspection and maintenance work.

[0132] Step S2140: According to the spatio-temporal entropy weight coefficient of each point and the surface features of the point cloud data, perform entropy value weighted coding on the point cloud data to generate point cloud spatio-temporal coding.

[0133] Specifically, the core task of step S2140 is to perform entropy value weighted coding on the point cloud data according to the spatio-temporal entropy weight coefficient of each point and the surface features of the point cloud data to generate point cloud spatio-temporal coding, aiming to highlight the information contribution degree of key points, suppress redundant information, and provide a more valuable point cloud data representation form for subsequent multi-modal fusion and highway inspection and maintenance image analysis.

[0134] When performing entropy value weighted coding, first, it is clear that the spatio-temporal entropy weight coefficient is calculated based on the point cloud acquisition timestamp and spatial coordinates, which reflects the information uniqueness and importance of each point in the time and space dimensions. The surface features of the point cloud data, such as the reflection intensity gradient, normal vector offset, and local curvature mutation threshold, describe the characteristics of the point cloud data surface from different angles. For each point cloud data point, its spatio-temporal entropy weight coefficient and surface features are comprehensively considered. Suppose the spatio-temporal entropy weight coefficient of the point is its reflection intensity gradient is the normal vector offset is and the local curvature mutation threshold is . To highlight the information contribution degree of key points, these features are weighted and fused. The comprehensive feature value is obtained After that, entropy value weighted coding is performed on each point cloud data point according to its size. The coding method can adopt a quantization-based method, dividing the comprehensive feature value into different hierarchical intervals, and each interval corresponds to a specific coding value. For example, the comprehensive feature value is divided into intervals , C min represents the minimum value in the comprehensive feature value division interval, that is, the minimum value among all the comprehensive feature values of the point cloud data points , C max represents the maximum value in the comprehensive feature value division interval, that is, the maximum value among all the comprehensive feature values of the point cloud data points . If falls within the interval , then the point cloud data point is encoded as . In this way, the larger the comprehensive feature value of a point, the more its coding value can reflect its importance and will receive more attention in subsequent processing, thus achieving the purpose of highlighting key point information. In addition, for points with relatively small comprehensive feature values, since the information they carry is relatively less or more redundant, their importance in the data is reduced after coding, playing a role in suppressing redundant information. Through this entropy value weighted coding method, the generated point cloud spatio-temporal coding can better retain key information, reduce the data volume, and improve the data processing efficiency.

[0135] The beneficial effects of this step are reflected at multiple levels. In terms of data processing, the point cloud spatio-temporal coding reduces the data volume and lowers the storage and transmission costs. For example, in large-scale highway inspection projects, after a large amount of point cloud data is subjected to entropy value weighted coding, the data volume is significantly reduced, facilitating data storage and transmission and improving the efficiency of data management. In multi-modal fusion, the point cloud spatio-temporal coding that highlights key points can better fuse with other modal data (such as visual images, vibration signals). For example, when fusing with visual images, the key point information in the point cloud spatio-temporal coding can better match the key features in the visual images, improving the accuracy and effect of the fusion and providing richer and more accurate information for subsequent highway inspection and maintenance image analysis. In terms of highway inspection and maintenance decision-making, the encoded data is more convenient for analysis and understanding, helping technicians quickly and accurately identify road problems, such as damage to road facilities, pavement diseases, etc., so as to timely formulate reasonable maintenance strategies, improve the quality and efficiency of highway inspection and maintenance work, and ensure the safety and normal operation of the highway. For example, when analyzing the road surface condition, through the point cloud spatio-temporal coding, the area with sudden local curvature change can be quickly located to judge whether there are pavement potholes and other diseases, providing an accurate basis for road maintenance.

[0136] Step S2200: Obtain the visual image of the highway patrol area, perform cross-modal correlation analysis on the six-axis vibration signal and the visual image, and obtain the weighted visual feature map.

[0137] The purpose of Step S2200 is to obtain the visual image of the highway patrol area, perform cross-modal correlation analysis on it and the six-axis vibration signal, obtain the weighted visual feature map, provide key visual information support for constructing the multi-modal fusion feature vector, and thus improve the accuracy of subsequent highway patrol and maintenance image analysis. In the highway patrol and maintenance scenario, the visual image contains rich information such as road surface conditions and road facilities, while the six-axis vibration signal reflects the motion state of the equipment. Combining the two can comprehensively understand the actual situation during highway patrol. Through cross-modal correlation analysis, the equipment vibration information can be associated with the specific scenarios in the visual image.

[0138] Furthermore, Step S2200 includes:

[0139] Step S2210: Calculate the gradient direction of each pixel point of the visual image to form the gradient direction of the visual image; then calculate the mutual information between each characteristic frequency band of the six-axis vibration signal and the gradient direction of the visual image to obtain the pixel-frequency mutual information matrix I pf ;

[0140] Specifically, mutual information is a concept in information theory, used to measure the degree of mutual dependence between two random variables. In this step, by calculating the mutual information, the correlation degree between the characteristic frequency band of the vibration signal and the gradient direction of the visual image can be quantified.

[0141] For the six-axis vibration signal, the characteristic frequency bands related to the movement speed of the maintenance equipment (such as 0.5 - 20 Hz) have been determined in the previous steps. For the visual image, the gradient direction reflects the direction of brightness change in the image, which can highlight the edge and texture information of the image. When calculating the mutual information, first, each characteristic frequency band of the six-axis vibration signal needs to be processed to convert it into a form corresponding to the pixels of the visual image. For example, the amplitude of the vibration signal at each time point can be mapped to the pixel position of the image (assuming the vibration signal is synchronized with the image acquisition time). Then, for the visual image, calculate the gradient direction of each pixel point. When calculating the gradient direction, common methods are to use edge detection operators such as the Sobel operator. Taking the Sobel operator as an example, it performs convolution operations with the image through convolution kernels in the horizontal and vertical directions to obtain the horizontal direction gradient and the vertical direction gradient , and then calculate the gradient amplitude and the gradient direction . Through the above calculations, the gradient direction of each pixel point of the visual image is obtained, and the gradient direction of the visual image is formed.

[0142] Next, calculate the mutual information between the characteristic frequency band of the vibration signal and the gradient direction of the visual image. Assume that the amplitude sequence of a certain characteristic frequency band of the vibration signal within a certain period of time is , and the gradient direction sequence of a certain pixel point of the visual image is (assuming that their time lengths are the same and synchronous here), and the mutual information is calculated based on the probability distribution. First, count the and joint probability distribution and their respective marginal probability distributions and . Then, use the conventional mutual information formula to calculate and to calculate . Perform the above calculations for all pixel points and the characteristic frequency bands of the vibration signal to obtain the pixel-frequency band mutual information matrix . The element in the matrix represents the mutual information between the -th pixel point and the -th characteristic frequency band of the vibration signal.

[0143] The beneficial effect of this step is significant. The pixel-frequency band mutual information matrix obtained by calculating the mutual information can accurately reflect the internal connection between the vibration signal and the visual image. In the subsequent multi-modal fusion process, can be used as an important reference basis to help determine which pixel points are more closely related to which characteristic frequency bands of the vibration signal. For example, when analyzing road diseases, if the mutual information between the vibration signal of a certain characteristic frequency band and the pixels in a certain area of the image is large, it indicates that this area may be closely related to the vibration change of the equipment, and there may be road diseases or other abnormal conditions, thus guiding technicians to analyze these areas more targeted and improving the accuracy and efficiency of highway inspection and maintenance image analysis. In addition, provides key data for constructing the pixel vibration sensitivity weight model, which helps to achieve more reasonable multi-modal fusion.

[0144] Step S2220, construct a non-linear mapping model w = F(m1, n1, I pf ) of the rigid coefficient m1 of the maintenance equipment, the road surface roughness n1, the pixel-frequency band mutual information matrix I pf and the pixel vibration sensitivity weight w; where F is a non-linear function based on a deep neural network;

[0145] Specifically, the rigidity coefficient m1 of the maintenance equipment reflects the equipment's ability to resist vibration. It is usually related to factors such as the equipment's structural materials and mechanical design and can be obtained through the equipment's technical parameters. The road surface roughness n1 reflects the unevenness of the road surface and can be measured by on-vehicle sensors (such as laser profilometers). The deep neural network (DNN) has powerful non-linear modeling capabilities and can learn complex input-output relationships. When constructing the model w = F(m1, n1, I pf ), m1, n1, and I pf are used as the inputs of the network. The structure of the network can include multiple hidden layers, and each hidden layer consists of multiple neurons. Neurons are connected by weights, and when signals are transmitted between neurons, they are transformed by a non-linear activation function (such as the ReLU function), enabling the network to learn complex non-linear relationships between the input data. When training a deep neural network, a large amount of sample data is required. These sample data include highway inspection data under different equipment rigidity coefficients and road surface roughness, as well as the corresponding pixel-frequency band mutual information matrix and known pixel vibration sensitivity weights (which can be obtained through expert annotation or other reliable methods). Through the backpropagation algorithm, the weights between neurons in the network are continuously adjusted to minimize the error between the pixel vibration sensitivity weight w output by the network and the known true value. For example, using the mean square error (MSE) as the loss function, through multiple iterations of training, the network gradually learns the non-linear mapping relationship between m1, n1, and I pf and w.

[0146] The beneficial effects of this step are reflected in multiple aspects. From the perspective of multimodal fusion, the pixel vibration sensitivity weight w calculated through this non-linear mapping model can more reasonably reflect the impact of the equipment, road surface conditions, and the correlation between vibration and visual images on the pixels. For example, when the equipment rigidity coefficient is large, the equipment has a strong ability to resist vibration, the amplitude of road surface vibration transmitted to the equipment is small, and correspondingly, the image pixels are less affected by vibration, and the weight w is low; conversely, when the road surface roughness is large, the equipment vibration intensifies, the pixels are more affected by vibration, and the weight w is high. Such weights assigned according to the actual situation enable more accurate fusion of vibration signals and visual image information during multimodal fusion, highlighting vibration-related visual features and improving the accuracy of multimodal fusion. In the analysis of highway inspection and maintenance images, accurate pixel vibration sensitivity weights help to more accurately identify road surface diseases and abnormal conditions. For example, in the areas of the image that are closely related to vibration (i.e., the pixel areas with higher weights), there may be road surface diseases such as potholes and cracks. By focusing on the analysis of these areas, the accuracy and efficiency of disease detection can be improved, providing a more reliable basis for highway maintenance decisions.

[0147] Step S2230: Determine m1, n1, and I through the Analytic Hierarchy Process pf The influence weight on w, and substitute it into the model w = F(m1, n1, I pf ), and calculate the pixel vibration sensitivity weight of each pixel point;

[0148] Specifically, the Analytic Hierarchy Process (AHP) is a decision-making method that decomposes elements related to decision-making into levels such as goals, criteria, and solutions, and conducts qualitative and quantitative analysis on this basis. In this step, AHP is used to determine the rigidity coefficient of maintenance equipment , the road surface roughness , the pixel-frequency band mutual information matrix The relative importance of the pixel vibration sensitivity weight ;

[0149] First, construct a hierarchical structure model. Take calculating the pixel vibration sensitivity weight as the target layer; take as the criterion layer; and take each pixel point as the solution layer. Next, construct a judgment matrix. The judgment matrix is the key to AHP, which reflects the decision-maker's judgment on the relative importance of each factor. For every two factors in the criterion layer, determine their relative importance to the target layer (calculating the pixel vibration sensitivity weight ) through pairwise comparison. For example, compare and 's influence on . If it is considered that has a slightly more important influence on than , according to the 1-9 scale method of AHP (where 1 means the two factors are equally important, 3 means one factor is slightly more important than the other, 5 means significantly important, 7 means strongly important, 9 means extremely important, and 2, 4, 6, 8 represent intermediate values between adjacent judgments), assign the corresponding element in the judgment matrix the value of 3; conversely, when comparing with , assign the corresponding element the value of . In this way, a judgment matrix '' can be obtained, where represents the importance of the th factor relative to the th factor to the target layer. Then, calculate the eigenvector and the maximum eigenvalue of the judgment matrix. By calculating the eigenvector of the judgment matrix , the relative weight vector of each factor can be obtained. The specific calculation method is to solve the equation , where is the matrix 's largest eigenvalue. Generally, methods such as the eigenvalue method or the sum-product method are used for calculation. Finally, substitute the calculated weight vector into the model w = F(m1, n1, I pf ). Assume that F is a specific function based on a deep neural network. For each pixel point, is weighted with the corresponding weight (the calculation method here depends on 's specific form, and it may actually be more complex, involving the specific operations of the neural network), so as to obtain the pixel vibration sensitivity weights of each pixel point. The beneficial effects of this step are reflected in many aspects. From the perspective of model accuracy, the weights determined by AHP can more scientifically reflect 's influence degree on , avoiding subjectivity; in terms of multimodal fusion, accurate weights make the calculation of pixel vibration sensitivity weights more in line with the actual situation. Furthermore, when fusing the visual feature map and vibration information, it can more reasonably highlight the visual features related to vibration and improve the quality of multimodal fusion. In the analysis of highway inspection and maintenance images, accurate pixel vibration sensitivity weights help to more accurately locate and analyze abnormal situations such as road surface diseases. For example, for diseases such as road surface cracks, through the reasonably weighted pixel vibration sensitivity weights, these disease areas can be more accurately highlighted in the visual image, providing a more reliable basis for subsequent disease assessment and maintenance decisions.

[0150] Step S2240: Extract features from the visual images of the highway inspection area to obtain a visual feature map;

[0151] Specifically, the visual feature map contains rich semantic information in the image and is an important basis for subsequent multi-modal fusion and image analysis. In this step, a convolutional neural network (CNN) is used to extract features from the visual image. A convolutional neural network is a deep learning model specifically designed for processing image data. It automatically extracts image features through a series of convolutional layers, pooling layers, and activation functions. Before feature extraction, the visual image needs to be preprocessed, including operations such as grayscale conversion (if it is a color image) and normalization. Grayscale conversion is to convert a color image into a grayscale image, reducing the amount of data and facilitating subsequent processing; normalization is to map the pixel values of the image to a specific range (such as [0,1] or [−1,1]), making different images comparable and helping to improve the training effect and stability of the model. The preprocessed image is input into the pre-constructed convolutional neural network. The convolutional layer is one of the core components of the CNN. It performs a convolution operation by sliding a convolution kernel over the image. The convolution kernel is a small matrix, usually with a size of 3×3 or 5×5, etc. The process of the convolution operation is to multiply the elements of the convolution kernel with the local region of the image and sum them up to obtain the convolved feature map. For example, for a 3×3 convolution kernel and a 3×3 local region in the image, each element of the convolution kernel is multiplied by the corresponding element in the image region, and then these products are added together to obtain a pixel value at the corresponding position in the convolved feature map. By using different convolution kernels, different types of features in the image, such as edges and textures, can be extracted. The convolutional layer usually has multiple convolution kernels, thus generating multiple feature maps, and each feature map corresponds to a specific image feature.

[0152] After the convolutional layer, an activation function layer is generally connected. The activation function introduces non-linearity into the neural network, enabling the network to learn more complex functional relationships. Commonly used activation functions include the ReLU function, the Sigmoid function, etc. Taking the ReLU function as an example, it sets the values less than 0 in the output of the convolutional layer to 0 and keeps the values greater than 0 unchanged, which can effectively avoid the vanishing gradient problem and accelerate the training speed of the network. The pooling layer is also an important part of the CNN. It is mainly used to reduce the dimension of the feature map, reduce the computational amount and prevent overfitting. Common pooling methods include max pooling and average pooling. Max pooling selects the maximum value in a local area (such as a 2×2 area) as the output after pooling; average pooling calculates the average value of the local area as the output. For example, when performing max pooling in a 2×2 area, the maximum pixel value in the area is selected as the result after pooling. Through the alternating operations of multiple convolutional layers, activation function layers and pooling layers, the convolutional neural network gradually extracts semantic information at different levels of the image. Shallow convolutional layers can extract low-level features of the image, such as edges, lines, etc.; deep convolutional layers can extract more high-level semantic features, such as the shape and structure of objects. Finally, the network outputs a visual feature map with semantic information at different levels.

[0153] The beneficial effects of this step are reflected at multiple levels. In terms of multi-modal fusion, the visual feature map provides rich visual information for subsequent fusion with other modal data such as vibration signals and point cloud data. For example, when fusing the visual feature map with the vibration frequency domain fingerprint spectrum, features such as edges and textures in the visual feature map can be associated with relevant features in the vibration signal, enhancing the effect of multi-modal fusion. In the analysis of highway inspection and maintenance images, the visual feature map helps to more accurately identify targets such as road surface diseases and road facilities. For example, through the features extracted by the convolutional neural network, the shape, size and location of road surface cracks, potholes and other diseases can be more clearly distinguished, improving the accuracy of disease detection. In addition, the visual feature map can also be used as the input for subsequent tasks such as image classification and object recognition, providing more comprehensive support for highway inspection and maintenance work, and improving the scientificity and accuracy of highway maintenance decisions.

[0154] Step S2250, multiply each pixel vibration sensitivity weight of each pixel point with the visual feature map element by element to obtain a weighted visual feature map.

[0155] Specifically, the purpose of this step is to weight each pixel in the visual feature map according to the pixel vibration sensitivity weight, highlight the visual features related to device vibration, and suppress the information unrelated to vibration, so as to provide more targeted visual information for subsequent multimodal fusion and highway inspection and maintenance image analysis. Before performing element-wise multiplication, the pixel vibration sensitivity weight w of each pixel point and the visual feature map have been calculated through the previous steps. The pixel vibration sensitivity weight w reflects the degree to which each pixel is affected by device vibration, and its value range is usually within a certain interval, such as [0, 1]. The larger the value, the greater the impact of the pixel on vibration and the closer the association with vibration. The visual feature map is obtained by a convolutional neural network extracting features from visual images. It contains rich image semantic information, and each pixel value represents the intensity of the image features at the corresponding position.

[0156] This step has beneficial effects in many aspects. From the perspective of multimodal fusion, the weighted visual feature map can better fuse with other modal data (such as vibration frequency domain fingerprint maps, point cloud spatio-temporal encoding). Since the visual features related to vibration are highlighted, when constructing a multimodal fusion feature vector subsequently, these features can more effectively interact and fuse with other modal information, improving the accuracy and effectiveness of multimodal fusion. For example, when analyzing highway pavement diseases, areas such as pavement cracks that are closely related to vibration are strengthened in the weighted visual feature map. When fused with the vibration frequency domain fingerprint map, it can more accurately reflect the relationship between device vibration and pavement diseases. In terms of highway inspection and maintenance image analysis, the weighted visual feature map helps to improve the recognition accuracy of pavement diseases and anomalies. Because the visual features related to vibration are enhanced, technicians can more clearly observe the areas that may have problems when analyzing the images. For example, when detecting pavement potholes, since the pothole areas often cause device vibration, the corresponding pixels in the visual feature map are weighted, and the features are more obvious, facilitating technicians to quickly discover and locate these disease areas, providing a more reliable basis for highway maintenance decisions. In addition, the weighted visual feature map can also reduce the interference of background information unrelated to vibration, enabling subsequent image analysis to focus more on key areas, improving analysis efficiency, helping to timely discover potential highway safety hazards, and ensuring the normal operation and traffic safety of highways.

[0157] Step S2300: Construct a multimodal fusion feature vector based on the calibrated vibration frequency domain fingerprint map, point cloud spatio-temporal encoding, and weighted visual feature map.

[0158] Step S2300 aims to construct a multi-modal fusion feature vector by fusing the calibrated vibration frequency-domain fingerprint map, point cloud spatio-temporal encoding, and weighted visual feature map, providing key data support for the subsequent generation of high-quality enhanced road inspection and maintenance images. In the road inspection and maintenance scenario, single-modal data often cannot comprehensively and accurately reflect the actual situation. Multi-modal fusion can integrate the advantages of different data sources and improve the accuracy and reliability of road condition analysis. The vibration frequency-domain fingerprint map reflects the vibration characteristics of maintenance equipment, the point cloud spatio-temporal encoding contains the spatial structure information of the road inspection area, and the weighted visual feature map presents the visual details of the road surface and surrounding environment. Fusing these three-modal data can provide richer and more comprehensive information for road inspection and maintenance image analysis.

[0159] Furthermore, as Figure 7 shown, step S2300 includes:

[0160] Step S2310, align the calibrated vibration frequency-domain fingerprint map, point cloud spatio-temporal encoding, and weighted visual feature map in a common reference coordinate system to form a three-dimensional heterogeneous modal feature tensor T; the three-dimensional heterogeneous modal feature tensor T has three channels, namely the visual sub-tensor, the vibration sub-tensor, and the point cloud sub-tensor;

[0161] Specifically, this step is the basis of multimodal fusion. By integrating data from different modalities in the same coordinate system, subsequent unified processing and analysis of this data can be carried out. The common reference coordinate system is a unified coordinate framework used to determine the spatial position relationships of data from different modalities. In actual operation, first, the coordinate transformation relationships of each modality of data in the common reference coordinate system need to be determined. For the vibration frequency-domain fingerprint spectrum, it does not directly correspond to spatial coordinates, but its position relationship in the common reference coordinate system can be indirectly determined by associating it with the motion state of the maintenance equipment and the acquisition position information. For example, assuming the driving trajectory of the maintenance equipment on the road is known, through the positioning information of the equipment (such as GPS data) and the time sequence of vibration signal acquisition, the data in the vibration frequency-domain fingerprint spectrum can be corresponding to specific positions on the road. The point cloud spatio-temporal encoding contains the time stamps and spatial coordinate information of point cloud acquisition. When converting it to the common reference coordinate system, coordinate transformation needs to be carried out according to the parameters of the acquisition device and the characteristics of the acquisition scene. For example, if the installation position and attitude of the laser scanning device are known, the original coordinates of the point cloud data can be converted to the common reference coordinate system through the corresponding rotation and translation transformation matrices. The weighted visual feature map is usually generated based on the images obtained by the image acquisition device, and its coordinate system is related to the image pixel positions. During the alignment process, coordinate transformation needs to be carried out according to the parameters of the image acquisition device (such as focal length, viewing angle, etc.) and the relative position relationship with the common reference coordinate system. For example, through techniques such as camera calibration, the mapping relationship between the image pixel coordinates and the common reference coordinate system can be determined, and each pixel point in the visual feature map can be corresponding to the corresponding position in the common reference coordinate system. After completing the coordinate transformation, the calibrated vibration frequency-domain fingerprint spectrum, point cloud spatio-temporal encoding, and weighted visual feature map are combined according to the channel dimension to form a three-dimensional heterogeneous modality feature tensor T.

[0162] The beneficial effects of this step are significant. From the perspective of data processing, aligning data of different modalities in a common reference coordinate system makes the data spatially consistent, facilitating subsequent unified processing and analysis. For example, when performing feature extraction and fusion operations, calculations can be based on a unified coordinate system, improving the calculation efficiency and accuracy. In terms of multimodal fusion, this alignment method provides a basis for the interaction and fusion between data of different modalities. In the same coordinate system, data of different modalities can be more effectively correlated and fused, uncovering potential connections between the data and enhancing the effect of multimodal fusion. For example, when analyzing highway pavement diseases, the aligned point cloud data and visual feature maps can more accurately match the spatial structure and visual features of the pavement, thus more clearly identifying the location and shape of the diseases. In addition, in subsequent image generation and analysis, multimodal data in a unified coordinate system can provide strong support for generating more realistic and practical images, helping to improve the accuracy and reliability of highway inspection and maintenance image analysis.

[0163] Step S2320: Using an adaptive channel attention mechanism, dynamically assign weights to the visual sub-tensor, vibration sub-tensor, and point cloud sub-tensor of the three-dimensional heterogeneous modal feature tensor T to obtain a weighted feature tensor.

[0164] Specifically, the weighted feature tensor is used to improve the effect of multimodal fusion. The adaptive channel attention mechanism is a method that can automatically adjust the channel weights according to the importance of the features in each channel, enabling the model to pay more attention to important feature information and suppress unimportant information.

[0165] When implementing the adaptive channel attention mechanism, first calculate the mean and variance of the features on each channel. For the visual sub-tensor, assuming its feature representation is , with a size of ( represents the height, represents the width, represents the number of channels , then the calculation method of its mean is: ; where is the index of the height , is the index of the width , represents the channel vector at the position ([[]] , ). The calculation method of the variance is: . For the vibration sub-tensor and point cloud sub-tensor, calculate the mean and variance using a similar method.

[0166] Judge the importance of the channel features according to the magnitude of the mean and variance. Generally speaking, for channels with larger mean and variance, it indicates that these channels contain more important information, and this information may be more crucial for accurately describing the highway inspection and maintenance scenario in multimodal fusion. For example, when analyzing road surface diseases, if the mean and variance of a certain channel in the visual subtensor are large in the disease area, it means that this channel may contain the key visual features of the disease, such as the edge information of cracks or the texture features of potholes. Calculate the weight of each channel based on the mean and variance. Assume that the weight vector of all channels is where , correspond to the weights of the visual subtensor, vibration subtensor, and point cloud subtensor respectively. Taking the weight of the visual subtensor as an example, the weight is calculated using the following formula: ; where represents the vibration subtensor in the three-dimensional heterogeneous modal feature tensor T, represents the point cloud subtensor in the three-dimensional heterogeneous modal feature tensor T, is a set containing three elements; is 's index, used to traverse the elements in the set . is a mean variable that changes with the index , and takes values from the set . For example, when = , represents the mean of the visual subtensor ; is the same reason, and it is a variance variable that changes with the index . After obtaining the weight of each channel, perform a weighting operation on each subtensor of the three-dimensional heterogeneous modal feature tensor T to obtain the weighted feature tensor.

[0167] ​​​The beneficial effects of this step are reflected at multiple levels. In terms of multimodal fusion, through the adaptive channel attention mechanism, the model can automatically focus on important feature channels, enhance the utilization of key information, suppress redundant information, and thus improve the quality of multimodal fusion. For example, when fusing vibration, point cloud, and visual information for road surface disease analysis, it can highlight features related to diseases, such as the features reflecting abnormal vibration of equipment in the vibration sub-tensor, the spatial structure features of the disease area in the point cloud sub-tensor, and the texture features of the disease in the visual sub-tensor, making the fused features more representative. In the analysis of highway inspection and maintenance images, the weighted feature tensor helps to more accurately identify targets such as road surface diseases and road facilities. Since important features are enhanced, technicians can more clearly observe the details and features of the targets when analyzing images, improving the accuracy and efficiency of disease detection. In addition, this dynamic weight allocation method can also improve the adaptability of the model, enabling it to more effectively fuse multimodal data in different highway inspection scenarios and providing a more reliable basis for highway maintenance decision-making.

[0168] Step S2330: Superimpose the spatio-temporal entropy weight mask and perform pooling processing on the weighted feature tensor to generate a multimodal fusion feature vector.

[0169] Specifically, this step aims to further optimize the result of multimodal fusion. By suppressing redundant information and highlighting key information, it generates a multimodal fusion feature vector with stronger representation ability, providing high-quality data input for subsequent adversarial generative network training and highway inspection and maintenance image analysis. The spatio-temporal entropy weight mask is generated based on the spatio-temporal entropy information of the point cloud data. In step S2120, the spatio-temporal entropy weight coefficients of each point cloud data point have been calculated, and these coefficients reflect the information importance of the point cloud data in the time and space dimensions. The spatio-temporal entropy weight mask is to expand these weight coefficients to the same dimension and structure as the weighted feature tensor, so that each element at a position corresponds to a spatio-temporal entropy weight value. The operation of superimposing the spatio-temporal entropy weight mask is to multiply the weighted feature tensor and the spatio-temporal entropy weight mask element by element to obtain the feature tensor after superimposing the mask. In this way, the positions with higher entropy values (i.e., more important information) in space and time are enhanced, while the positions with lower entropy values (relatively redundant information) are suppressed. For example, in a highway inspection scenario, if the point cloud data in a certain area has a high entropy value in time and space, it indicates that this area contains important information, such as the key parts of road facilities or disease areas. After superimposing the spatio-temporal entropy weight mask, the features of these areas in the feature tensor will be enhanced.

[0170] Pooling processing is a commonly used dimensionality reduction operation, aiming to reduce the dimension of data while retaining important feature information. In this step, for the feature tensor T after superimposing the spatio-temporal entropy weight mask maskedPerform pooling. Common pooling methods include max pooling and average pooling. Through pooling, the dimension of the feature tensor is reduced to generate a multi-modal fusion feature vector.

[0171] The beneficial effects of this step are reflected in multiple aspects. In terms of data processing, superimposing the spatio-temporal entropy weight mask and pooling can effectively suppress redundant information, reduce the data volume, lower the computational complexity, and improve the efficiency of subsequent model training and analysis. For example, when processing large-scale highway inspection data, the data volume is significantly reduced after these operations, while key information is retained, enabling the subsequent training of the adversarial generation network to converge faster. In terms of multi-modal fusion, by highlighting key information and suppressing redundant information, the generated multi-modal fusion feature vector has stronger representation ability and can more accurately reflect the characteristics of the highway inspection and maintenance scenario. For example, when analyzing road surface diseases, the multi-modal fusion feature vector can more concentratedly reflect the key features of the diseases, helping to improve the accuracy of disease identification. In the analysis of highway inspection and maintenance images, this high-quality multi-modal fusion feature vector provides a solid foundation for subsequent image generation and analysis, enabling the generation of more realistic and detailed enhanced images, improving the accuracy of highway inspection and maintenance condition analysis, and providing a more reliable basis for highway maintenance decision-making.

[0172] Step S3000: Based on the multi-modal fusion feature vector, construct and train an adversarial generation network including a generator and a discriminator; obtain a new multi-modal fusion feature vector, and generate an enhanced highway inspection and maintenance image according to the new multi-modal fusion feature vector and the trained adversarial generation network.

[0173] Furthermore, step S3000 includes:

[0174] Step S3100: Construct and train an adversarial generation network including a generator and a discriminator;

[0175] The core purpose of step S3100 is to construct and train an adversarial generation network including a generator and a discriminator, which is a key link for generating high-quality enhanced highway inspection and maintenance images. The adversarial generation network (GAN) consists of a generator and a discriminator, which confront and cooperate with each other in training. Through this mechanism, it can learn the distribution of real data and thus generate more realistic image data. In the application scenario of highway inspection and maintenance images, due to problems such as insufficient data volume and uneven image quality in the actually collected images, the adversarial generation network can be used to enhance the existing image data, generate more images with rich details and features, and provide more sufficient and high-quality data support for subsequent highway condition analysis.

[0176] Furthermore, step S3100 includes:

[0177] Step S3110: Construct the generator of the generative adversarial network;

[0178] Specifically, step S3110 aims to construct the generator of the generative adversarial network, select the structure of a multi-layer convolutional neural network (CNN), and optimize the design for the multi-modal fusion feature vector to achieve the transformation from the multi-modal fusion feature vector to high-quality highway inspection and maintenance images. The multi-layer convolutional neural network has powerful feature extraction and image generation capabilities in the field of image processing. Through the combination of components such as convolutional layers, activation function layers, and pooling layers, it can automatically learn the feature representation of images.

[0179] When constructing the convolutional layer of the generator, refer to the DCGAN (Deep Convolutional Generative Adversarial Network) architecture and make adjustments according to the characteristics of the multi-modal fusion feature vector. The DCGAN architecture is a convolutional neural network architecture successfully applied to image generation. It maps low-dimensional vectors to the high-resolution image space through a series of transposed convolutional operations. For the multi-modal fusion feature vector in this embodiment, it contains calibrated vibration frequency-domain fingerprint maps, point cloud spatio-temporal encoding, and weighted visual feature maps and other information. These information have different modal characteristics, and the convolution kernel size and stride need to be set specifically. For the vibration frequency-domain fingerprint map features, because they contain fine frequency-related information, use a smaller convolution kernel (such as 3×3). The smaller convolution kernel can more finely capture the changes in frequency features in the local area. For example, it can accurately extract details such as the energy distribution changes in a specific frequency band. Suppose the energy change of the vibration frequency-domain fingerprint map shows a local peak feature in a certain frequency band. The smaller convolution kernel can more sensitively perceive this change, so that when generating images, it can more accurately reflect the image details related to this frequency feature, such as the performance of the equipment vibration caused by the corresponding road conditions in the image. For the point cloud spatio-temporal encoding, since the spatial structure features of the point cloud data are relatively macroscopic, use a slightly larger convolution kernel (such as 5×5). The larger convolution kernel can capture the structure information of the point cloud data in a wider spatial range, which helps to better restore the scene spatial layout represented by the point cloud when generating images. For example, when processing point cloud data containing roads and surrounding facilities, the larger convolution kernel can integrate the point cloud information in a wider area, making the spatial structure features such as the shape, width of the road and its relative position relationship with the surrounding facilities in the generated image more accurate. For the weighted visual feature map, combine the texture and detail characteristics of the image and select an appropriate convolution kernel (such as 4×4). This size of convolution kernel can not only capture the local texture details of the image, but also consider the context information of the surrounding pixels to a certain extent. For example, when generating an image containing road surface texture, the 4×4 convolution kernel can effectively extract the local features of the texture, such as the texture orientation of road cracks and the edge features of potholes, while taking into account the information of the surrounding pixels, making the generated image texture more natural and coherent.

[0180] In this way, the input layer of the generator receives the multi-modal fusion feature vectors constructed in step S2000. Each convolutional layer gradually processes and transforms these features, and finally generates a highway inspection and maintenance image with high resolution. The beneficial effects of this step are remarkable. In terms of image generation quality, the targeted convolutional kernel design enables the generator to fully exploit the information in the multi-modal fusion feature vectors, and the generated images are richer and more accurate in terms of details and features. For example, the generated images can more realistically present the texture of the road surface, the shape of diseases, and the details of road facilities, improving the visualization effect and readability of the images. In terms of multi-modal fusion applications, this design helps to better fuse the data features of different modalities, so that the generated images not only contain visual information, but also can reflect the information reflected by vibration and point cloud data, enhancing the overall expression ability of the images for highway inspection and maintenance scenarios, and providing a more comprehensive and accurate data basis for subsequent image analysis.

[0181] Step S3120: Embed a vibration-blur conversion layer at the front end of the generator in the adversarial generation network;

[0182] Specifically, the main task of step S3120 is to embed a vibration-blur conversion layer at the front end of the generator in the adversarial generation network. This layer works based on the vibration frequency domain fingerprint map calibrated in step S1000, aiming to generate a realistic motion blur effect according to the vibration state of the device, making the generated highway inspection and maintenance images more realistic. During the working process, first, the motion state parameters of the device at different times, such as acceleration, angular velocity, etc., are calculated according to the characteristic information in the vibration frequency domain fingerprint map. The vibration frequency domain fingerprint map contains rich information about the vibration of the device. By analyzing the characteristics such as the main frequency band energy distribution, harmonic distortion rate, and inter-axis coupling coefficient of the vibration frequency domain fingerprint map, the motion state of the device at different times can be inferred. For example, according to the change of vibration energy in a specific frequency band, it can be judged whether the device is in an accelerating, decelerating, or uniform motion state; through the inter-axis coupling coefficient, the vibration correlation of the device in different directions can be understood, and then the change of its motion posture can be inferred. After obtaining the motion state parameters of the device, these parameters are used as inputs, and combined with the Newton-Euler kinematic model, a blur kernel function is dynamically generated. The Newton-Euler kinematic model is a classical mechanics model for describing the motion of objects. It can calculate the motion trajectory and posture change of an object in space according to parameters such as the acceleration and angular velocity of the object. In this step, this model is used to simulate the vibration and displacement of the device during the motion process, so as to generate a blur kernel function that matches the actual motion state of the device.

[0183] The generation process of the blur kernel takes into account the movement direction, speed, and vibration conditions of the device. When the device moves at a relatively high speed and has a large vibration amplitude, the generated blur kernel size increases accordingly to simulate a more obvious motion blur effect. For example, when the device is driving at a high speed and the road surface is bumpy, resulting in large vibrations, the image will be blurred to a greater extent. At this time, the generated blur kernel size is large, making the generated image show a more obvious blur effect in the corresponding area and more realistically reflecting the actual shooting situation. Conversely, when the device moves at a relatively low speed and has small vibrations, the blur kernel size decreases. To achieve the linkage between the kernel size and the device speed error threshold, a speed error threshold range is set. When the calculated device speed error exceeds this threshold, the size of the blur kernel is adjusted according to a certain proportional relationship. Assume that the speed error threshold range is [-ε, ε]. When the calculated device speed error is greater than ε, according to the pre-set proportional coefficient k z , the size of the blur kernel is increased; when the speed error is less than -ε, the size of the blur kernel is also decreased according to the proportional coefficient k z . This can ensure that the generated blur effect closely matches the actual motion state of the device and improve the authenticity of the generated image.

[0184] The beneficial effects of this step are reflected at multiple levels. In terms of image authenticity, the motion blur effect generated by the vibration-blur conversion layer makes the generated highway inspection and maintenance images more in line with the actual shooting scene. For example, when analyzing images taken by maintenance equipment driving on the road, the real image will be blurred to a certain extent due to the movement of the equipment. The blur effect simulated by this conversion layer can restore this real situation, helping subsequent analysts to more accurately judge the information in the image. In terms of image analysis accuracy, reasonable motion blur simulation can avoid analysis errors caused by ignoring blur factors. For example, when detecting road diseases, if the image does not correctly simulate motion blur, it may cause deviations in the edges and details of the diseases, affecting the detection accuracy. The real blur image generated by this step can improve the accuracy of disease detection and other image analysis tasks, providing a more reliable basis for highway maintenance decisions.

[0185] Step S3130: Construct a discriminator for the generative adversarial network. The discriminator adopts a two-stream verification architecture, and the two-stream verification architecture includes a video stream part and a point cloud stream part.

[0186] Specifically, the core of step S3130 is to construct a discriminator for the generative adversarial network. This discriminator adopts a two-stream verification architecture, including a video stream part and a point cloud stream part. The purpose is to more accurately judge the quality of the generated images through a multi-dimensional evaluation method, and then guide the generator to generate road inspection and maintenance images closer to the real ones. In the visual stream part of the discriminator, an asymmetric wavelet decomposition loss function is used to evaluate the texture similarity between the generated image and the real image. Wavelet decomposition is a mathematical method of decomposing a signal into different frequency sub-bands. Asymmetric wavelet decomposition can more effectively extract the texture features of an image. During operation, first perform asymmetric wavelet decomposition on the generated image and the real image respectively, decomposing the images into sub-bands of different frequencies. Selecting an appropriate wavelet basis function (such as the db4 wavelet basis) is crucial, as it determines the effect of wavelet decomposition on extracting the texture features of the image. The db4 wavelet basis can better capture the edge and texture information of different scales in the image when dealing with image textures. For each sub-band, calculate the difference between the corresponding sub-bands of the generated image and the real image. Here, the mean square error (MSE) is used to measure the pixel difference between the sub-bands. The mean square error quantifies the difference degree between two images by calculating the average of the sum of the squares of the differences between the corresponding pixel values of the two images. Then, according to the importance of different sub-bands to the image texture, assign weights to the differences of each sub-band. High-frequency sub-bands usually contain detailed texture information of the image, such as cracks on the road surface, details of signs, etc. These information are crucial for judging the authenticity and quality of the image, so higher weights are given; low-frequency sub-bands mainly reflect the general outline of the image, such as the overall shape of the road, regional distribution, etc., and the weights are relatively low. Finally, sum up the weighted differences of all sub-bands to obtain the loss value of the visual stream, and use this to evaluate the texture similarity between the generated image and the real image.

[0187] In the point cloud stream part, a joint probability model of reflection - shadow is constructed based on the Markov random field to verify the topological consistency between the shadow area of the generated image and the reflection intensity distribution of the point cloud. First, the reflection intensity information in the point cloud data is associated with the shadow area of the generated image. For each pixel point in the generated image, its corresponding position in the point cloud data is determined, and the reflection intensity value at this position is obtained. The point cloud data records information such as the three - dimensional coordinates and reflection intensity of the object surfaces in the scene. By corresponding the generated image pixels with the point cloud data, a connection between the two can be established. Then, using the Markov random field model, considering the spatial neighborhood relationship between pixel points, a joint probability model of reflection - shadow is constructed. The Markov random field is a probability - based model that assumes the state of a pixel point is only related to the states of its neighboring pixel points. During the model construction process, the parameters of the model, such as the potential function of the nodes and the weights of the edges, are determined to describe the mutual influence relationship between pixel points. The potential function defines the state energy of a single pixel point, and the edge weight represents the interaction intensity between adjacent pixel points. By calculating the joint probability of the shadow area of the generated image and the reflection intensity distribution of the point cloud under this model, the topological consistency between the two is evaluated. If the joint probability is high, it indicates that the topological consistency between the shadow area of the generated image and the reflection intensity distribution of the point cloud is good; otherwise, the consistency is poor.

[0188] The evaluation results of the visual stream and the point cloud stream are fused to obtain the final discrimination result of the discriminator. The fusion method uses weighted summation. According to the importance of visual texture and point cloud topological consistency in the actual application scenario, different weights are assigned to the visual stream loss value and the point cloud stream joint probability. For example, in a scenario that emphasizes image texture details, such as accurately detecting road surface diseases, the accuracy of the road surface texture is crucial for judging the type and degree of diseases. At this time, the weight of the visual stream loss value can be appropriately increased; in a scenario with high requirements for the accuracy of the scene structure, such as analyzing the spatial layout of road facilities, the point cloud topological consistency is more critical, and the weight of the point cloud stream joint probability should be increased. By adjusting the weights, the discriminator can more accurately judge the quality of the generated image.

[0189] The beneficial effects of this step are significant. In terms of image quality assessment, through the dual evaluation of the visual flow and the point cloud flow, the discriminator can comprehensively judge the quality of the generated image from multiple perspectives, which is more accurate and reliable than a single evaluation method. For example, relying solely on the texture evaluation of visual images may ignore the spatial structure information of objects in the image, while combining the topological consistency evaluation of point cloud data can more comprehensively judge the authenticity of the image. In terms of guiding the training of the generator, accurate discrimination results can provide more effective feedback for the training of the generator, prompting the generator to continuously adjust the strategy of generating images and generate highway inspection and maintenance images closer to the real situation, improving the quality and practicality of the generated images. In terms of analyzing highway inspection and maintenance images, high-quality generated images help to more accurately identify targets such as road surface diseases and road facilities, improving the accuracy and efficiency of analysis and providing a more reliable basis for highway maintenance decisions.

[0190] Step S3140: Train the constructed adversarial generation network according to the multi-modal fusion feature vector.

[0191] Specifically, the main task of step S3140 is to train the constructed adversarial generation network according to the multi-modal fusion feature vector. Through the continuous confrontation and optimization of the generator and the discriminator, the image generated by the generator becomes closer and closer to the real highway inspection and maintenance image, achieving the purpose of adversarial training. During the training process, the multi-modal fusion feature vector is used as the input of the generator. The multi-modal fusion feature vector contains rich information such as the calibrated vibration frequency domain fingerprint map, the spatio-temporal encoding of the point cloud, and the weighted visual feature map. These information provide a basis for the generator to generate high-quality images. The generator generates simulated highway inspection and maintenance images according to the input feature vector. Since the generator is constructed based on a multi-layer convolutional neural network structure, it gradually processes and transforms the multi-modal fusion feature vector. Under the action of components such as the convolutional layer and the activation function layer, the abstract feature vector is transformed into an image with a certain resolution.

[0192] The discriminator receives the generated simulated image and the real highway inspection and maintenance image. Among them, the video stream part analyzes the dynamic features of the image and evaluates the texture similarity between the generated image and the real image through the asymmetric wavelet decomposition loss function. As mentioned above, the asymmetric wavelet decomposition decomposes the image into sub-bands of different frequencies. By calculating the mean square error between the sub-bands and combining the weights of different sub-bands, the loss value of the visual flow is obtained to measure the difference in texture between the generated image and the real image. The point cloud stream part judges the authenticity of the image from the perspective of the relevant features of the point cloud data. Based on the reflection-shadow joint probability model constructed by the Markov random field, it verifies the topological consistency between the shadow area of the generated image and the point cloud reflection intensity distribution, and calculates the joint probability to evaluate the degree of consistency between the two.

[0193] By continuously adjusting the parameters of the generator and the discriminator, the images generated by the generator become closer and closer to real images. In the initial stage of training, the images generated by the generator may have a large difference from real images. The discriminator can easily identify these differences and give a low evaluation score (for the judgment of the quality of the generated images, the evaluation results such as the visual flow loss value and the joint probability of the point cloud flow can be mapped to an evaluation score through a certain function). At this time, according to the feedback of the discriminator, the generator adjusts its own parameters through the backpropagation algorithm. The backpropagation algorithm is an optimization algorithm for training neural networks. It calculates the gradient of the loss function with respect to the network parameters, and propagates the gradient information from the output layer back to the input layer, gradually adjusting the weights and biases of each layer in the network, so that the images generated by the generator are closer to real images in subsequent iterations.

[0194] At the same time, the discriminator is also continuously optimizing its own parameters to better identify the differences between the generated images and real images. For example, in each training iteration, the discriminator updates its internal parameters according to the new generated images and real images, improving the accuracy of the judgment of image quality. As the training progresses, the generator gradually learns the features and distribution laws of real images, and the quality of the generated images continues to improve. The judgment difficulty of the discriminator also gradually increases. When the images generated by the generator can deceive the discriminator and make it difficult to distinguish the differences between the generated images and real images, the purpose of adversarial training is achieved.

[0195] The beneficial effects of this step are reflected in many aspects. In terms of improving the quality of image generation, through adversarial training, the generator can learn the rich features and distribution laws of real images, and the generated highway inspection and maintenance images are getting closer and closer to the real situation in terms of details, textures, structures, etc. For example, the generated images can more accurately present the shape, size, and location of pavement diseases, as well as the appearance and status of road facilities, improving the authenticity and usability of the images. In terms of enhancing the adaptability of the model, the training process of the adversarial generation network enables the model to adapt to different highway inspection and maintenance scenarios and data characteristics. Since the actual collected data may be diverse and complex, through adversarial training, the generator and the discriminator can continuously adjust and optimize, improving the model's processing ability and adaptability to various data. In terms of the application of highway inspection and maintenance image analysis, high-quality generated images provide more reliable data support for subsequent analysis tasks. For example, in tasks such as image-based pavement disease detection and road facility assessment, the generated enhanced images can help the detection model learn more comprehensive features, improving the accuracy and efficiency of detection, and providing a more accurate basis for highway maintenance decisions.

[0196] Step S3200, use the trained adversarial generation network to input a new multi-modal fusion feature vector to generate enhanced highway inspection and maintenance images.

[0197] Specifically, the purpose of step S3200 is to use the trained adversarial generative network to input new multi-modal fusion feature vectors and generate enhanced highway inspection and maintenance images. These enhanced images are intended to provide richer and more accurate information for subsequent analysis of highway inspection and maintenance conditions. When performing this step, new multi-modal fusion feature vectors are first obtained. These new feature vectors also contain information such as calibrated vibration frequency domain fingerprint maps, point cloud spatio-temporal encoding, and weighted visual feature maps. They may come from highway inspection data at different times and different sections, or be obtained after different processing of existing data. For example, in a new highway inspection, the point cloud data, visual images, and corresponding equipment vibration data collected are processed through the same process as before, that is, the point cloud data is encoded with entropy value weighting, and the visual images are subjected to cross-modal correlation analysis, etc., to obtain new multi-modal fusion feature vectors.

[0198] The new multi-modal fusion feature vectors are input into the generator of the trained adversarial generative network. Since the generator has learned how to generate patterns and rules for generating highway inspection and maintenance images close to the real ones during the previous training process, it can now generate enhanced images based on the new input. The multi-layer convolutional neural network structure inside the generator processes the input feature vectors step by step according to the trained parameters and weights. For example, for the vibration frequency domain fingerprint map features, a smaller convolutional kernel (such as 3×3) will finely capture its frequency-related features and transform these features into relevant details in the image; for the point cloud spatio-temporal encoding, a slightly larger convolutional kernel (such as 5×5) will integrate its spatial structure information to make the spatial layout of the road and the surrounding environment in the generated image more reasonable; for the weighted visual feature map, a suitable convolutional kernel (such as 4×4) will highlight the texture and detail characteristics of the image, making the road surface texture, road signs, etc. in the generated image clearer. These enhanced images are richer than the original images in terms of details and features. For example, when detecting road surface diseases, the enhanced images may more clearly show the direction, width of cracks, and the depth and edge details of potholes. In terms of road facility assessment, features such as the text, color of road signs, and the shape, material of guardrails in the image will also be more obvious. This is because during the training process of the adversarial generative network, the generator is continuously optimized to generate images closer to the real ones, and the feedback of the discriminator prompts the generator to pay attention to various details and features of the image, resulting in a significant improvement in the quality of the generated images.

[0199] In terms of beneficial effects, in the analysis of highway inspection and maintenance conditions, the enhanced images contribute to improving the accuracy of the analysis. Taking pavement disease detection as an example, clearer image details enable the detection algorithm to more accurately identify the type, scope, and severity of diseases. Traditional image analysis may have difficulty accurately judging some minor diseases or diseases in complex backgrounds due to image quality problems, while the enhanced images can provide more abundant information, reducing the situations of misjudgment and missed judgment. For road facility assessment, rich image features help assessors more comprehensively understand the status of facilities, such as determining whether road signs are clear and distinguishable and whether they need to be replaced, and whether there are damages or deformations in guardrails, etc., so as to timely formulate maintenance plans and ensure the safety and normal use of the road. In terms of data supplementation, the enhanced images can expand the image dataset of highway inspection and maintenance, provide more diverse data samples for subsequent model training, algorithm optimization, etc., further improve the overall level of highway inspection and maintenance image analysis technology, and promote the efficient development of highway maintenance work.

[0200] Embodiment 2

[0201] Based on Embodiment 1, this embodiment provides an adversarial generation data enhancement system for highway inspection and maintenance images, as Figure 8 shown, including:

[0202] Spectrum construction module: used to collect the six-axis vibration signals of maintenance equipment in real time and construct a vibration frequency-domain fingerprint spectrum; perform frequency-domain decomposition on the collected six-axis vibration signals based on a sliding time window, extract the characteristic frequency bands related to the movement speed of the maintenance equipment, and generate a vibration-speed mapping relationship matrix; dynamically correct the vibration-speed mapping relationship matrix and output the calibrated vibration frequency-domain fingerprint spectrum;

[0203] Feature fusion module: used to collect the point cloud data of the highway inspection area, perform entropy value weighted coding on the point cloud data to generate point cloud spatio-temporal coding; obtain the visual image of the highway inspection area, perform cross-modal correlation analysis on the six-axis vibration signals and the visual image to obtain the weighted visual feature map; construct a multi-modal fusion feature vector according to the calibrated vibration frequency-domain fingerprint spectrum, point cloud spatio-temporal coding, and weighted visual feature map;

[0204] Adversarial generation module: construct and train an adversarial generation network including a generator and a discriminator according to the multi-modal fusion feature vector; obtain a new multi-modal fusion feature vector, and generate enhanced highway inspection and maintenance images according to the new multi-modal fusion feature vector and the trained adversarial generation network.

[0205] The methods and systems of the present application can be implemented in many ways. For example, the methods and systems of the present application can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present application are not limited to the order specifically described above, unless otherwise specifically stated.

[0206] In addition, parts of the above technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive elaboration.

[0207] As described above in the specific embodiments, the objectives, technical solutions, and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An adversarial generation data enhancement method for highway inspection and maintenance images, characterized in that, The method includes: Collecting six-axis vibration signals of the maintenance equipment in real time, and constructing a vibration frequency-domain fingerprint map; performing frequency-domain decomposition on the collected six-axis vibration signals based on a sliding time window, extracting characteristic frequency bands related to the movement speed of the maintenance equipment, and generating a vibration-speed mapping relationship matrix; dynamically correcting the vibration-speed mapping relationship matrix, and outputting a calibrated vibration frequency-domain fingerprint map; Collecting point cloud data of the highway inspection area, performing entropy-weighted encoding on the point cloud data to generate a point cloud spatio-temporal code; obtaining a visual image of the highway inspection area, performing cross-modal correlation analysis on the six-axis vibration signal and the visual image to obtain a weighted visual feature map; constructing a multi-modal fusion feature vector according to the calibrated vibration frequency-domain fingerprint map, the point cloud spatio-temporal code, and the weighted visual feature map; The generating of the point cloud spatio-temporal code includes: collecting point cloud data of the highway inspection area, recording the point cloud acquisition timestamp and spatial coordinates; calculating the spatio-temporal entropy weight coefficient of each point based on the point cloud acquisition timestamp and spatial coordinates; extracting the surface features of the point cloud data, where the surface features include reflection intensity gradient, normal vector offset, and local curvature mutation threshold; performing entropy-weighted encoding on the point cloud data according to the spatio-temporal entropy weight coefficient of each point and the surface features of the point cloud data to generate a point cloud spatio-temporal code; Constructing and training an adversarial generation network including a generator and a discriminator according to the multi-modal fusion feature vector; obtaining a new multi-modal fusion feature vector, and generating an enhanced highway inspection and maintenance image according to the new multi-modal fusion feature vector and the trained adversarial generation network.

2. The adversarial generation data augmentation method for highway inspection and maintenance images according to claim 1, wherein The six-axis vibration signals of the real-time acquisition and maintenance equipment include: installing an inertial measurement unit on the highway inspection and maintenance equipment, and according to the installed inertial measurement unit, real-time collecting the six-axis vibration signals of the maintenance equipment during driving; the six-axis vibration signals include acceleration signals a x (t), a y (t), a z (t) in the three directions of the X, Y, and Z axes and angular velocity signals ω x (t), ω y (t), ω z (t).

3. The adversarial generation data enhancement method for highway inspection and maintenance images according to claim 2, characterized in that The constructing of the vibration frequency-domain fingerprint map includes: Performing frequency-domain conversion on the collected six-axis vibration signals to obtain the main frequency band energy distribution and harmonic distortion rate; analyzing the mutual relationship between the acceleration signals in the three directions of the X, Y, and Z axes and the angular velocity signals in the three directions of the X, Y, and Z axes to obtain the inter-axis coupling coefficient; integrating the main frequency band energy distribution, the harmonic distortion rate, and the inter-axis coupling coefficient to construct a vibration frequency-domain fingerprint map.

4. The adversarial generation data augmentation method for highway inspection and maintenance images according to claim 3, wherein The obtaining of the inter-axis coupling coefficient includes: Align the acceleration signals a x (t), a y (t), a z (t) in the three directions of the X, Y, and Z axes and the angular velocity signals ω x (t), ω y (t), ω z (t) with timestamps; According to the aligned a x (t), a y (t), a z (t), ω x (t), ω y (t), ω z (t), construct a multi - variable signal space matrix; Obtaining a correlation coefficient matrix R and a non-linear correlation matrix N according to a multi-signal space matrix; Fuse the correlation coefficient matrix and the non - linear correlation matrix to obtain the comprehensive correlation matrix C, and extract the inter - axis coupling coefficient from the comprehensive correlation matrix C.

5. The adversarial generation data enhancement method for highway inspection and maintenance images according to claim 4, characterized in that The constructing of the multi-signal space matrix includes: The aligned a x (t), a y (t), a z (t), ω x (t), ω y (t), ω z are combined into a six-dimensional signal vector ; Over a period of time Collect six-dimensional signal vectors at multiple moments to construct a multi-dimensional signal space matrix wherein is the six-dimensional signal vector at the nth moment.

6. The adversarial generation data enhancement method for highway inspection and maintenance images according to claim 5, characterized in that, The obtaining of the correlation coefficient matrix R includes: For the multi - variable signal space matrix For any two components and calculate and the correlation coefficients between all different axes, obtaining a correlation coefficient matrix , the element in the correlation coefficient matrix represents the correlation between the signals of the -th axis and the -th axis; where is the six - dimensional signal vector at the i - th moment, is the six - dimensional signal vector at the j - th moment, 1 ≤ i ≤ n, 1 ≤ j ≤ n, .

7. The adversarial generation data enhancement method for highway inspection and maintenance images according to claim 6, characterized in that, Obtaining the non - linear correlation matrix N includes: calculating and the mutual information value of , and constructing the non - linear correlation matrix N according to the mutual information value .

8. The adversarial generation data augmentation method for highway inspection and maintenance images according to claim 1, wherein The generating of the vibration-speed mapping relationship matrix includes: Perform time-frequency representation on the six-axis vibration signals of each characteristic frequency band, and extract the amplitude statistics of each characteristic frequency band; obtain the moving speed of the maintenance equipment, correlate the amplitude statistics of each characteristic frequency band with the moving speed of the maintenance equipment, and establish a vibration-speed mapping relationship matrix M vs ; among them, the matrix element M vs (f i' ,v j' ) represents the amplitude statistic corresponding to the i'-th characteristic frequency band f i' and the j'-th speed value v j' .

9. The adversarial generation data augmentation method for highway inspection and maintenance images according to claim 8, characterized in that, The outputting of the calibrated vibration frequency-domain fingerprint map includes: Dynamically correct the vibration-velocity mapping relationship matrix through a Kalman filter to obtain the corrected M vs matrix; according to the corrected M vs matrix, calibrate the vibration frequency-domain fingerprint spectrum to obtain the calibrated vibration frequency-domain fingerprint spectrum.

10. The adversarial generation data augmentation method for highway inspection and maintenance images according to claim 1, wherein, The obtaining of the weighted visual feature map includes: Calculate the gradient direction of each pixel of the visual image to form the gradient direction of the visual image; calculate the mutual information between each characteristic frequency band of the six-axis vibration signal and the gradient direction of the visual image to obtain the pixel-frequency band mutual information matrix I pf ; construct the nonlinear mapping model w = F(m1, n1, I pf ) of the rigidity coefficient m1 of the maintenance equipment, the road surface roughness n1, the pixel-frequency band mutual information matrix I pf and the pixel vibration sensitivity weight w; where F is a nonlinear function based on a deep neural network; determine the influence weights of m1, n1, and I pf on w through the analytic hierarchy process, and substitute them into the model w = F(m1, n1, I pf ) to calculate the pixel vibration sensitivity weight of each pixel point; extract features from the visual image of the highway inspection area to obtain the visual feature map; multiply the pixel vibration sensitivity weight of each pixel point element-wise with the visual feature map to obtain the weighted visual feature map.

11. An adversarial generation data augmentation system for highway inspection and maintenance images, which is used to implement the adversarial generation data augmentation method for highway inspection and maintenance images described in any one of claims 1-10, and is characterized in that, The system includes: A map construction module: used for collecting six-axis vibration signals of the maintenance equipment in real time, constructing a vibration frequency-domain fingerprint map; performing frequency-domain decomposition on the collected six-axis vibration signals based on a sliding time window, extracting characteristic frequency bands related to the movement speed of the maintenance equipment, generating a vibration-speed mapping relationship matrix; dynamically correcting the vibration-speed mapping relationship matrix, and outputting a calibrated vibration frequency-domain fingerprint map; Feature Fusion Module: It is used to collect the point cloud data of the highway inspection area, perform entropy value weighted coding on the point cloud data to generate point cloud spatio-temporal coding; obtain the visual image of the highway inspection area, perform cross-modal correlation analysis on the six-axis vibration signal and the visual image to obtain the weighted visual feature map; construct a multi-modal fusion feature vector according to the calibrated vibration frequency domain fingerprint spectrum, point cloud spatio-temporal coding and weighted visual feature map. Adversarial Generation Module: According to the multi-modal fusion feature vector, construct and train an adversarial generation network including a generator and a discriminator; obtain a new multi-modal fusion feature vector, and generate an enhanced highway inspection and maintenance image according to the new multi-modal fusion feature vector and the trained adversarial generation network.

Citation Information

Patent Citations

  • Visualized highway maintenance decision-making system

    CN107153928B

  • Road patrol maintenance system

    CN118333608A

  • Pavement crack defect detection method based on generative adversarial network

    CN110120038A

  • Intelligent road inspection method and equipment based on multi-dimensional vision fusion

    CN119274030A