A highway pavement disease detection method and system based on image recognition

CN122676431APending Publication Date: 2026-09-01YIXIN (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610786268.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0002]在高速公路车载移动巡检作业中,巡检车辆通常以固定的速度在起伏变坡路段连续行驶,受底盘悬架动态形变、轮胎径向跳动及路面横坡变化影响,安装于车辆底部的车载成像模组会产生高频俯仰与横滚姿态波动,导致连续采集的路面原始图像序列出现严重的非线性透视畸变与动态视场偏移;现有基于图像识别的路面病害检测方法多采用固定相机标定参数对原始图像直接进行特征提取与像素级分割,未构建车辆底盘刚性结构参考点、相机光学中心与实时动态位姿偏角之间的空间几何映射与动态畸变补偿机制,致使图像中的微裂缝、坑槽等病害实体在投影成像过程中发生各向异性拉伸或压缩,病害边界特征与背景干扰发生空间混叠,且缺乏基于连续帧时空轨迹的视差一致性校验与物理尺寸精准反演逻辑,最终导致病害定位偏移率高、瞬态水雾与阴影易被误识别,难以输出满足高等级公路精细化养护决策要求的高置信度结构化检测报告

Benefits of technology

通过预设于巡检车辆底盘左侧轮拱内衬、右侧轮拱内衬及车载成像模组光学中心处的三维空间参考锚点,结合视场衰减梯度与车辆动态位姿偏角拟合构建虚拟自适应包络椭圆面,通过极坐标圆周状态映射轨迹与方位角连续旋转变换解算分量极值包络路径以生成像素位移修正量进行空间坐标系重映射与局部几何形变补偿,并将病害初步定位区块投影至桩号里程与横向偏移量基准的连续帧二维展开参考系执行跨帧轨迹关联与视差一致性校验,最终结合多尺度特征加权投票与双阈值判决融合处理的技术手段,所以克服了现有检测方法在车辆高频俯仰与横滚姿态波动下依赖固定相机标定参数导致的非线性透视畸变无法动态补偿、病害实体各向异性拉伸与背景干扰空间混叠、瞬态水雾与阴影伪影易误识别以及缺乏连续帧物理尺寸精准反演逻辑的技术问题,进而实现了高动态巡检工况下病害几何畸变的自适应精准校正、病害定位时空轨迹的稳定关联与抗干扰鲁棒性的提升,以及病害类型标签与实际物理尺寸的精确量化输出,最终达到了直接输出满足高等级公路精细化养护决策要求的高置信度结构化检测报告的技术效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676431A_ABST
    Figure CN122676431A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for detecting highway pavement defects based on image recognition, belonging to the field of intelligent operation and maintenance technology for road infrastructure. The method includes: acquiring an original pavement image sequence; inputting the original pavement image sequence into a preset multi-scale convolutional feature extraction architecture to obtain a defect candidate feature map; extracting three-dimensional spatial reference anchor points preset at the inner lining of the left and right wheel arches of the inspection vehicle chassis and the optical center of the onboard imaging module based on the defect candidate feature map; and constructing a virtual adaptive envelope ellipse by fitting the field-of-view attenuation gradient of the three-dimensional spatial reference anchor points with the vehicle's dynamic pose angle. This invention achieves adaptive geometric distortion compensation, cross-frame spatiotemporal stable correlation, and high-confidence structured data output for pavement defects under high-dynamic driving environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent operation and maintenance technology for road infrastructure, and in particular to a method and system for detecting highway pavement defects based on image recognition. Background Technology

[0002] In highway mobile inspection operations, inspection vehicles typically travel continuously at a fixed speed on undulating and sloped road sections. Affected by the dynamic deformation of the chassis suspension, radial runout of the tires, and changes in the cross slope of the road surface, the on-board imaging module installed at the bottom of the vehicle will generate high-frequency pitch and roll attitude fluctuations, resulting in severe nonlinear perspective distortion and dynamic field of view shift in the continuously acquired original road image sequence. Existing image recognition-based road defect detection methods mostly use fixed camera calibration parameters to directly extract features and perform pixel-level segmentation on the original images. They do not construct a spatial geometric mapping and dynamic distortion compensation mechanism between the rigid structural reference point of the vehicle chassis, the optical center of the camera, and the real-time dynamic pose angle. This causes the defect entities such as microcracks and potholes in the image to undergo anisotropic stretching or compression during the projection imaging process. The defect boundary features and background interference are spatially mixed. Furthermore, there is a lack of parallax consistency verification and accurate physical size inversion logic based on continuous frame spatiotemporal trajectory. Ultimately, this results in a high defect positioning offset rate, and transient water mist and shadows are easily misidentified, making it difficult to output a high-confidence structured inspection report that meets the requirements of refined maintenance decision-making for high-grade highways. Summary of the Invention

[0003] This invention provides a method and system for detecting road surface defects based on image recognition, which realizes adaptive geometric distortion compensation, cross-frame spatiotemporal stable correlation, and high-confidence structured data output for road surface defects under high dynamic driving environment.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a method for detecting highway pavement defects based on image recognition, the method comprising: Obtain the original road surface image sequence; input the original road surface image sequence into the preset multi-scale convolutional feature extraction architecture to obtain the disease candidate feature map; Based on the defect candidate feature mapping map, three-dimensional spatial reference anchor points are extracted from the inner lining of the left wheel arch, the inner lining of the right wheel arch, and the optical center of the vehicle imaging module on the chassis of the inspection vehicle. A virtual adaptive envelope ellipse is constructed by fitting the field attenuation gradient of the three-dimensional spatial reference anchor points with the dynamic pose angle of the vehicle. Orthogonal distortion gradient components are extracted from the virtual adaptive envelope ellipse, and polar coordinate circular state mapping trajectory is constructed. The extreme value envelope path of the components is solved by continuous azimuth angle rotation transformation. The distortion vector is decoupled and the principal distortion axis direction and orthogonal compensation amplitude are locked. Based on this, the geometric state vector orthogonal decomposition operation is performed to obtain the pixel displacement correction amount of the anisotropic projection distortion law. The pixel displacement correction is applied to the disease candidate feature map to perform spatial coordinate system remapping and local geometric deformation compensation. After pixel-level conditional random field semantic segmentation and polygon edge morphological closure reconstruction, a preliminary set of disease localization blocks is obtained. The preliminary location of the disease blocks is projected onto a continuous two-dimensional unfolded reference system based on the highway station mileage and lane lateral offset. Cross-frame block trajectory association and disparity consistency verification are performed to obtain the spatiotemporal confidence sequence of the disease. Based on the spatiotemporal confidence sequence of the pavement defects, a multi-scale feature weighted voting and dual-threshold decision fusion process is performed to obtain a structured pavement defect detection report.

[0005] Secondly, a highway pavement defect detection system based on image recognition includes: The acquisition module is used to acquire the original road surface image sequence; the original road surface image sequence is input into the preset multi-scale convolutional feature extraction architecture to obtain the disease candidate feature map; The module is used to extract three-dimensional spatial reference anchor points based on the defect candidate feature map, which are preset at the inner lining of the left wheel arch, the inner lining of the right wheel arch, and the optical center of the vehicle imaging module on the chassis of the inspection vehicle; and to construct a virtual adaptive envelope ellipse by fitting the field of view attenuation gradient of the three-dimensional spatial reference anchor points with the dynamic pose angle of the vehicle. The calculation module is used to extract orthogonal distortion gradient components from the virtual adaptive envelope ellipse, construct the polar coordinate circular state mapping trajectory, solve the extreme value envelope path of the components through continuous azimuth angle rotation transformation, decouple the distortion vector and lock the main distortion axis direction and orthogonal compensation amplitude, and perform geometric vector orthogonal decomposition operation accordingly to obtain the pixel displacement correction amount of the anisotropic projection distortion law. The compensation module is used to apply the pixel displacement correction amount to the disease candidate feature map to perform spatial coordinate system remapping and local geometric deformation compensation. After pixel-level conditional random field semantic segmentation and polygon edge morphological closure reconstruction, a set of preliminary disease location blocks is obtained. The verification module is used to project the set of preliminary location blocks of the defects onto a continuous frame two-dimensional unfolded reference system based on the highway station mileage and lane lateral offset, and to perform cross-frame block trajectory association and disparity consistency verification to obtain the spatiotemporal confidence sequence of the defects. The processing module is used to perform multi-scale feature weighted voting and dual-threshold decision fusion processing based on the spatiotemporal confidence sequence of the pavement defects to obtain a structured pavement defect detection report.

[0006] Thirdly, a computing device, comprising: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0007] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0008] The above-described solution of the present invention has at least the following beneficial effects: By pre-setting three-dimensional spatial reference anchor points at the inner linings of the left and right wheel arches of the inspection vehicle chassis and the optical center of the onboard imaging module, and combining the field-of-view attenuation gradient with the vehicle's dynamic pose angle fitting, a virtual adaptive envelope ellipse is constructed. The extreme value envelope path of the components is calculated through polar coordinate circular state mapping trajectory and continuous azimuth angle rotation transformation to generate pixel displacement correction for spatial coordinate system remapping and local geometric deformation compensation. The initial defect location block is projected onto a continuous frame two-dimensional unfolded reference system based on the station mileage and lateral offset benchmarks to perform cross-frame trajectory association and disparity consistency verification. Finally, a multi-scale feature weighted voting and dual-threshold decision fusion processing technique is used. This invention overcomes the technical problems of existing detection methods, such as the inability to dynamically compensate for nonlinear perspective distortion caused by relying on fixed camera calibration parameters under high-frequency pitch and roll attitude fluctuations of vehicles, the anisotropic stretching of the defect entity and spatial aliasing of background interference, the easy misidentification of transient water mist and shadow artifacts, and the lack of accurate inversion logic for physical dimensions of continuous frames. It achieves adaptive and accurate correction of defect geometric distortion under high-dynamic inspection conditions, stable correlation of defect location spatiotemporal trajectory and improved anti-interference robustness, as well as accurate quantitative output of defect type labels and actual physical dimensions. Ultimately, it achieves the technical effect of directly outputting high-confidence structured inspection reports that meet the requirements of refined maintenance decision-making for high-grade highways. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating a method for detecting highway pavement defects based on image recognition, provided by an embodiment of the present invention.

[0010] Figure 2 This is a schematic diagram of a highway pavement defect detection system based on image recognition, provided by an embodiment of the present invention. Detailed Implementation

[0011] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art.

[0012] like Figure 1 As shown in the figure, an embodiment of the present invention proposes a method for detecting highway pavement defects based on image recognition. The method includes the following steps: Obtain the original road surface image sequence; input the original road surface image sequence into the preset multi-scale convolutional feature extraction architecture to obtain the disease candidate feature map; Based on the defect candidate feature mapping map, three-dimensional spatial reference anchor points are extracted from the inner lining of the left wheel arch, the inner lining of the right wheel arch, and the optical center of the vehicle imaging module on the chassis of the inspection vehicle. A virtual adaptive envelope ellipse is constructed by fitting the field attenuation gradient of the three-dimensional spatial reference anchor points with the dynamic pose angle of the vehicle. Orthogonal distortion gradient components are extracted from the virtual adaptive envelope ellipse, and polar coordinate circular state mapping trajectory is constructed. The extreme value envelope path of the components is solved by continuous azimuth angle rotation transformation. The distortion vector is decoupled and the principal distortion axis direction and orthogonal compensation amplitude are locked. Based on this, the geometric state vector orthogonal decomposition operation is performed to obtain the pixel displacement correction amount of the anisotropic projection distortion law. The pixel displacement correction is applied to the disease candidate feature map to perform spatial coordinate system remapping and local geometric deformation compensation. After pixel-level conditional random field semantic segmentation and polygon edge morphological closure reconstruction, a preliminary set of disease localization blocks is obtained. The preliminary location of the disease blocks is projected onto a continuous two-dimensional unfolded reference system based on the highway station mileage and lane lateral offset. Cross-frame block trajectory association and disparity consistency verification are performed to obtain the spatiotemporal confidence sequence of the disease. Based on the spatiotemporal confidence sequence of the pavement defects, a multi-scale feature weighted voting and dual-threshold decision fusion process is performed to obtain a structured pavement defect detection report.

[0013] In this embodiment of the invention, based on three-dimensional spatial reference anchor points preset at the inner linings of the left and right wheel arches of the inspection vehicle chassis and the optical center of the vehicle-mounted imaging module, a virtual adaptive envelope ellipse is constructed by combining the field-of-view attenuation gradient and the vehicle's dynamic pose angle fitting. A polar coordinate circular state mapping trajectory is constructed by extracting orthogonal distortion gradient components, and the extreme value envelope path of the components is solved using continuous azimuth angle rotation transformation to decouple the distortion vector and lock the principal distortion axis direction and orthogonal compensation amplitude. Based on this, a geometrical vector orthogonal decomposition operation is performed to obtain the pixel displacement correction amount. This correction amount is applied to the defect candidate feature mapping map for spatial coordinate system remapping and local geometric deformation compensation. After pixel-level conditional random field semantic segmentation and polygon edge morphological closure reconstruction, a preliminary defect location block set is obtained, and this set is projected onto a... The method uses a continuous frame two-dimensional unfolded reference system based on highway mileage and lane lateral offset to perform cross-frame block trajectory association and disparity consistency verification to obtain a spatiotemporal confidence sequence of defects. Finally, based on this sequence, a multi-scale feature weighted voting and dual-threshold decision fusion processing technique is performed. This overcomes the technical problems of existing detection methods under high-speed dynamic inspection conditions, such as the inability to dynamically compensate for nonlinear perspective distortion caused by high-frequency fluctuations in vehicle pose, spatial mixing of anisotropic stretching of defect entities and background interference, easy misidentification of transient light and shadow artifacts, and the lack of cross-frame spatiotemporal continuity verification and accurate physical size inversion mechanism. This achieves adaptive and accurate correction of geometric distortion in high dynamic imaging, stable association of spatiotemporal trajectory of defect location and improved anti-interference robustness, and accurate quantitative matching of defect type label and actual physical size.

[0014] In a preferred embodiment of the present invention, step 1 above may include: Step 1.1 involves performing dynamic illumination compensation and motion blur deconvolution processing on consecutive image frames in the original road surface image sequence to eliminate local overexposure and motion blur interference under high-speed driving conditions, resulting in a time-aligned standardized road surface image sequence. Specifically, this includes: analyzing the acquired original road surface image sequence frame by frame to identify abnormal illumination regions in each frame. These regions are mainly caused by local overexposure or underexposure due to tree shade, backlighting, or strong light reflection during high-speed driving. The pixel grayscale values ​​in the locally overexposed regions are close to 255, while the pixel grayscale values ​​in the locally underexposed regions are close to 0. For these illumination anomalies, a dynamic illumination compensation algorithm is used to process them. The overall grayscale mean and grayscale variance of each frame are calculated, and the illumination compensation coefficient is determined based on the grayscale mean and variance. The pixel grayscale values ​​in the abnormal illumination regions are adaptively adjusted, specifically by reducing the pixel grayscale values ​​in overexposed regions and increasing the pixel grayscale values ​​in underexposed regions, while preserving the detail information in the normally illuminated regions of the image to avoid image distortion caused by overcompensation.

[0015] To address motion blur interference caused by high-speed driving, motion blur deconvolution processing is performed. The characteristics of motion blur are analyzed, revealing that it is primarily linear in high-speed inspection, with the direction of motion aligned with the vehicle's direction of travel. The length of the motion blur is related to the vehicle's speed and the frame rate of the onboard imaging module. The length is calculated by converting the vehicle speed to pixels per second and multiplying it by the time interval between two adjacent frames. Based on this analysis, the motion blur kernel parameters for each frame are determined through pixel displacement matching between adjacent frames, including the blur direction and blur length. Then, a deconvolution algorithm is applied. Based on a defined blur kernel, a reverse operation is performed on the blurred image to gradually restore the road surface details obscured by motion blur and eliminate motion blur interference. , These are the two-dimensional Fourier transform and the inverse Fourier transform, respectively. For fuzzy kernel Fourier transform; for The complex conjugate; Noise power spectrum; This is the noise smoothing coefficient, used to balance restoration accuracy and noise suppression.

[0016] Temporal alignment is performed on consecutive image frames after dynamic illumination compensation and motion blur deconvolution processing. By extracting road feature points in each frame, such as lane line edges and manhole cover edges, the displacement of feature points between adjacent frames is calculated. Based on the displacement, the image frames are translated and rotated for fine-tuning to ensure that the road area positions in consecutive image frames are aligned, avoiding inter-frame road misalignment caused by slight vehicle displacement. The final result is a standardized road image sequence that is temporally aligned, uniformly illuminated, and without obvious ghosting.

[0017] Step 1.2 involves superimposing the standardized road surface image sequence with inter-frame features according to a preset sliding step size to construct a multi-dimensional channel feature tensor. This multi-dimensional channel feature tensor is then injected into the bottom convolutional branch of the multi-scale convolutional feature extraction architecture. Specifically, this includes: setting a preset sliding step size, which needs to be combined with the frame rate of the vehicle imaging module and the vehicle's driving speed to ensure that there are overlapping road surface features between adjacent superimposed frames without generating too much redundant information. Typically, the sliding step size is set to 1 or 2 frames, meaning that feature superposition is performed once for every 1 or 2 adjacent images. The specific process of inter-frame feature superposition is as follows: according to the preset sliding step size, consecutive image frames are selected sequentially for feature fusion. For example, when the sliding step size is 1 frame, the 1st frame is superimposed with the 2nd frame, the 2nd frame with the 3rd frame, the 3rd frame with the 4th frame, and so on. The superposition method is a weighted summation of the gray values ​​of corresponding pixels. The weights are allocated according to the clarity of the two images. The image frame with higher clarity has a larger weight, thereby strengthening the continuity and integrity of the road surface features and reducing the impact of noise in a single image.

[0018] After completing the inter-frame feature overlay, a multi-dimensional channel feature tensor is constructed. The multi-dimensional channels are designed to meet the needs of pavement distress detection, primarily including grayscale feature channels, texture feature channels, and edge feature channels. The grayscale feature channel reflects the brightness differences of pavement pixels, the texture feature channel reflects the discrete texture information of the asphalt mixture surface, and the edge feature channel reflects the edge contour information of pavement micro-cracks, potholes, and other distresses. The three types of features are extracted from each overlaid image and integrated to form a multi-dimensional channel feature tensor. This tensor contains both local features of a single frame and continuous features from adjacent frames, comprehensively reflecting the overall state of the pavement. The constructed multi-dimensional channel feature tensor is then injected into the bottom-level convolutional branch of a pre-defined multi-scale convolutional feature extraction architecture.

[0019] Step 1.3: In the multi-scale convolutional feature extraction architecture, road surface texture response signals at different receptive field scales are captured through a parallel dilated convolutional structure. Cross-layer feature pooling and channel weight recalibration operations are performed to aggregate the discrete features of the asphalt mixture surface layer and the response information of microcrack edges to obtain a candidate feature map of the road surface. Specifically, this includes: building a parallel dilated convolutional structure, which contains multiple parallel convolutional branches. Each branch is set with a different void ratio. The value of the void ratio is determined according to the scale of the road surface defects. Typically, three parallel branches are set with void ratios of 1, 3, and 5. The branch with a void ratio of 1 is used to capture fine road surface texture features, such as the discrete distribution of asphalt particles; the branch with a void ratio of 3 is used to capture medium-scale road surface features, such as the edges of microcracks; and the branch with a void ratio of 5 is used to capture larger-scale road surface features, such as the outline of potholes. By simultaneously operating multiple parallel dilated convolution branches, road surface texture response signals at different receptive field scales are captured, ensuring that no damage features at different scales are missed. After processing by the parallel dilated convolution structure, different convolution branches will output feature maps at different scales. The purpose of cross-layer feature pooling is to fuse these feature maps at different scales and retain the key information of features at each scale. Specifically, the feature map output by each branch is pooled using max pooling, selecting the maximum pixel value within each pooling window as the output value of that window to reduce the dimensionality of the feature map while preserving the saliency of the features. The feature maps after pooling from different branches are then concatenated according to the channel dimension to obtain the cross-layer fused feature map, thus integrating road surface features at different scales.

[0020] Because different channels have varying importance for disease detection—for example, edge feature channels are far more effective than background texture channels in identifying microcracks and potholes—it is necessary to recalibrate the weights of each channel in the fused feature map to enhance useful features and suppress useless background features. Specifically, the global mean and global variance of each feature channel are calculated. Based on the mean and variance, the importance weight of each channel is determined, assigning higher weights to high-importance channels such as edge feature channels and lower weights to low-importance channels such as background texture channels. The feature map of each channel is multiplied by its corresponding weight to obtain the recalibrated feature map. The recalibrated feature map then undergoes feature aggregation processing, focusing on aggregating discrete features of the asphalt mixture surface and edge response information of microcracks. This integrates scattered disease-related features into continuous feature regions, removes background interference features, and finally yields a disease candidate feature map that clearly reflects the candidate areas of pavement diseases.

[0021] In this embodiment of the invention, by performing dynamic illumination compensation and motion blur deconvolution processing on continuous image frames, constructing multi-dimensional channel feature tensors by superimposing inter-frame features according to a preset sliding step size and injecting them into the bottom convolutional branch, and capturing road surface texture response signals at different receptive field scales through parallel dilated convolutional structures and performing cross-layer feature pooling and channel weight recalibration operations, the technical problems of severe local overexposure and motion blur interference under high-speed driving conditions, easy omission of minor defects and background texture confusion due to single-scale feature extraction, and insufficient feature space expression are overcome. Thus, the invention achieves efficient aggregation and accurate capture of multi-scale road surface discrete features and micro-crack edge response information by eliminating dynamic imaging environment and motion interference.

[0022] In a preferred embodiment of the present invention, step 2 above may include: Step 2.1: By analyzing the pixel spatial distribution matrix of the defect candidate feature map, and combining it with the preset camera intrinsic parameter calibration matrix and the vehicle chassis rigid mounting coordinate system, the spatial coordinates of the three-dimensional spatial reference anchor points at the inner lining of the left and right wheel arches of the inspection vehicle chassis and the optical center of the vehicle imaging module are located. Specifically, this includes: analyzing the pixel spatial distribution matrix of the obtained defect candidate feature map. The pixel spatial distribution matrix is ​​a matrix composed of the coordinate information, grayscale feature information and defect candidate feature information of all pixels in the map. By analyzing this matrix, the imaging spatial position corresponding to each pixel in the map can be clearly determined. The system calls upon a preset camera intrinsic parameter calibration matrix, which contains core parameters such as the camera's focal length, pixel size, and principal point coordinates. This matrix is ​​crucial for converting image pixel coordinates into three-dimensional spatial coordinates. These parameters are obtained by pre-calibrating the vehicle-mounted imaging module using a standard calibration board before the inspection operation, ensuring the accuracy and stability of the parameters. Simultaneously, the system calls upon the vehicle chassis rigid mounting coordinate system. This coordinate system uses the vehicle chassis geometric center as the origin, the vehicle's driving direction as the longitudinal axis, the lane width direction as the lateral axis, and the direction perpendicular to the road surface as the vertical axis. This system is used to unify the spatial coordinate reference between the chassis components and the imaging module.

[0023] The spatial coordinates of three anchor points are determined separately: For the left and right wheel arch liners of the inspection vehicle chassis, their relative positions in the rigid mounting coordinate system of the vehicle chassis are fixed. Their initial coordinates in this coordinate system are determined by the chassis design parameters. Combined with the camera intrinsic parameter calibration matrix, these initial coordinates are mapped to the pixel coordinates of the defect candidate feature mapping map. The coordinate deviation is adjusted by reverse verification through the pixel spatial distribution matrix to determine the three-dimensional spatial coordinates of the two wheel arch liner anchor points. For the optical center of the vehicle imaging module, its position in the rigid mounting coordinate system of the vehicle chassis is determined by the installation parameters of the imaging module. Combined with the camera intrinsic parameter calibration matrix, its installation position coordinates are converted into three-dimensional spatial coordinates to ensure that these coordinates are completely consistent with the actual optical center position of the imaging module. The three-dimensional spatial coordinates of the three anchor points are integrated to form a complete set of three-dimensional spatial reference anchor points.

[0024] Step 2.2: Based on the spatial coordinates of the three-dimensional reference anchor points, calculate the illumination attenuation gradient in the edge region of the continuous imaging frames. Simultaneously acquire pitch, roll, and yaw angle deflection data output by the vehicle's inertial measurement unit, and synthesize the vehicle's dynamic pose deflection vector representing the vehicle's attitude fluctuations. Specifically, this includes: determining the field of view of the continuous imaging frames based on the obtained three-dimensional reference anchor point coordinates. The field of view is defined with the optical center of the vehicle imaging module as the vertex and the projections of the left and right wheel arch liner anchor points onto the imaging plane as boundaries, delineating the effective imaging area for each frame; focusing on analyzing the illumination changes in the edge region of the continuous imaging frames, selecting the appropriate... The pixel region at the edge of the field of view is typically a width of 10 to 20 pixels on each of the top, bottom, left, and right edges of the imaging frame. The illuminance value of each pixel in this region is calculated. The illuminance value is obtained by converting the pixel's grayscale value. The higher the grayscale value, the stronger the illuminance. The field of view attenuation gradient is calculated. Specifically, the illuminance values ​​of adjacent pixels in the edge region are selected, the illuminance difference between two adjacent pixels is calculated, and then divided by the spatial distance between the two pixels to obtain the rate of change of illuminance in a single direction. The rates of change of illuminance in the horizontal and vertical directions are calculated separately, and then integrated to obtain the field of view attenuation gradient of the edge region of the continuous imaging frame. This gradient reflects the attenuation law of illumination at the edge of the field of view.

[0025] The attitude data output from the vehicle's inertial measurement unit (IMU) is collected synchronously. The IMU is the core component for real-time detection of the vehicle's attitude, outputting pitch, roll, and yaw angle data during vehicle movement. Pitch angle reflects the vehicle's forward and backward tilt angle, roll angle reflects the vehicle's left and right tilt angle, and yaw angle reflects the vehicle's deviation from the preset route. During acquisition, data is collected at the same time interval as the imaging frames to ensure a one-to-one correspondence between attitude data and imaging frames, avoiding parameter mismatches caused by time differences. The collected pitch, roll, and yaw angle data are vector-synthesized to obtain the vehicle's dynamic attitude deviation vector. The synthesis process is as follows: using the rigid mounting coordinate system of the vehicle chassis as a reference, the deviation data of the three angles are converted into corresponding vector components, where pitch angle corresponds to the vertical vector component, roll angle corresponds to the lateral vector component, and yaw angle corresponds to the longitudinal vector component. The three vector components are superimposed to obtain the vehicle's dynamic attitude deviation vector that comprehensively reflects the fluctuations in the vehicle's attitude.

[0026] Step 2.3: Map the field-of-view attenuation gradient and the vehicle's dynamic pose angle vector to the 3D imaging projection space. Use nonlinear surface interpolation iterative calculations to fit a virtual adaptive envelope ellipse with continuously transitioning boundary curvature. This allows the principal axis direction and major and minor axis radii of the virtual adaptive envelope ellipse to dynamically adapt to the spatial deformation trend of the road imaging field of view, completing the spatial surface construction of the virtual adaptive envelope ellipse. Specifically, this includes mapping the obtained field-of-view attenuation gradient and the vehicle's dynamic pose angle vector to the 3D imaging projection space. This space is a 3D space connecting the optical center of the vehicle imaging module and the road imaging area, realistically reflecting the projection imaging process of the road image. Let the 3D space reference anchor points be the inner lining anchor points of the left wheel arch of the chassis. Right wheel arch inner lining anchor point Optical center anchor point of vehicle imaging module During the mapping process, using this set of three-dimensional space reference anchor points as a benchmark, the field-of-view attenuation gradient is converted into a lighting change vector in three-dimensional space. The vehicle's dynamic pose angle vector is converted into an attitude offset vector in three-dimensional space. Satisfying coordinate uniformity constraints ,in This is a 3D imaging projection space coordinate system, ensuring that the two parameters work together in the same space coordinate system.

[0027] Elliptical surface fitting is performed using nonlinear surface interpolation iterative computation. The core purpose of this computation is to construct a virtual adaptive envelope elliptical surface with continuously transitioning boundary curvature, ensuring that the elliptical surface dynamically adapts to the spatial deformation trend of the road imaging field of view. The specific computation process is as follows: using three three-dimensional spatial reference anchor points... Based on this, a uniform grid of nodes is divided in the three-dimensional imaging projection space to obtain the grid node set. ,in Each node is indexed into a 3D spatial grid and its corresponding node is calculated by traversing the grid. Corresponding field attenuation gradient value With vehicle dynamic pose deflection vector value A nonlinear interpolation method is employed, based on the parameter values ​​of adjacent grid nodes, through an interpolation function. Calculate the parameter values ​​of the intermediate grid nodes. , fill the gaps in the grid, among which For adjacent known mesh nodes, the parameters of the mesh nodes are continuously adjusted through iterative calculations, and a boundary curvature continuity constraint function is introduced. The iterative process satisfies ,in For the number of iterations, To allow for continuous curvature error, the curvature of the fitted surface boundary is made continuous without abrupt changes, gradually forming the initial outline of the elliptical surface.

[0028] During the fitting process, the key is to ensure that the principal axis direction of the virtual adaptive envelope ellipse dynamically adapts to the spatial deformation trend of the road imaging field of view with respect to the major and minor axis radii: let the initial principal axis tilt angle of the ellipse be... The vehicle pitch attitude fluctuation angle is Then, the spindle tilt angle after dynamic adjustment satisfies This ensures that the principal axis of the elliptical surface always aligns with the central axis of the road surface imaging field of view; let the initial major axis radius of the elliptical surface be... The initial minor axis radius is The vehicle roll attitude fluctuation adaptation coefficient is The road surface cross slope variation adaptation coefficient is: Then the major axis radius is dynamically adjusted. minor axis radius To ensure the elliptical surface completely covers the road imaging field of view, while conforming to the spatial deformation law of the field of view, it avoids omissions in the field of view or mismatch between the surface and the field of view. After multiple rounds of nonlinear surface interpolation iteration calculations, when the boundary curvature of the elliptical surface is continuous, the principal axis direction and the radii of the major and minor axes stably adapt to the spatial deformation trend of the road imaging field of view, and the elliptical surface completely covers the entire road imaging area, the iteration calculation stops, and the spatial surface construction of the virtual adaptive envelope elliptical surface is completed.

[0029] In this embodiment of the invention, the pixel spatial distribution matrix of the defect candidate feature map is analyzed and combined with the preset camera intrinsic parameter calibration matrix and the vehicle chassis rigid mounting coordinate system to accurately locate the spatial coordinates of the three-dimensional spatial reference anchor point. Based on this coordinate, the field of view attenuation gradient of the edge region of the continuous imaging frame is calculated. The pitch angle, roll angle and yaw angle deflection data output by the vehicle inertial measurement unit are collected simultaneously to synthesize the vehicle dynamic pose deflection vector. The field of view attenuation gradient and the vehicle dynamic pose deflection vector are mapped to the three-dimensional imaging projection space and a virtual adaptive envelope with continuous transition of boundary curvature is fitted by nonlinear surface interpolation iterative calculation. The elliptical surface is used to dynamically adapt the principal axis direction and the radii of the major and minor axes to the spatial deformation trend of the road imaging field of view. This overcomes the technical problems of existing detection methods that rely on fixed camera calibration parameters under high-speed dynamic inspection conditions, cannot perceive the nonlinear spatial distortion caused by the high-frequency deformation and pose fluctuation of the vehicle chassis suspension on the imaging field of view in real time, and lack of dynamic geometric mapping reference, which leads to the inaccuracy of the distortion compensation model. In this way, a dynamic three-dimensional geometric reference reference that is highly coupled with the real-time running posture and optical imaging characteristics of the vehicle is constructed, so that the virtual surface can adaptively track and accurately fit the continuous spatial deformation trend of the imaging field of view.

[0030] In a preferred embodiment of the present invention, step 3 above may include: Step 3.1: On the 3D surface mesh of the virtual adaptive envelope ellipse, collect surface curvature change rate data along the radial and tangential orthogonal directions, and extract orthogonal distortion gradient components. Specifically, this includes: calling the constructed virtual adaptive envelope ellipse, which has achieved continuous boundary curvature transition through nonlinear interpolation iteration calculations and can dynamically adapt to the spatial deformation of the road imaging field of view; defining the 3D surface mesh structure of the ellipse, which consists of uniformly distributed mesh nodes, each corresponding to a spatial point on the ellipse. The mesh density is set according to the imaging accuracy requirements to ensure accurate capture of surface curvature changes; typically, the mesh node spacing is set to a spatial distance corresponding to 1 to 2 pixels; and collecting surface curvature change rate data for each mesh node along the radial and tangential orthogonal directions of the 3D surface mesh. Among them, radial direction refers to the direction from the geometric center of the virtual adaptive envelope ellipse to the grid node. This direction corresponds to the depth direction of the road surface imaging and reflects the distortion change from the center of the imaging field of view to the edge. Tangential direction refers to the direction perpendicular to the radial direction and along the circumference of the ellipse. This direction corresponds to the width direction of the road surface imaging and reflects the distortion change in the lateral direction of the imaging field of view.

[0031] The process of collecting curvature change rate data is as follows: For each grid node, the curvature values ​​of its two adjacent grid nodes in the radial and tangential directions are calculated. The curvature difference between the two adjacent nodes is divided by the spatial distance between the two nodes to obtain the curvature change rate of that node in the corresponding direction. This process is repeated for all grid nodes in both the radial and tangential directions to collect the curvature change rate data, forming a complete curvature change rate dataset. Based on the collected radial and tangential curvature change rate data, orthogonal distortion gradient components are extracted. Specifically, the radial curvature change rate is used as one gradient component, and the tangential curvature change rate is used as another orthogonal gradient component. The two components are perpendicular to each other and do not interfere with each other, together forming the orthogonal distortion gradient component.

[0032] Step 3.2: Based on the orthogonal distortion gradient components, using the geometric center of the virtual adaptive envelope ellipse as the origin of polar coordinates, map the gradient magnitude and phase angle of each grid node to polar coordinate space to construct a polar coordinate circular state mapping trajectory. Specifically, this includes: determining the origin of polar coordinates, using the geometric center of the virtual adaptive envelope ellipse as the origin. This geometric center is obtained by calculating the mean spatial coordinates of all grid nodes on the ellipse, ensuring it is located at the center of the ellipse, serving as the reference point for polar coordinate mapping; and extracting the gradient magnitude and phase angle of the orthogonal distortion gradient components corresponding to each grid node. The gradient magnitude refers to the magnitude of the orthogonal distortion gradient component, obtained by calculating the square root of the sum of the squares of the radial and tangential gradient components, reflecting the distortion intensity at that grid node. The phase angle refers to the angle between the orthogonal distortion gradient component and the polar coordinate horizontal axis, with the direction of vehicle travel from the geometric center of the ellipse as the positive direction of the horizontal axis. This is obtained by calculating the arctangent of the tangential and radial gradient components, reflecting the direction of distortion.

[0033] The gradient magnitude and phase angle of each grid node are mapped to polar coordinate space. The mapping process is as follows: taking the origin of polar coordinates as the reference, the phase angle is used as the polar angle of polar coordinates, and the gradient magnitude is used as the polar radius of polar coordinates. Each grid node corresponds to a point in polar coordinate space. The polar angle determines the direction of the point, and the polar radius determines the distance of the point from the origin. After mapping all grid nodes to polar coordinate space, the corresponding polar coordinate points are connected sequentially in order of increasing phase angle to form a continuous circular trajectory, i.e., the polar coordinate circular state mapping trajectory.

[0034] Step 3.3: Perform a continuous azimuth rotation transformation with a preset angular step size along the polar coordinate circular state mapping trajectory. Track the peak values ​​of the gradient component responses under each rotation phase and fit a smooth and continuous component extremum envelope path. Specifically, this includes: setting the preset azimuth rotation transformation angular step size. The angular step size needs to balance computational efficiency and detection accuracy, and is usually set to 1 to 2 degrees. The smaller the angular step size, the higher the detection accuracy, but the greater the computational load. Considering the real-time requirements of high-speed inspection, 1 degree is preferentially selected as the preset angular step size. Perform a continuous azimuth rotation transformation along the constructed polar coordinate circular state mapping trajectory. The rotation transformation is centered on the polar coordinate origin. According to the preset angular step size, starting from 0 degrees, the azimuth angle is gradually increased until a 360-degree rotation is completed. Each rotation step size yields a rotation phase, and each rotation phase corresponds to a direction in polar coordinate space.

[0035] At each rotation phase, the peak value of the gradient component response in that direction is tracked. Specifically, the tracking method is as follows: in the polar angle direction corresponding to the current rotation phase, the gradient magnitudes of all polar coordinate points in that direction are selected. All magnitudes are compared, and the largest magnitude is selected as the peak value of the gradient component response at that rotation phase. This process is repeated for all rotation phases to obtain a set of peak data, with each peak corresponding to one rotation phase. The obtained peak data is then fitted using a smooth curve fitting algorithm to connect all peak points into a smooth and continuous curve, eliminating random fluctuations and noise interference in the peak data, thus obtaining a smooth and continuous component extremum envelope path.

[0036] Step 3.4: Based on the phase difference characteristics of the component extreme value envelope path, the original coupled distortion vector is separated by spatial orthogonal projection to decouple the cross-interference components and lock the principal distortion axis direction and orthogonal compensation amplitude. Specifically, this includes: analyzing the phase difference characteristics of the obtained component extreme value envelope path; the phase difference characteristics refer to the phase angle difference between two adjacent peak points on the component extreme value envelope path. By calculating the phase angle difference between adjacent peak points, the phase difference data is obtained. This data reflects the correlation of distortion in different directions, and the phase difference corresponding to the principal distortion direction will show a regular distribution; based on the phase difference characteristics, the original coupled distortion vector is separated by spatial orthogonal projection. The original coupled distortion vector is a mixed vector formed by the combined effects of multiple factors such as nonlinear perspective distortion and dynamic field of view shift during road imaging. It includes the principal distortion component, which is mainly caused by perspective distortion due to vehicle attitude fluctuations, and the cross-interference component, which is mainly caused by illumination changes and road texture interference, as secondary distortions. The specific process of spatial orthogonal projection separation is as follows: taking the peak direction of the component extreme value envelope path as the reference, two mutually orthogonal projection axes are constructed, and the original coupled distortion vector is projected onto the two orthogonal axes respectively. One axis corresponds to the main distortion direction, and the other axis corresponds to the interference direction, thereby realizing the separation of the main distortion component and the cross interference component.

[0037] To decouple cross-interference components, an interference threshold is set. Components projected onto the interference direction with amplitudes below the threshold are identified as cross-interference components and removed. The main distortion component projected onto the main distortion direction is retained, thus decoupling the coupled distortion vector. The main distortion axis direction and orthogonal compensation amplitude are then locked. The main distortion axis direction refers to the projection axis direction corresponding to the main distortion component, i.e., the direction with the largest peak value and most regular phase difference in the component's extreme value envelope path. This direction corresponds to the most significant distortion direction in road surface imaging. The orthogonal compensation amplitude refers to the maximum or minimum value of the main distortion component along the main distortion axis direction, reflecting the severity of the main distortion.

[0038] Step 3.5: Using the principal distortion axis as the reference coordinate axis, perform geometric vector orthogonal decomposition on the orthogonal distortion gradient components to solve the surface distortion response into lateral and longitudinal offset components in the two-dimensional image plane, and synthesize the pixel displacement correction amount of the anisotropic projection distortion law. Specifically, this includes: constructing a coordinate system of the two-dimensional image plane using the locked principal distortion axis as the reference coordinate axis. This coordinate system is consistent with the coordinate system of the road surface imaging image. The lateral direction corresponds to the width direction of the image, and the longitudinal direction corresponds to the height direction of the image, ensuring that the decomposed offset components are directly applied to image correction; and performing geometric vector orthogonal decomposition on the extracted orthogonal distortion gradient components. The specific process of orthogonal decomposition is as follows: the orthogonal distortion gradient component of each grid node is decomposed into the principal distortion axis direction of the reference coordinate axis and the orthogonal direction perpendicular to the reference coordinate axis, resulting in gradient components in two directions; then the gradient components in these two directions are mapped to the horizontal and vertical directions of the two-dimensional image plane, respectively, where the gradient component corresponding to the principal distortion axis direction is mapped as the vertical offset component, and the gradient component corresponding to the orthogonal direction is mapped as the horizontal offset component.

[0039] The surface distortion response is solved as lateral and longitudinal offset components within the two-dimensional image plane. The calculation process is as follows: Based on the gradient component magnitude of each grid node and the projection ratio between the virtual adaptive envelope ellipse and the two-dimensional image plane, the pixel offset distance corresponding to each grid node is calculated. The lateral offset component corresponds to the pixel's offset distance in the image width direction, and the longitudinal offset component corresponds to the pixel's offset distance in the image height direction. The positive and negative signs of the offset distance represent the offset direction: positive directions are the right and bottom sides of the image, and negative directions are the left and top sides. The lateral and longitudinal offset components corresponding to all grid nodes are integrated to synthesize the pixel displacement correction amount for anisotropic projection distortion. This correction amount includes the lateral and longitudinal offset parameters of each pixel in the image, accurately matching the anisotropic stretching or compression distortion of the road surface image.

[0040] In this embodiment of the invention, orthogonal distortion gradient components are extracted by collecting curvature change rate data along radial and tangential orthogonal directions on a three-dimensional surface mesh of a virtual adaptive envelope ellipse. A polar coordinate circular state mapping trajectory is constructed with the geometric center as the origin of polar coordinates. The gradient response peaks under each rotation phase are tracked by continuous azimuth rotation transformation with a preset angular step size, and a smooth and continuous component extreme value envelope path is fitted. Spatial orthogonal projection separation is performed using phase difference characteristics to decouple the coupled distortion vector and lock the principal distortion axis direction and orthogonal compensation amplitude. Then, geometrical vector orthogonal decomposition operation is performed with the principal distortion axis direction as the reference coordinate axis to solve the surface distortion response. This technique calculates the horizontal and vertical offset components within the two-dimensional image plane and synthesizes them into pixel displacement correction values. Therefore, it overcomes the technical problems of the high coupling of anisotropic projection distortion vectors under complex dynamic imaging conditions, which makes direct quantization difficult; the inability of traditional single-direction compensation strategies to effectively separate cross-interference components, leading to inaccurate local geometric deformation correction; and the lack of continuous envelope tracking in polar coordinate space, which results in directional blind spots and phase jumps in the calculation of compensation parameters. As a result, it achieves the precise decoupling of multi-dimensional coupled spatial surface distortion into quantifiable compensation components in independent orthogonal directions, realizing accurate locking of the principal axis direction of the main distortion and high-precision calculation of the orthogonal compensation amplitude.

[0041] In a preferred embodiment of the present invention, step 4 above may include: Step 4.1: Superimpose the horizontal and vertical offset components of the pixel displacement correction onto the original pixel grid coordinates of the disease candidate feature map to construct a spatial coordinate system remapping transformation matrix. Specifically, this includes: calling the obtained pixel displacement correction, clarifying that the correction contains two core components, namely the horizontal offset component and the vertical offset component. The horizontal offset component corresponds to the horizontal offset distance of the pixel in the width direction of the disease candidate feature map, and the vertical offset component corresponds to the vertical offset distance of the pixel in the height direction. The positive and negative values ​​of the two components represent the offset direction, with positive values ​​corresponding to the right and bottom of the image and negative values ​​corresponding to the left and top of the image. Extract the original pixel grid coordinates of the disease candidate feature map. The original pixel grid coordinates are composed of the position information of all pixels in the map. Each pixel corresponds to a unique two-dimensional coordinate. The horizontal coordinate represents the position of the pixel in the width direction of the image, and the vertical coordinate represents the position of the pixel in the height direction of the image. The coordinates of all pixels constitute a complete original pixel grid coordinate system.

[0042] The horizontal and vertical offset components of the pixel displacement correction are superimposed onto the corresponding original pixel grid coordinates. Specifically, for each pixel's original horizontal coordinate, the horizontal offset component is added to obtain the corrected horizontal coordinate; for each pixel's original vertical coordinate, the vertical offset component is added to obtain the corrected vertical coordinate. In this way, the mapping of all pixels from distorted coordinates to normal coordinates is completed. Based on the original and corrected coordinates of all pixels, a spatial coordinate system remapping transformation matrix is ​​constructed.

[0043] Step 4.2: Based on the spatial coordinate system remapping transformation matrix, perform bilinear interpolation resampling and local geometric deformation compensation operations on the disease candidate feature map to eliminate feature stretching distortion caused by surface projection and obtain a geometrically corrected feature map. Specifically, this includes: calling the constructed spatial coordinate system remapping transformation matrix; based on this matrix, determining the original pixel position corresponding to each corrected pixel; since the coordinates of the corrected pixels may not be integers and cannot directly correspond to the integer coordinates of the original pixels, interpolation is needed to calculate the pixel value corresponding to that position; bilinear interpolation resampling... Sampling is the preferred method that balances computational efficiency and correction accuracy. The specific process of bilinear interpolation resampling is as follows: For the coordinates of each corrected pixel, find the coordinates of its four neighboring original integer pixels, calculate the distance from the corrected pixel to these four neighboring pixels, and assign different weights according to the distance, with the closer the distance, the greater the weight; multiply the gray values ​​of the four neighboring pixels by the corresponding weights, and then add the products to obtain the gray value of the corrected pixel. The gray value calculation of all corrected pixels is completed in sequence to realize resampling and ensure the clarity and continuity of the corrected image.

[0044] Perform local geometric deformation compensation calculations. Although the distortion has been initially corrected through coordinate remapping, some areas may still have local geometric deformations, mainly manifested as incomplete elimination of local stretching or compression of the defect features. The compensation process is as follows: analyze the local feature distribution of the resampled image, compare it with the spatial deformation law of the virtual adaptive envelope ellipse, and for areas with local deformation, fine-tune the coordinates and gray values ​​of the pixels in that area according to the local change trend of the pixel displacement correction amount, so that the feature distribution of the local area is consistent with the feature distribution of the normal road surface, completely eliminating the feature stretching distortion caused by the curved surface projection. After bilinear interpolation resampling and local geometric deformation compensation, a geometrically corrected feature map is obtained.

[0045] Step 4.3: Input the geometrically corrected feature map into the pixel-level conditional random field semantic segmentation model, calculate the feature similarity potential function and label compatibility potential function between adjacent pixel nodes, and optimize the probability distribution of the disease area boundary by iterative energy minimization to obtain a binary semantic segmentation mask. Specifically, this includes: inputting the obtained geometrically corrected feature map into the pixel-level conditional random field semantic segmentation model. This model achieves accurate pixel-level segmentation based on the feature information of pixels and the correlation between adjacent pixels, distinguishing between disease areas and background areas.

[0046] Calculate the feature similarity potential function between adjacent pixel nodes ,node The multidimensional feature vector is denoted as ,node The multidimensional feature vector is denoted as , The feature smoothing parameter controls the sensitivity to feature differences; the label compatibility potential function is also relevant. The semantic label of node i is denoted as The semantic label of node j is denoted as , The label incompatibility penalty coefficient is a fixed positive number. The feature similarity potential function reflects the degree of feature similarity between two adjacent pixels. Specifically, it extracts the grayscale, texture, and edge features of two adjacent pixels, calculates the difference between these features, and combines this with a preset similarity coefficient to obtain a value reflecting the feature similarity between the two pixels. The more similar the features, the larger this value, indicating a higher probability that the two pixels belong to the same region and are both defects or both background. The label compatibility potential function constrains the label assignment of adjacent pixels, ensuring consistency in labels and avoiding isolated abnormal labels. Specifically, it calculates the penalty value for assigning different labels to two adjacent pixels based on preset label compatibility rules. The larger the penalty value, the lower the rationality of assigning different labels to two pixels.

[0047] The probability distribution of the diseased area boundary is optimized by iteratively minimizing the energy. The core of energy minimization is minimizing the energy function of the entire image, which is composed of the feature similarity potential function and the label compatibility potential function. During the iteration process, the label of each pixel is continuously adjusted and the corresponding energy value is calculated until the energy value reaches the minimum value. At this point, the label allocation result is the most reasonable and the probability distribution of the disease area boundary is the most accurate.

[0048] Based on the optimized probability distribution of the diseased area boundaries, a binarized semantic segmentation mask is obtained. This mask is a two-dimensional image with the same size as the geometric correction feature map. Pixels with a probability greater than a preset threshold are marked as 1, representing that the pixel belongs to the diseased area; pixels with a probability less than or equal to the preset threshold are marked as 0, representing that the pixel belongs to the background area. Through this binarization process, the location and range of all suspected diseased areas are clearly identified.

[0049] Step 4.4: Based on the binarized semantic segmentation mask, perform polygon edge morphological closure reconstruction. Perform line segment fitting and hole filling operations on the fracture boundaries, extract the complete connected component contours, and merge blocks according to spatial adjacency to obtain a preliminary set of disease location blocks. Specifically, this includes: based on the obtained binarized semantic segmentation mask, performing polygon edge morphological closure operations. The core of the morphological closure operation is to dilate and erode the mask. The purpose of dilation is to fill the small holes in the disease area and connect the fractured boundary lines; the purpose of erosion is to eliminate small noise points in the background area of ​​the mask and restore the true contour of the disease area, avoiding contour distortion caused by dilation. Through the closure operation, the boundary and hole problems of the disease area are initially repaired. For the mask after morphological closure, perform fracture boundary line segment fitting operations. For disease boundaries that still have fractures, extract the endpoints of the fracture boundaries, and use a line segment fitting algorithm to connect adjacent endpoints into smooth line segments, filling the boundary fractures and ensuring the continuity and integrity of the disease boundaries. The specific process of line segment fitting is as follows: calculate the coordinates of the endpoints of the fracture boundary, and fit a smooth line segment that accurately matches the distribution of the endpoints according to the distribution trend of the endpoints, so that the fracture boundary forms a complete closed contour; for holes that still exist in the disease area in the mask, that is, small areas that are mistakenly marked as background, after analyzing the pixel labels around the holes and confirming that the holes belong to part of the disease area, the pixel labels inside the holes are changed from 0 to 1 to complete the hole filling, ensure the integrity of the disease area, and avoid the deviation in the calculation of disease size caused by the existence of holes.

[0050] Extract complete connected component contours. A connected component is a region in the mask consisting of all adjacent pixels labeled 1. Each connected component corresponds to a suspected disease area. Using a connected component extraction algorithm, the entire mask is traversed to identify all connected components, extracting the contour information of each component, including its coordinates, shape, and size. Blocks are merged based on spatial adjacency. For connected components that are close together and spatially adjacent, it is determined whether they belong to the same disease entity, mainly based on the shape, size, and spatial relationship of the connected components. If they belong to the same disease entity, these connected components are merged into a complete disease block; if they do not belong to the same disease entity, their individual blocks are retained, ultimately resulting in a preliminary set of disease location blocks.

[0051] In this embodiment of the invention, the lateral and vertical offset components of the pixel displacement correction are respectively superimposed onto the original pixel grid coordinates of the disease candidate feature map to construct a spatial coordinate system remapping transformation matrix. Based on this matrix, bilinear interpolation resampling and local geometric deformation compensation operations are performed on the feature map. The geometrically corrected feature map is then input into a pixel-level conditional random field semantic segmentation model to calculate the feature similarity potential function and label compatibility potential function between adjacent pixel nodes. This optimizes the probability distribution of the disease region boundary through iterative energy minimization. Finally, based on a binarized semantic segmentation mask, polygon edge morphological closure reconstruction and fracture boundary segment fitting are performed. The technique of filling holes and merging blocks according to spatial adjacency overcomes the technical problems of the difficulty in completely eliminating residual feature stretching distortion after dynamic compensation, the tendency of traditional segmentation algorithms to break and fragment the disease boundary and overlap isolated artifacts under complex road textures and background noise interference, and the lack of topological closure and spatial aggregation mechanism that leads to the discreteness of subsequent temporal correlation targets. As a result, it achieves the technical effects of accurately restoring the true geometric shape and size of the disease, realizing smooth and continuous segmentation of the disease area boundary and seamless filling of internal holes, effectively removing background interference and aggregating according to spatial continuity to form topologically complete and continuously distributed high-confidence positioning blocks.

[0052] In a preferred embodiment of the present invention, step 5 above may include: Step 5.1: The pixel boundary coordinates of the preliminary defect location block set, combined with the real-time mileage data of the inspection vehicle and the lane line lateral offset parameters, are mapped to a continuous frame two-dimensional unfolded reference system with the highway mileage as the longitudinal reference and the lane lateral offset as the lateral reference through perspective projection inverse transformation, to obtain the block spatial coordinate mapping sequence. Specifically, this includes: extracting the pixel boundary coordinates of each block in the obtained preliminary defect location block set. The pixel boundary coordinates of each defect block are composed of the two-dimensional coordinates of all pixels on the edge of the block, forming a closed boundary contour. The position and range of each defect block in the geometric correction feature mapping map can be clearly defined through the coordinates; and calling the real-time mileage data of the inspection vehicle and the lane line lateral offset parameters. Among them, the real-time mileage data is collected in real time by the odometer on the inspection vehicle, in kilometers. The collection frequency is consistent with the imaging frame frequency to ensure that each frame of the image has corresponding mileage data. This data reflects the specific driving position of the inspection vehicle on the highway. The lane line lateral offset parameter is obtained by lane line recognition through geometric correction feature mapping map. With the center line of the highway lane as the reference, the lateral distance from each defect block to the center line is calculated, in meters, reflecting the specific lateral position of the defect in the lane.

[0053] An inverse perspective projection transformation is performed to map the pixel boundary coordinates to a continuous frame two-dimensional unfolded reference system. This reference system uses highway mileage markers as the vertical reference, with each marker corresponding to the real-time mileage of the inspection vehicle; the markers increase proportionally with each additional mileage. The lateral offset of the lane is used as the horizontal reference; a larger lateral coordinate value indicates a greater distance between the defect and the lane centerline. The specific process of the inverse perspective projection transformation is as follows: based on a preset camera intrinsic parameter calibration matrix and the vehicle chassis rigid mounting coordinate system, the pixel boundary coordinates of the defect area are inversely converted to highway physical space coordinates. These physical space coordinates are then mapped to the continuous frame two-dimensional unfolded reference system, obtaining the spatial coordinates of each defect area within this reference system. Following the temporal order of the imaging frames, the spatial coordinates of the defect areas in all frames are arranged sequentially, forming a sequence of mapped spatial coordinates.

[0054] Step 5.2: Based on the block spatial coordinate mapping sequence, calculate the centroid displacement vector and morphological overlap of corresponding blocks between adjacent consecutive frames. Use a temporal state machine to perform cross-frame block trajectory association and unique identifier binding to obtain a continuous disease movement trajectory chain. Specifically, this includes: based on the obtained block spatial coordinate mapping sequence, performing matching analysis on disease blocks in adjacent consecutive frames one by one, and calculating the centroid displacement vector and morphological overlap of corresponding blocks between adjacent frames. The calculation process for the centroid displacement vector is as follows: calculate the centroid coordinates of corresponding disease blocks in two adjacent frames respectively. The centroid displacement vector is obtained by subtracting the centroid coordinates of the previous frame from the centroid coordinates of the next frame block. This vector reflects the magnitude and direction of the displacement of the lesion block between adjacent frames. The morphological overlap is calculated by dividing the overlapping area of ​​corresponding blocks in two adjacent frames by the sum of the areas of the two blocks. The closer this value is to 1, the more similar the shapes and positions of the two blocks are, and the more likely they are to be the same lesion. A temporal state machine is used to perform cross-frame block trajectory association. The temporal state machine uses preset association thresholds, including centroid displacement thresholds and morphological overlap thresholds, to determine whether lesion blocks in adjacent frames are the same lesion: if the centroid displacement vectors of two blocks in adjacent frames are less than the preset centroid displacement threshold and the morphological overlap is greater than the preset morphological overlap threshold, they are determined to be the same lesion and associated; if the threshold requirements are not met, they are determined to be different lesions and tracked separately.

[0055] A unique identifier is bound to each successfully associated disease block. A unique identifier is assigned to each associated disease, and this identifier remains unchanged throughout all consecutive frames, regardless of the position of the disease block in the consecutive frames. This ensures that each disease can be tracked individually and avoids confusion in association between different diseases. The consecutive associated blocks of each disease are connected in chronological order to form a chain of continuous movement trajectories of the diseases.

[0056] Step 5.3: Along the continuous motion trajectory chain of the defect, extract the disparity change features and scale scaling factors of each frame block. Perform disparity consistency verification based on the preset static road surface disparity evolution constraint rules and spatial geometric consistency judgment criteria. Remove artifact blocks caused by dynamic obstructions and transient light and shadow interference to obtain a subset of static defect candidate trajectories. Specifically, this includes: along each obtained continuous motion trajectory chain of the defect, extracting the disparity change features and scale scaling factors of each frame block. The disparity change feature extraction process is as follows: using the optical center of the vehicle-mounted imaging module as a reference, calculate the disparity change features of each frame block... The disparity value of the defect area in a frame is calculated as the difference in the projected distance from the area to the optical center. The change in disparity value between adjacent frames is then calculated to form a disparity change feature, which reflects the disparity change pattern of the defect area during imaging. The extraction process of the scale scaling factor involves calculating the area ratio of corresponding defect areas in adjacent frames, dividing the area of ​​the defect in the later frame by the area in the previous frame to obtain the scale scaling factor. This factor reflects the size change of the defect area in consecutive frames. Pre-defined static road surface disparity evolution constraint rules and spatial geometric consistency judgment criteria are applied. The static road surface disparity evolution constraint rules specify the range of disparity change for static defects, meaning the disparity change between adjacent frames should be less than a preset disparity threshold, and the disparity change trend should be consistent with the vehicle's driving direction, conforming to perspective imaging rules. The spatial geometric consistency judgment criteria specify the range of the scale scaling factor for static defects, meaning the scale scaling factor should be close to 1, with a deviation not exceeding a preset scale deviation threshold, because the actual size of static defects is fixed, and the imaging size should not change significantly in consecutive frames.

[0057] Based on the aforementioned rules and criteria, disparity consistency verification is performed. Each continuous motion trajectory chain of a defect is verified one by one. If the disparity change characteristics of a trajectory chain conform to the static pavement disparity evolution constraint rules, and the scale scaling factor conforms to the spatial geometric consistency judgment criteria, then the defect corresponding to the trajectory chain is determined to be a static defect and is retained; otherwise, the defect corresponding to the trajectory chain is determined to be an artifact block, such as transient water mist, shadows, dynamic occlusions, etc., and is removed. All static defect trajectory chains that pass the disparity consistency verification are integrated to obtain a subset of static defect candidate trajectories. This subset only contains the actual static pavement defect trajectories.

[0058] Step 5.4: Calculate the spatiotemporal correlation confidence weights based on the proportion of consecutively existing frames, trajectory spatial smoothness, and disparity check matching degree of the static disease candidate trajectory subset. Combine these weights in chronological order to obtain the disease spatiotemporal confidence sequence. Specifically, this includes: determining the three core calculation indicators for the spatiotemporal correlation confidence weights, namely the proportion of consecutively existing frames, trajectory spatial smoothness, and disparity check matching degree. These three indicators jointly determine the confidence level of the disease; the higher the indicator value, the higher the confidence level of the disease. Calculate the three indicators for each static disease candidate trajectory. The calculation process for the percentage of consecutively existing frames is as follows: The actual number of frames existing in the trajectory within consecutive frames is counted, and this number is divided by the total number of imaging frames inspected to obtain the percentage of consecutively existing frames. This indicator reflects the stability of the disease throughout the inspection process; the more frames existing, the higher the reliability. The calculation process for trajectory spatial smoothness is as follows: The difference between the centroid displacement vectors of adjacent frames in the trajectory is calculated, and the average of all differences is taken. The smaller the average, the smoother the spatial change of the trajectory, the more it conforms to the motion law of static diseases, and the higher the reliability. The calculation process for disparity verification matching degree is as follows: The number of frames in the trajectory that conform to the disparity consistency verification rules is counted, and this number is divided by the total number of existing frames in the trajectory to obtain the disparity verification matching degree. This indicator reflects the regularity of the disparity changes in the disease; the higher the matching degree, the higher the reliability.

[0059] The spatiotemporal correlation confidence weight is calculated by weighted summation of the three indicators. Preset weight coefficients for the three indicators, based on their importance, with disparity verification matching degree having the highest weight, followed by trajectory spatial smoothness, and the proportion of consecutively existing frames having the lowest weight. The sum of the three weight coefficients is 1. The specific calculation method is as follows: multiply the proportion of consecutively existing frames by its corresponding weight coefficient, add the trajectory spatial smoothness multiplied by its corresponding weight coefficient, and add the disparity verification matching degree multiplied by its corresponding weight coefficient to obtain the spatiotemporal correlation confidence weight of the static defect. The weight value ranges from 0 to 1; the closer the value is to 1, the higher the authenticity and credibility of the defect. The spatiotemporal correlation confidence weights of each static defect in each frame are sequentially combined according to the timeline to obtain the defect spatiotemporal confidence sequence.

[0060] In this embodiment of the invention, the pixel boundary coordinates of the preliminary location block set of the defect are combined with the real-time mileage data of the inspection vehicle and the lane lateral offset parameters, and mapped through perspective projection inverse transformation to a continuous frame two-dimensional unfolded reference system with the highway mileage as the longitudinal reference and the lane lateral offset as the lateral reference. The centroid displacement vector and morphological overlap of the corresponding blocks between adjacent consecutive frames are calculated, and a temporal state machine is used to perform cross-frame block trajectory association and unique identifier binding. The disparity change features and scale factor of each frame block are extracted along the continuous motion trajectory chain of the defect, and disparity consistency is checked according to the preset static road surface disparity evolution constraint rules and spatial geometric consistency judgment criteria to remove artifact blocks caused by dynamic obstructions and transient light and shadow interference. The technique of combining the proportion of consecutive frames of static defect candidate trajectory subsets, trajectory spatial smoothness, and disparity verification matching degree in a time-axis order to calculate spatiotemporal correlation confidence weights overcomes the technical problems of existing detection methods. These problems include the lack of a unified physical coordinate reference in continuous frame processing, which leads to defect location drift with vehicle pose; cross-frame target correlation is easily affected by instantaneous occlusion and sudden changes in light and shadow, resulting in false alarms or target loss; and the lack of a stability verification mechanism based on spatiotemporal continuity and physical disparity laws, which makes it difficult to accurately separate static real defects from dynamic environmental artifacts. This technique achieves accurate mapping and stable tracking of defect targets from two-dimensional pixel space to the actual physical coordinate system of highways, and constructs a confidence sequence that quantitatively reflects the stability and spatial consistency of defect existence.

[0061] In a preferred embodiment of the present invention, step 6 above may include: Step 6.1: Extract local feature blocks from the geometrically corrected feature map corresponding to each confidence node in the spatiotemporal confidence sequence of the disease. Divide the feature blocks into multi-scale feature levels according to the disease morphology scale. Perform multi-scale feature weighted voting operation on each scale feature level in combination with the confidence weight of the spatiotemporal confidence sequence of the disease to obtain a preliminary classification label set for the disease type. Specifically, this includes: extracting local feature blocks from the geometrically corrected feature map corresponding to each confidence node in the spatiotemporal confidence sequence of the disease. Each confidence node corresponds to a time point in the spatiotemporal confidence sequence of the disease, that is, it corresponds to a frame of geometrically corrected feature map. The local feature block refers to the feature region of the preliminary localized disease block in the frame, which contains the core information of the disease such as grayscale features, texture features, and edge features. During extraction, it is necessary to accurately correspond to each confidence node to ensure that the feature block and the confidence weight correspond one-to-one and avoid the situation of feature and weight misalignment. Divide the extracted local feature blocks into multi-scale feature levels according to the disease morphology scale. Based on the common types and morphological characteristics of highway pavement distresses, they are typically divided into three scale levels: small-scale, medium-scale, and large-scale. The small-scale level mainly corresponds to minor distresses such as microcracks, and its feature blocks primarily contain subtle edge textures and local grayscale variations. The medium-scale level mainly corresponds to medium-sized distresses such as potholes and minor damage, and its feature blocks contain clear boundary outlines and a certain range of texture anomalies. The large-scale level mainly corresponds to large-area damage and cracking, and its feature blocks contain extensive abnormal areas and obvious morphological features. During the classification process, each local feature block is assigned to its corresponding scale level based on its pixel span and outline size, ensuring that each feature block accurately matches the appropriate scale classification.

[0062] By combining the confidence weights of the spatiotemporal confidence sequence of the disease, a multi-scale feature weighted voting operation is performed on each scale feature level. The core of the voting is to assign a corresponding voting weight to each scale feature level. This voting weight is related to the confidence weight of the spatiotemporal confidence sequence of the disease. The higher the confidence weight, the greater the voting weight of the corresponding scale feature level, and the stronger the influence of the voting result. The specific voting process is as follows: Common types of highway pavement defects are preset, including micro-cracks, potholes, large-area damage, and alligator cracks. For each scale feature level, local feature blocks are matched with the closest defect type and voted on. The voting weight of each feature block is equal to its corresponding spatiotemporal confidence weight, multiplied by the weight coefficient of that scale level. The weight coefficients for small, medium, and large scale levels are set according to the defect identification accuracy requirements to ensure accurate identification of defects at different scales. The voting results of all feature blocks are statistically analyzed. The total number of votes for each defect type is the sum of the voting weights of all feature blocks supporting that type. The defect type with the highest total number of votes is the preliminary classification label for that defect. All preliminary classification labels for static defects are organized according to the defect's unique identifier to form a preliminary defect type classification label set.

[0063] Step 6.2 involves setting a primary threshold for morphological discrimination and a secondary threshold for temporal stability. A dual-threshold decision fusion process is then performed on the preliminary disease type classification label set. The primary threshold filters out low-confidence noise classifications, while the secondary threshold verifies the logical consistency of classification across consecutive frames, ultimately locking in the final disease type label. Specifically, this includes setting a primary threshold for morphological discrimination and a secondary threshold for temporal stability. These two thresholds work together to optimize classification. The primary threshold for morphological discrimination is a confidence threshold, ranging from 0 to 1, set according to the accuracy requirements of highway pavement disease detection, primarily used to filter out low-confidence noise classifications. The secondary threshold for temporal stability is a frame count threshold, primarily used to verify the logical consistency of disease classification labels in consecutive frames, ensuring that the classification results conform to the characteristics of static diseases. The preliminary disease type classification label set is then subjected to dual-threshold decision fusion processing, with the primary threshold for morphological discrimination filtering out low-confidence noise classifications. Each disease's preliminary classification label is verified one by one, and the spatiotemporal correlation confidence weight corresponding to the disease is extracted. If the confidence weight is greater than or equal to the morphological discrimination master threshold, it indicates that the preliminary classification label has high credibility and is retained. If the confidence weight is less than the morphological discrimination master threshold, it indicates that the preliminary classification label may be a noisy classification, such as an incorrect classification caused by artifact residue or feature mismatch, and is removed. At the same time, the disease is marked as pending confirmation and further judgment is made in conjunction with time-series verification.

[0064] The consistency of classification logic across consecutive frames is verified using a temporally stable secondary threshold. For diseases that pass the primary threshold verification, their preliminary classification label sequence in consecutive frames is extracted, and the number of consecutive frames with the same classification label in this sequence is counted. If the number of consecutive frames is greater than or equal to the temporally stable secondary threshold, it indicates that the classification label of the disease remains stable in consecutive frames, conforming to the classification logic of static diseases, and the classification label is confirmed as a valid label. If the number of consecutive frames is less than the temporally stable secondary threshold, it indicates that the classification label of the disease changes frequently, and there may be a classification error. The multi-scale features and confidence weights of the disease need to be re-verified. If a stable classification label cannot be determined after re-verification, the disease is marked as pending classification. The classification labels of diseases that pass the dual threshold decision are integrated to lock the final disease type label for each disease. For diseases with pending classification, a pending label is marked and relevant feature information is recorded.

[0065] Step 6.3: Based on the spatial mapping relationship between the final disease type label and the spatiotemporal confidence sequence of the disease, extract the pixel span of the disease boundary contour. Combined with the projection scale parameters of the continuous frame two-dimensional unfolded reference system, calculate the actual longitudinal length, lateral width, and equivalent projected area of ​​the disease. Specifically, this includes: determining the boundary contour position of each disease in the continuous frame two-dimensional unfolded reference system based on the locked spatial mapping relationship between the final disease type label and the disease spatiotemporal confidence sequence. The spatial mapping relationship of the disease spatiotemporal confidence sequence has clearly defined the spatial coordinates of each disease in the continuous frame two-dimensional unfolded reference system. Combined with the final disease type label, the boundary contour range of the disease can be accurately located; extract the pixel span of the disease boundary contour. Pixel span refers to the number of horizontal and vertical pixels in the geometric correction feature map of the lesion boundary contour. The specific extraction process is as follows: traverse all pixel coordinates of the lesion boundary contour, find the maximum and minimum horizontal coordinates of the contour in the horizontal direction (image width direction), and subtract the minimum horizontal coordinate from the maximum horizontal coordinate to obtain the horizontal pixel span; find the maximum and minimum vertical coordinates of the contour in the vertical direction (image height direction), and subtract the minimum vertical coordinate from the maximum vertical coordinate to obtain the vertical pixel span; the horizontal pixel span corresponds to the horizontal width pixel value of the lesion, and the vertical pixel span corresponds to the vertical length pixel value of the lesion.

[0066] The projection scale parameter of the continuous frame 2D unfolded reference frame is invoked. This parameter is a preset conversion relationship between pixels and actual physical size, in pixels per meter, that is, how many pixels correspond to one meter of actual road surface. Its value is calculated in advance through the camera intrinsic parameter calibration matrix, the installation height of the vehicle-mounted imaging module, and the driving parameters of the inspection vehicle to ensure the accuracy of the conversion. The process of determining the projection scale parameter is as follows: divide the pixel span of a reference object of actual known length, such as the standard lane width in the image, by the actual length of the reference object to obtain the projection scale parameter.

[0067] Calculate the actual longitudinal length, transverse width, and equivalent projected area of ​​the disease. The specific calculation process is as follows: Divide the longitudinal pixel span by the projection scale parameter to obtain the actual longitudinal length of the disease (in meters); divide the transverse pixel span by the projection scale parameter to obtain the actual transverse width of the disease (in meters); the equivalent projected area is calculated by multiplying the actual longitudinal length by the actual transverse width to obtain the actual equivalent projected area of ​​the disease (in square meters). For irregularly shaped diseases, such as irregular pits and cracks, the boundary contour needs to be fitted first, the equivalent rectangle size calculated, and then the actual physical dimensions calculated in the above manner to ensure the accuracy of the dimensional calculation and meet the needs of refined maintenance for disease dimensional data.

[0068] Step 6.4 involves structured data encapsulation and standardized formatting of the final defect type label, calculated actual physical dimensions, corresponding station coordinates, and lane assignment information to obtain a structured pavement defect detection report. This includes collecting all defect-related detection data, including the locked final defect type label, calculated actual physical dimensions of the defect (longitudinal length, transverse width, equivalent projected area), the obtained station coordinates and lane assignment information, and the obtained spatiotemporal association confidence weight. Simultaneously, basic inspection information is supplemented, including inspection time, inspected road section, inspection vehicle number, and vehicle-mounted imaging module parameters, ensuring the completeness of the report data. All collected data is then structured data encapsulated. The core of structured encapsulation is organizing scattered data according to fixed fields to form a standardized data structure. Each defect corresponds to a complete data record, and each record contains the following core fields: unique defect identifier, final defect type, actual longitudinal length, actual transverse width, equivalent projected area, station coordinates, lane assignment, spatiotemporal association confidence weight, inspection time, and inspected road section. During the encapsulation process, ensure that the data format of each field is consistent. For example, length, width, and area should be retained to two decimal places, station number and mileage coordinates should be retained to three decimal places, and confidence weights should be retained to three decimal places to avoid data format confusion.

[0069] The packaged structured data is standardized and formatted. The formatting must adhere to the requirements of refined maintenance decision-making for high-grade highways, employing a clear and standardized format to ensure the report is easy to read, understand, and directly applicable to maintenance decisions. The formatting structure mainly includes three parts: The first part is basic inspection information, including inspection time, inspected road section, inspection vehicle number, imaging module parameters, etc., centrally displaying the basic inspection situation; the second part is a summary statistical analysis of defects, including the total number of defects, the number of defects of each type, the distribution of defects in each lane, and the distribution of defect confidence levels, facilitating a quick understanding of the overall defect situation of the inspected road section; the third part is detailed information on individual defects, arranged sequentially according to the defect's unique identifier. Each defect displays all its structured data fields in detail, along with a defect boundary contour diagram extracted from the geometric correction feature map.

[0070] The standardized report is reviewed to verify the accuracy and completeness of the data, ensuring there are no data errors, missing fields, or disordered formats. Once the review is approved, the final structured pavement defect detection report is generated.

[0071] In this embodiment of the invention, local feature blocks of the geometrically corrected feature map corresponding to each confidence node in the spatiotemporal confidence sequence of the disease are extracted and multi-scale feature levels are divided according to the morphological scale of the disease. A preliminary classification label set is obtained by performing multi-scale feature weighted voting operation combined with confidence weights. A preset morphological discrimination primary threshold and a temporally stable secondary threshold are used to perform dual-threshold decision fusion processing on the label set to filter low-confidence noise classification and verify the consistency of classification logic across consecutive frames to lock the final disease type label. Based on the final label and spatial mapping relationship, the pixel span of the disease boundary contour is extracted, and the actual longitudinal length, lateral width, and equivalent projected area of ​​the disease are calculated by combining the projection scale parameters of the two-dimensional unfolded reference system of consecutive frames. Finally, the final disease type label, the calculated actual physical dimensions, the corresponding station mileage coordinates, and lane affiliation information are encapsulated in structured data and standardized in layout. Therefore, this method overcomes the limitations of traditional disease identification methods. The traditional single-scale feature classification method is susceptible to misjudgment due to local texture noise interference; the lack of temporal logic consistency verification makes dynamic transient artifacts easy to be misclassified; image pixel features are difficult to accurately invert into actual engineering physical dimensions; and the fragmented output data dimensions lack a unified structured standard, making it difficult to directly connect with maintenance engineering acceptance and decision-making systems. This paper addresses these technical problems by achieving deep fusion and stable judgment of multi-scale disease morphological features and high-confidence spatiotemporal weights, ensuring the noise resistance and robustness of disease type identification and the logical consistency of continuous frames. It completes the accurate quantitative inversion from image pixel space to actual road surface physical dimensions, and ultimately directly outputs a high-confidence structured pavement disease detection report that integrates accurate disease classification, actual quantitative dimensions, station spatial coordinates, and lane assignment information. This achieves the technical effect of comprehensively supporting the refined maintenance planning of high-grade highways, accurate calculation of engineering quantities, and digital management of the entire life cycle of infrastructure.

[0072] like Figure 2 As shown, embodiments of the present invention also provide a highway pavement defect detection system based on image recognition, comprising: The acquisition module is used to acquire the original road surface image sequence; the original road surface image sequence is input into the preset multi-scale convolutional feature extraction architecture to obtain the disease candidate feature map; The module is used to extract three-dimensional spatial reference anchor points based on the defect candidate feature map, which are preset at the inner lining of the left wheel arch, the inner lining of the right wheel arch, and the optical center of the vehicle imaging module on the chassis of the inspection vehicle; and to construct a virtual adaptive envelope ellipse by fitting the field of view attenuation gradient of the three-dimensional spatial reference anchor points with the dynamic pose angle of the vehicle. The calculation module is used to extract orthogonal distortion gradient components from the virtual adaptive envelope ellipse, construct the polar coordinate circular state mapping trajectory, solve the extreme value envelope path of the components through continuous azimuth angle rotation transformation, decouple the distortion vector and lock the main distortion axis direction and orthogonal compensation amplitude, and perform geometric vector orthogonal decomposition operation accordingly to obtain the pixel displacement correction amount of the anisotropic projection distortion law. The compensation module is used to apply the pixel displacement correction amount to the disease candidate feature map to perform spatial coordinate system remapping and local geometric deformation compensation. After pixel-level conditional random field semantic segmentation and polygon edge morphological closure reconstruction, a set of preliminary disease location blocks is obtained. The verification module is used to project the set of preliminary location blocks of the defects onto a continuous frame two-dimensional unfolded reference system based on the highway station mileage and lane lateral offset, and to perform cross-frame block trajectory association and disparity consistency verification to obtain the spatiotemporal confidence sequence of the defects. The processing module is used to perform multi-scale feature weighted voting and dual-threshold decision fusion processing based on the spatiotemporal confidence sequence of the pavement defects to obtain a structured pavement defect detection report.

[0073] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0074] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0075] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0076] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for detecting highway pavement defects based on image recognition, characterized in that, The method includes: Obtain the original road surface image sequence; input the original road surface image sequence into the preset multi-scale convolutional feature extraction architecture to obtain the disease candidate feature map; Based on the defect candidate feature mapping map, three-dimensional spatial reference anchor points are extracted from the inner lining of the left wheel arch, the inner lining of the right wheel arch, and the optical center of the vehicle imaging module on the chassis of the inspection vehicle. A virtual adaptive envelope ellipse is constructed by fitting the field attenuation gradient of the three-dimensional spatial reference anchor points with the dynamic pose angle of the vehicle. Orthogonal distortion gradient components are extracted from the virtual adaptive envelope ellipse, and polar coordinate circular state mapping trajectory is constructed. The extreme value envelope path of the components is solved by continuous azimuth angle rotation transformation. The distortion vector is decoupled and the principal distortion axis direction and orthogonal compensation amplitude are locked. Based on this, the geometric state vector orthogonal decomposition operation is performed to obtain the pixel displacement correction amount of the anisotropic projection distortion law. The pixel displacement correction is applied to the disease candidate feature map to perform spatial coordinate system remapping and local geometric deformation compensation. After pixel-level conditional random field semantic segmentation and polygon edge morphological closure reconstruction, a preliminary set of disease localization blocks is obtained. The preliminary location of the disease blocks is projected onto a continuous two-dimensional unfolded reference system based on the highway station mileage and lane lateral offset. Cross-frame block trajectory association and disparity consistency verification are performed to obtain the spatiotemporal confidence sequence of the disease. Based on the spatiotemporal confidence sequence of the pavement defects, a multi-scale feature weighted voting and dual-threshold decision fusion process is performed to obtain a structured pavement defect detection report.

2. The method for detecting highway pavement defects based on image recognition according to claim 1, characterized in that, Obtain the original road surface image sequence; The original road surface image sequence is input into a pre-defined multi-scale convolutional feature extraction architecture to obtain a disease candidate feature map, including: Dynamic illumination compensation and motion blur deconvolution processing are performed on consecutive image frames in the original road surface image sequence to eliminate local overexposure and motion blur interference under high-speed driving conditions, resulting in a time-aligned standardized road surface image sequence. Standardized road image sequences are superimposed with inter-frame features according to a preset sliding step size to construct a multi-dimensional channel feature tensor, and the multi-dimensional channel feature tensor is injected into the bottom convolutional branch of the multi-scale convolutional feature extraction architecture. In the multi-scale convolutional feature extraction architecture, the road texture response signals at different receptive field scales are captured by a parallel dilated convolutional structure. Cross-layer feature pooling and channel weight recalibration operations are performed to aggregate the discrete features of the asphalt mixture surface and the response information of microcrack edges to obtain a disease candidate feature map.

3. The method for detecting highway pavement defects based on image recognition according to claim 2, characterized in that, Based on the disease candidate feature mapping map, three-dimensional spatial reference anchor points are extracted from the inner lining of the left wheel arch, the inner lining of the right wheel arch, and the optical center of the vehicle imaging module of the inspection vehicle chassis. A virtual adaptive envelope ellipse is constructed by fitting the field-of-view attenuation gradient of the three-dimensional space reference anchor point with the vehicle's dynamic pose angle, including: By analyzing the pixel spatial distribution matrix of the candidate feature map of the disease, and combining the preset camera intrinsic parameter calibration matrix with the vehicle chassis rigid installation coordinate system, the spatial coordinates of the three-dimensional spatial reference anchor point at the inner lining of the left wheel arch, the inner lining of the right wheel arch, and the optical center of the vehicle imaging module of the inspection vehicle chassis are located. Based on the spatial coordinates of the three-dimensional reference anchor point, the illumination attenuation gradient of the edge region of the continuous imaging frame is calculated. The pitch angle, roll angle and yaw angle deflection data output by the vehicle inertial measurement unit are collected simultaneously to synthesize the vehicle dynamic pose deflection vector of the vehicle running attitude fluctuation. The field-of-view attenuation gradient and the vehicle dynamic pose deflection vector are mapped to the three-dimensional imaging projection space. Nonlinear surface interpolation iterative calculation is used to fit a virtual adaptive envelope ellipse with a continuous transition of boundary curvature. This allows the principal axis direction and major and minor axis radii of the virtual adaptive envelope ellipse to dynamically adapt to the spatial deformation trend of the road imaging field of view, thus completing the spatial surface construction of the virtual adaptive envelope ellipse.

4. The method for detecting highway pavement defects based on image recognition according to claim 3, characterized in that, Orthogonal distortion gradient components are extracted from the virtual adaptive envelope ellipse, and a polar coordinate circular state mapping trajectory is constructed. The extreme value envelope path of the components is solved through continuous azimuth rotation transformation. The distortion vector is decoupled and the principal distortion axis direction and orthogonal compensation amplitude are locked. Based on this, a geometric state vector orthogonal decomposition operation is performed to obtain the pixel displacement correction amount according to the anisotropic projection distortion law, including: On the three-dimensional surface mesh of the virtual adaptive envelope ellipse, surface curvature change rate data are collected along the radial and tangential orthogonal directions, and orthogonal distortion gradient components are extracted. Based on orthogonal distortion gradient components, with the geometric center of the virtual adaptive envelope ellipse as the origin of polar coordinates, the gradient magnitude and phase angle of each grid node are mapped to polar coordinate space to construct a polar coordinate circular state mapping trajectory. Perform a continuous azimuth rotation transformation with a preset angular step size along the polar coordinate circular state mapping trajectory, track the peak values ​​of the gradient component response under each rotation phase, and fit to obtain a smooth and continuous component extremum envelope path. Based on the phase difference characteristics of the component extreme value envelope path, the original coupled distortion vector is separated by spatial orthogonal projection, decoupled cross interference components, and locked the main distortion principal axis direction and orthogonal compensation amplitude. Using the principal distortion axis as the reference coordinate axis, the geometric vector orthogonal decomposition operation is performed on the orthogonal distortion gradient components to solve the surface distortion response into horizontal and vertical offset components in the two-dimensional image plane, and the pixel displacement correction amount of the anisotropic projection distortion law is synthesized.

5. The method for detecting highway pavement defects based on image recognition according to claim 4, characterized in that, The pixel displacement correction is applied to the disease candidate feature map for spatial coordinate system remapping and local geometric deformation compensation. After pixel-level conditional random field semantic segmentation and polygon edge morphological closure reconstruction, a preliminary set of disease localization blocks is obtained, including: The horizontal and vertical offset components of the pixel displacement correction are superimposed onto the original pixel grid coordinates of the disease candidate feature map to construct a spatial coordinate system remapping transformation matrix. Based on the spatial coordinate system remapping transformation matrix, bilinear interpolation resampling and local geometric deformation compensation operations are performed on the disease candidate feature mapping map to eliminate feature stretching distortion caused by surface projection and obtain geometrically corrected feature mapping map. The geometrically corrected feature map is input into the pixel-level conditional random field semantic segmentation model. The feature similarity potential function and label compatibility potential function between adjacent pixel nodes are calculated. The probability distribution of the disease area boundary is optimized by iterative energy minimization to obtain the binarized semantic segmentation mask. Based on the binary semantic segmentation mask, polygon edge morphological closure reconstruction is performed. Line segment fitting and hole filling operations are performed on the fracture boundary. The complete connected domain contour is extracted and the blocks are merged according to the spatial adjacency relationship to obtain a set of preliminary disease location blocks.

6. The method for detecting highway pavement defects based on image recognition according to claim 5, characterized in that, The initial location of the defects is projected onto a continuous two-dimensional unfolded reference frame based on highway mileage and lane lateral offset. Cross-frame block trajectory correlation and disparity consistency verification are performed to obtain the spatiotemporal confidence sequence of the defects, including: The pixel boundary coordinates of the initial location block set of the disease are combined with the real-time mileage data of the inspection vehicle and the lateral offset parameter of the lane line, and mapped to a continuous frame two-dimensional unfolded reference system with the highway mileage as the longitudinal reference and the lane lateral offset as the lateral reference through perspective projection inverse transformation, so as to obtain the block space coordinate mapping sequence. Based on the block spatial coordinate mapping sequence, the centroid displacement vector and morphological overlap of corresponding blocks between adjacent consecutive frames are calculated. A time-series state machine is used to perform cross-frame block trajectory association and unique identifier binding to obtain the continuous motion trajectory chain of the disease. Along the continuous motion trajectory chain of the disease, the disparity change features and scale scaling factors of each frame block are extracted. Based on the preset static road disparity evolution constraint rules and spatial geometric consistency judgment criteria, disparity consistency is checked. Artifact blocks caused by dynamic occlusions and transient light and shadow interference are removed to obtain a subset of static disease candidate trajectories. Based on the proportion of consecutive frames in the static disease candidate trajectory subset, trajectory spatial smoothness, and disparity check matching degree, the spatiotemporal correlation confidence weight is calculated, and the disease spatiotemporal confidence sequence is obtained by combining them in time axis order.

7. The method for detecting highway pavement defects based on image recognition according to claim 6, characterized in that, Based on the spatiotemporal confidence sequence of pavement defects, a multi-scale feature-weighted voting and dual-threshold decision fusion process is performed to obtain a structured pavement defect detection report, including: Local feature blocks of the geometrically corrected feature map corresponding to each confidence node in the spatiotemporal confidence sequence of the disease are extracted, and the disease is divided into multi-scale feature levels according to the morphological scale. Multi-scale feature weighted voting operation is performed on each scale feature level in combination with the confidence weight of the spatiotemporal confidence sequence of the disease to obtain a preliminary classification label set of disease types. The system presets a primary threshold for morphological discrimination and a secondary threshold for temporal stability. It then performs dual-threshold decision fusion processing on the preliminary classification label set of disease types. The primary threshold filters out low-confidence noise classifications, and the secondary threshold verifies the consistency of classification logic in consecutive frames to lock in the final disease type label. Based on the spatial mapping relationship between the final disease type label and the spatiotemporal confidence sequence of the disease, the pixel span of the disease boundary contour is extracted. Combined with the projection scale parameters of the continuous frame two-dimensional unfolded reference system, the actual longitudinal length, transverse width and equivalent projected area of ​​the disease are calculated. The final pavement type label, the calculated actual physical dimensions, the corresponding station mileage coordinates and lane affiliation information are encapsulated in a structured data format and standardized in layout to obtain a structured pavement defect detection report.

8. A highway pavement defect detection system based on image recognition, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The acquisition module is used to acquire the original road surface image sequence; The original road surface image sequence is input into a preset multi-scale convolutional feature extraction architecture to obtain a disease candidate feature map. The module is used to extract three-dimensional spatial reference anchor points based on the disease candidate feature mapping map, which are preset at the inner lining of the left wheel arch of the inspection vehicle chassis, the inner lining of the right wheel arch, and the optical center of the vehicle imaging module. A virtual adaptive envelope ellipse is constructed by fitting the field-of-view attenuation gradient of the three-dimensional space reference anchor point with the vehicle's dynamic pose angle. The calculation module is used to extract orthogonal distortion gradient components from the virtual adaptive envelope ellipse, construct the polar coordinate circular state mapping trajectory, solve the extreme value envelope path of the components through continuous azimuth angle rotation transformation, decouple the distortion vector and lock the main distortion axis direction and orthogonal compensation amplitude, and perform geometric vector orthogonal decomposition operation accordingly to obtain the pixel displacement correction amount of the anisotropic projection distortion law. The compensation module is used to apply the pixel displacement correction amount to the disease candidate feature map to perform spatial coordinate system remapping and local geometric deformation compensation. After pixel-level conditional random field semantic segmentation and polygon edge morphological closure reconstruction, a set of preliminary disease location blocks is obtained. The verification module is used to project the set of preliminary location blocks of the defects onto a continuous frame two-dimensional unfolded reference system based on the highway station mileage and lane lateral offset, and to perform cross-frame block trajectory association and disparity consistency verification to obtain the spatiotemporal confidence sequence of the defects. The processing module is used to perform multi-scale feature weighted voting and dual-threshold decision fusion processing based on the spatiotemporal confidence sequence of the pavement defects to obtain a structured pavement defect detection report.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.