A method for three-dimensional reconstruction of asphalt pavement texture

By employing short-spacing narrow-baseline acquisition, frequency domain filtering enhancement, and feature matching optimization techniques, the efficiency and accuracy issues of 3D reconstruction of asphalt pavement during vehicle movement have been resolved, achieving efficient and accurate 3D modeling suitable for rapid detection of asphalt pavement.

CN122089946APending Publication Date: 2026-05-26SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2026-01-28
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and accurately reconstruct asphalt pavements in three dimensions while vehicles are in motion, especially due to the difficulty in matching feature points and data distortion caused by the weak texture characteristics of asphalt pavements.

Method used

By employing short-spacing, narrow-baseline continuous acquisition, frequency-domain low-pass filtering convolution enhancement, SIFT feature matching within a defined neighborhood, beam adjustment optimization, and dual noise reduction combined with projection transformation correction, rapid and high-precision 3D modeling of asphalt pavement is achieved.

Benefits of technology

Under the premise of normal vehicle traffic without disrupting traffic, high-precision 3D reconstruction of weak textured asphalt pavement was achieved, solving the global matching confusion problem caused by repetitive and single textures, improving the number of feature point detections and descriptor discrimination, and ensuring the geometric consistency and engineering-grade accuracy of the 3D model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089946A_ABST
    Figure CN122089946A_ABST
Patent Text Reader

Abstract

This invention discloses a method for 3D reconstruction of asphalt pavement texture, belonging to the fields of road detection and computer vision technology. The method first calibrates an industrial camera to obtain intrinsic parameter matrices and distortion coefficients, and then continuously acquires pavement images at short intervals of 1 to 2 centimeters on a moving vehicle. Next, the images undergo preprocessing including cropping, Gaussian kernel convolution for noise reduction, and 8 to 12 low-pass filtering convolution enhancements. Feature points are matched using the SIFT algorithm combined with a limited neighborhood range strategy of 80 to 120 pixels. Subsequently, based on the matched points, the fundamental and essential matrices are solved using the RANSAC algorithm and the 8-point method. The camera pose is optimized using bundle adjustment to obtain an initial point cloud, followed by kdtree clustering and statistical filtering for dual noise reduction. Finally, bending deformation is corrected through quadratic polynomial surface fitting, tilt is corrected using projection transformation, and the model scale is adjusted according to the scaling factor to obtain a standardized 3D model. This method effectively solves the problem of weak asphalt texture matching and achieves efficient and high-precision pavement texture reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of road inspection and computer vision technology, specifically to a method for three-dimensional reconstruction of asphalt pavement texture. Background Technology

[0002] Asphalt pavement surface texture is not only a key factor affecting driving safety and comfort, but also an important indicator for evaluating pavement skid resistance and drainage performance. Traditional pavement texture detection methods, such as the sand-spreading method, are simple to operate, but inefficient and require road closures, severely hindering traffic flow and failing to meet the practical needs of large-scale road network inspections. In recent years, although laser scanning technology can acquire high-precision pavement depth data, the equipment is expensive, maintenance costs are high, and data is easily distorted by vehicle vibrations during high-speed acquisition, making it difficult to widely promote and apply on ordinary inspection vehicles.

[0003] With the development of machine vision technology, 3D reconstruction technology based on multi-view geometry is gradually being applied to road surface detection. Most mainstream 3D reconstruction software uses motion recovery structure algorithms combined with multi-view stereo vision technology. These methods typically require acquiring a 360-degree surround multi-view image sequence of the target object, increasing the number of images to improve reconstruction accuracy. This type of method is highly dependent on scene feature points and is often applied to targets with significant geometric features, such as buildings or sculptures. It maintains high accuracy when dealing with buildings with obvious spatial surface transitions.

[0004] However, existing visual reconstruction techniques fall short when dealing with asphalt pavements. Asphalt pavements exhibit strong weak texture characteristics, with a uniform surface color, high repetition of granular structures, and sparse yet highly similar texture features. This makes it difficult for general motion recovery structure algorithms to extract sufficient and unique feature points. In traditional methods, feature point matching employs a global search strategy. Due to the high repetition of texture features on asphalt pavements, the algorithm is prone to confusion in similar regions, generating a large number of mismatched point pairs. This leads to failure in solving the fundamental matrix or a significant decrease in the accuracy of 3D reconstruction, and may even result in holes or distortions in the reconstructed model.

[0005] Furthermore, existing high-precision acquisition equipment often requires static fixed-point measurement, making it difficult to acquire data in real time while the vehicle is in motion. Existing vehicle-mounted dynamic acquisition solutions are often limited by motion blur and camera shake. When using a conventional rolling shutter camera for dynamic shooting, the rapid movement of the road surface relative to the camera during exposure produces rolling shutter effect and image ghosting, severely impacting the quality of subsequent feature point extraction. Simultaneously, the bumps and vibrations of the vehicle during travel cause uncontrollable changes in camera posture, resulting in geometric distortions such as bending and tilting in the reconstructed point cloud model, making it unsuitable for direct and accurate road texture depth measurement.

[0006] In summary, existing technologies struggle to balance detection efficiency and model quality, failing to meet the requirements of rapid detection without disrupting traffic flow, while also falling short of millimeter-level reconstruction accuracy. Therefore, a high-precision 3D reconstruction method is urgently needed that can adapt to high-speed vehicle environments and effectively overcome the defects of weak asphalt texture. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for three-dimensional reconstruction of asphalt pavement texture. By employing short-spacing narrow-baseline continuous acquisition, frequency domain low-pass filtering convolution enhancement, neighborhood-limited SIFT feature matching, bundle adjustment optimization, and dual noise reduction combined with projection transformation correction, this method achieves rapid and high-precision three-dimensional modeling of weakly textured asphalt pavement during vehicle travel, solving the problems of low detection efficiency and difficult feature matching in traditional methods.

[0008] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0009] A method for three-dimensional reconstruction of asphalt pavement texture includes the following steps:

[0010] Step S1: Calibrate the industrial camera to obtain the camera intrinsic parameter matrix and distortion coefficients. Use the industrial camera to continuously acquire asphalt pavement images on a moving vehicle at a preset time interval. The preset time interval makes the shooting position distance between adjacent images 1 to 2 centimeters, thereby obtaining a continuous digital image sequence.

[0011] Step S2: Perform image cropping, Gaussian kernel convolution noise reduction, and low-pass filtering convolution enhancement on each image in the continuous digital image sequence to obtain a preprocessed image sequence. Based on the adjacent images in the preprocessed image sequence, use the scale-invariant feature transformation algorithm to detect feature points. In the second image, limit the neighborhood search range to match the feature points of the first image. The neighborhood search range is the strip-shaped region within a preset pixel range above and below the corresponding point in the second image to obtain a set of matching point pairs.

[0012] Step S3: Solve the fundamental matrix based on the set of matching point pairs, derive the essential matrix in combination with the camera intrinsic parameter matrix, decompose the essential matrix to obtain the camera pose parameters, optimize the camera pose parameters and 3D point coordinates through the bundle adjustment method to obtain the initial 3D point cloud, and perform clustering denoising and statistical filtering denoising on the initial 3D point cloud in sequence to obtain the denoised 3D point cloud.

[0013] Step S4: Perform surface fitting to correct the bending deformation of the denoised 3D point cloud, extract the contour vertices of the corrected point cloud, solve the projection transformation matrix based on the contour vertices and apply it to the denoised 3D point cloud to obtain a tilted corrected point cloud, calculate the scaling adjustment coefficient based on the preset target side length and the contour side length of the tilted corrected point cloud, multiply the coordinates of each point in the tilted corrected point cloud by the scaling adjustment coefficient to obtain a standardized road surface texture 3D model.

[0014] Further, in step S3, the specific process of solving the essential matrix based on the set of matching point pairs includes: using a random sampling consensus algorithm to remove erroneous matching points from the set of matching point pairs to obtain a set of valid matching point pairs; and using an 8-point algorithm to solve the fundamental matrix based on the set of valid matching point pairs. Based on the aforementioned fundamental matrix and the camera intrinsic parameter matrix Derivation of the essential matrix The derivation formula is as follows:

[0015]

[0016] in, The camera intrinsic parameter matrix The transpose of .

[0017] Further, in step S1, the camera intrinsic parameter matrix The properties that reflect the camera itself are in the following form:

[0018]

[0019] in, The focal length of the industrial camera is in millimeters. express Number of pixels per millimeter in the direction. Use pixels to describe The length of the focal length along the axial direction, Use pixels to describe The length of the focal length in the direction, and This indicates the actual location of the principal point, and its unit is pixels.

[0020] Further, in step S3, the essential matrix is ​​decomposed. The specific process of obtaining camera pose parameters includes: processing the essential matrix. Perform singular value decomposition, the essential matrix It can be written as:

[0021]

[0022] in, and All orthogonal matrix, It is a diagonal matrix. The diagonal elements are respectively , , diagonal matrix, The singular values ​​of the diagonal matrix; for the diagonal matrix Make corrections to make it conform. Requirements;

[0023] Based on the singular value decomposition results, four candidate camera poses are obtained, and the projection matrices corresponding to the four candidate camera poses are as follows:

[0024]

[0025] in, and Given two candidate rotation matrices, and There are two candidate translation matrices; among the four candidate camera poses, only one is correct. Based on the constraint that the scene point is simultaneously in front of the focal planes of the two cameras, the unique correct camera pose parameter is selected.

[0026] Furthermore, in step S3, the specific process of solving for the 3D point coordinates using the camera pose parameters includes: given the corresponding image points between two images and the camera poses of the two images, constructing the back projection equation:

[0027]

[0028] in, , , The first The first row vector, second row vector, and third row vector of the camera projection matrix , , The first The first row vector, second row vector, and third row vector of the camera projection matrix , The first The x and y coordinates of the pixel points of a feature point in an image. , The first The x and y coordinates of the corresponding feature points in the image. The coordinates of a three-dimensional point in the world coordinate system. and Number the camera; perform singular value decomposition on the back projection equation to obtain the coordinates of the three-dimensional points in space. .

[0029] Further, in step S2, the specific process of the scale-invariant feature transformation algorithm for detecting feature points includes: constructing a scale space through Gaussian blur, wherein the scale space... Using Gaussian function Compared with the original image The convolution is obtained as follows: (Equation 2-6) where, and These are the x and y coordinates in the image coordinate system, respectively. This represents the convolution operation. The standard deviation of the Gaussian blur is used; extreme points are detected in the scale space using the difference of Gaussians operator, and the location, scale, and orientation information of key points are extracted; the key points are collected. Gradient magnitude of pixels within the neighborhood window and direction They are respectively:

[0030]

[0031] A feature vector descriptor is established based on the neighborhood gradient direction of the key point, and the feature point and the corresponding feature vector descriptor are output.

[0032] Further, in step S2, the preset pixel range is 80 to 120 pixels; the process of obtaining the matching point pair set includes: denoting the key point descriptor in the first image as... The candidate keypoint descriptors within the neighborhood search range in the second image are: Using Euclidean distance As a criterion for similarity:

[0033]

[0034] in, For the first image 128-dimensional feature vectors of key points For the second image, the first The 128-dimensional feature vectors of candidate key points and The eigenvectors are respectively the th One portion, Number the key points. The feature vector dimension index is used; feature points are considered successfully paired when the following conditions are met:

[0035]

[0036] in, A preset similarity threshold is set; candidate key points with the smallest Euclidean distance that is less than the preset similarity threshold are selected as matching points to obtain the set of matching point pairs. 8. The three-dimensional reconstruction method for asphalt pavement texture according to claim 1, characterized in that, in step S3, the specific process of optimizing the camera pose parameters and three-dimensional point coordinates by bundle adjustment includes: setting an objective function. The least squares sum of reprojection errors, for including Images, each image includes In a scenario with n feature points, the objective function is expressed as:

[0037]

[0038] in, , , , express A set of three-dimensional points Indicates the first A three-dimensional point, express The rotation matrix in the camera extrinsic parameters Indicates the first Rotation matrix of each camera, express Translation matrix in the extrinsic parameters of a camera Indicates the first Translation matrix of each camera, Indicates the first The intrinsic parameter matrix of each camera, For the first The three-dimensional point at the th t Predicted image point coordinates under each camera For the first The first image The detection pixel coordinates of each feature point and These are the x-axis and y-axis, respectively. For the first Zhang Image No. Reprojection error of each feature point Number the images. Number the feature points; minimize the objective function The camera pose parameters and the three-dimensional point coordinates are optimized.

[0039] Further, in step S2, the specific process of low-pass filtering convolution enhancement includes: performing a Fourier transform on the denoised image to convert the image from the spatial domain to the frequency domain; centering the frequency domain image to obtain a centered frequency domain image; performing 8 to 12 convolution operations on the centered frequency domain image using a low-pass filter to obtain a filtered frequency domain image; and performing an inverse Fourier transform on the filtered frequency domain image to convert it back to the spatial domain to obtain the enhanced image.

[0040] Further, in step S3, the clustering denoising uses the kdtree clustering algorithm, setting the clustering distance threshold to 0.5 to 2. Points whose distance to each other is less than the clustering distance threshold are divided into the same cluster. The cluster with the most points is retained as the main point cloud, and the remaining discrete clusters are removed to obtain a first-stage denoised point cloud. The statistical filtering denoising includes calculating the average distance from each point in the first-stage denoised point cloud to its neighboring points. and standard deviation Remove those with an average distance greater than point.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] (1) This invention constructs a narrow baseline image sequence by continuously acquiring asphalt pavement images at preset time intervals on a moving vehicle, controlling the spacing between adjacent image capture positions to be 1 to 2 cm, thus forming a visual constraint relationship with extremely high overlap. This short-interval acquisition method is similar to performing high-density scanning of the road surface, ensuring that the corresponding feature points between adjacent frames only undergo slight spatial shifts, thereby limiting the possible matching positions of feature points to a narrow range. This method effectively solves the global matching confusion problem caused by the repetitive and singular texture of asphalt pavement, significantly reduces the search space of the algorithm through physical constraints, and achieves continuous tracking and high-precision texture restoration of weakly textured pavement under the premise that the vehicle is driving normally and does not obstruct traffic.

[0043] (2) This invention employs a method of converting the denoised image from the spatial domain to the frequency domain using Fourier transform, then centers the frequency domain image and performs 8 to 12 convolution operations on the centered frequency domain image using a low-pass filter, and finally converts it back to the spatial domain using inverse Fourier transform. This frequency domain convolution enhancement method can suppress high-frequency noise while enhancing low-frequency texture information, significantly improving motion blur caused by slight shaking or insufficient lighting during vehicle dynamic shooting. This method is equivalent to deep cleaning and sharpening of the image, making the originally indistinct edges and unevenness of asphalt particles clearly distinguishable after enhancement, providing a high-quality feature extraction foundation for the SIFT algorithm, and effectively improving the number of feature point detections and descriptor discrimination in weak texture scenes.

[0044] (3) This invention limits the neighborhood search range in the second image to match the feature points of the first image, setting the neighborhood search range to a strip-shaped area within 80 to 120 pixels above and below the corresponding point in the second image. This neighborhood-limited matching method, combined with the spatial constraints established by narrow baseline acquisition, reduces the global search that originally needed to be performed in the entire image to a local strip-shaped area, greatly reducing the computational complexity of the algorithm. This method is like setting a search guide for the feature matching algorithm, so that the algorithm is no longer interfered with by a large number of similar asphalt particles in the image, fundamentally avoiding cross-regional mismatch caused by texture repetition, significantly improving the matching success rate and matching accuracy, and ensuring the geometric consistency of subsequent 3D reconstruction.

[0045] (4) This invention achieves dual cleaning of the initial 3D point cloud by using the kdtree clustering algorithm to set a clustering distance threshold for primary clustering noise reduction, and by using statistical filtering to calculate the average distance and standard deviation of each point to its neighboring points for secondary statistical filtering noise reduction. This dual noise reduction method first uses spatial clustering to quickly remove isolated noise clusters far from the main point cloud group, and then uses statistical characteristics to finely filter the remaining suspended noise points to ensure the smoothness of the point cloud surface. In addition, this invention completely solves the problem of point cloud bending caused by vehicle bumps, the problem of tilt caused by non-parallel camera installation, and the problem of scale uncertainty inherent in monocular vision by using surface fitting to correct the bending deformation of the denoised 3D point cloud, and solving the projection transformation matrix based on the contour vertices and applying it to the point cloud for tilt correction, combined with the method of scaling adjustment coefficient calculated based on the preset target side length and contour side length. This series of post-processing techniques refines the originally coarse and distorted initial point cloud into a smooth, clean, and dimensionally accurate standardized 3D model. The output model can be directly used for engineering-grade precision texture depth calculation and road surface skid resistance performance evaluation. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of the overall process of the method of the present invention;

[0047] Figure 2 This is a schematic diagram illustrating the principle of the narrow baseline image acquisition method in this invention;

[0048] Figure 3 This is a comparison of the effects of Gaussian kernel convolution noise reduction on images in this invention;

[0049] Figure 4 This is a schematic diagram illustrating the process of performing Fourier transform and frequency domain enhancement on images in this invention.

[0050] Figure 5 This is a schematic diagram illustrating the effect of using a low-pass filter for convolution operations in this invention;

[0051] Figure 6 This is a schematic diagram illustrating the effect of projection transformation on the denoised 3D point cloud in this invention.

[0052] Figure 7 This is a schematic diagram showing the scale comparison of the standardized model before and after the scale adjustment in this invention.

[0053] Figure 8 This is a schematic diagram of the camera calibration parameter output results in this invention;

[0054] Figure 9 This is a schematic diagram of the interior and exterior points of the random sampling consensus algorithm in this invention;

[0055] Figure 10 This is a schematic diagram of the poses of the four candidate cameras after the essential matrix decomposition in this invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. Unless otherwise specifically stated, the relative arrangement, expressions, and values ​​of components and steps set forth in these embodiments do not limit the scope of the present invention. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0057] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0058] like Figure 1 As shown, this invention discloses a method for three-dimensional reconstruction of asphalt pavement texture, comprising the following steps:

[0059] Step S1: Calibrate the industrial camera to obtain the camera intrinsic parameter matrix and distortion coefficients. Use the industrial camera to continuously acquire asphalt pavement images on a moving vehicle at preset time intervals. The preset time intervals make the shooting position spacing between adjacent images 1 to 2 centimeters, thereby obtaining a continuous digital image sequence.

[0060] A camera's imaging system comprises four coordinate systems: world coordinates, camera coordinates, image coordinates, and pixel coordinates. The coordinates of a point in the world coordinate system are transformed into the camera coordinate system using an extrinsic parameter matrix through rigid body transformation. Then, perspective projection and affine transformation are performed using an intrinsic parameter matrix to finally obtain the coordinates in the pixel coordinate system. Based on the pixel coordinates of two points in the image, the pixel distance between them can be determined. By reversing the above process, the positions of the two points in the world coordinate system can be obtained.

[0061] This embodiment uses Zhang Zhengyou's checkerboard calibration method for camera calibration. The regularity and high color contrast of the checkerboard pattern make it easy to extract the pixel coordinates of the corner points. The initial values ​​of the camera's intrinsic and extrinsic parameters are calculated using the homography matrix, and then the camera calibration is completed after coordinate optimization.

[0062] During calibration, approximately 20 checkerboard images need to be taken from different angles. In this embodiment, the checkerboard square size is 30 mm × 30 mm. The camera used should be consistent with the industrial camera used for subsequent texture image acquisition. Open the Camera Calibrator tool in MATLAB, click the Add Images button and import the captured checkerboard images. Enter the actual size of the squares in the Size of checkboard square field. After importing the images, the computer will automatically filter out images suitable for calibration, marking the origin, detected corner points, and coordinate axis directions.

[0063] For the captured chessboard image, detect all corner points where black and white squares intersect, including red, green, blue, and yellow points, and define the printed chessboard drawing as being located in the world coordinate system. On the plane, the yellow point in the upper left corner is designated as the origin. Since the spatial coordinates of all corner points in the world coordinate system of the checkerboard calibration drawing are known, the pixel coordinates of these corner points in the calibration image are also known. Through such matching point pairs, the homography matrix in the transformation process is obtained according to the optimization method.

[0064] By inputting the matching point pairs and specifying the calculation method, the corresponding homography matrix can be obtained.

[0065] When calibrating, select the radial distortion coefficient and click Calibrate to start the calibration. During the calibration process, a Mean Error in Pixels will be generated. To keep the error within a reasonable range, delete histograms with pixel errors greater than 0.6, and their corresponding checkerboard images will also be removed from the image set. Use the remaining images for recalibration to obtain more accurate parameters.

[0066] After camera calibration is completed, the camera and calibration board poses will be displayed in two ways. One shows the relative pose of the camera moving while the calibration board remains in place, and the other shows the relative pose of the camera moving while the calibration board remains in place. Both shooting methods can be used for subsequent calibration.

[0067] like Figure 8 As shown, after calibration, the camera parameters are output in the MATLAB workspace, where RadialDistortion is the camera distortion matrix and IntrinsicMatrix is ​​the camera intrinsic parameter matrix, and are saved as a .mat file for later use.

[0068] In this embodiment, the camera intrinsic parameter matrix The properties of the camera itself are shown in Equation 2-2:

[0069]

[0070] in The focal length of the industrial camera is in millimeters. express Number of pixels per millimeter in the direction. Use pixels to describe The length of the focal length along the axial direction, Use pixels to describe The length of the focal length in the direction, and This indicates the actual location of the principal point, and its unit is pixels.

[0071] This embodiment uses a Hikvision MVS-CS050-60UM monochrome industrial camera with a CMOS global shutter and a resolution of 2448×2048. The lens is a Hikvision MVL-MF1628M-8MP. To reduce shadows caused by occlusion of prominent textures, an additional fill light is required. The fill light source is a Hikvision MVS-LEBS-H-300-200-W surface light source. For detailed parameters of the camera, lens, and light source, please refer to [link to relevant documentation].

[0072] Table 2-1 Camera Lens and Light Source Related Parameters

[0073]

[0074] Unlike the progressive exposure mode of the rolling shutter, the global shutter exposes all pixels simultaneously by collecting light at the same time. While the rolling shutter can achieve higher frame rates, it suffers from partial exposure and blurring when objects are moving quickly, severely impacting image quality. In contrast, although the global shutter increases readout noise, it significantly shortens the exposure time, making it more suitable for photographing moving objects. Therefore, this embodiment uses the global shutter to acquire images during the movement of the test vehicle.

[0075] like Figure 2 As shown, during the data acquisition process, the test vehicle's speed was maintained at 30 kilometers per hour. The industrial camera was set to continuous image capture mode, with a time interval of 1 millisecond. A continuous set of images was acquired with a movement distance of approximately 1 centimeter. In this embodiment, two roads, Liangjiang North Road and Central Avenue, Jiangning District, Nanjing City, Jiangsu Province, were selected. A measurement point was selected every 5 meters along the wheel tracks of the vehicles, for a total of 20 measurement points for image acquisition. Seven sets of images were acquired at each measurement point, with five consecutive images in each set, resulting in a total of 35 images per measurement point. A total of 700 images were acquired for the tested road section.

[0076] Both Liangjiang North Road and Central Avenue are two-lane roads in both directions, with a design speed of 40 kilometers per hour and an average daily traffic volume of approximately 5,000 vehicles, mainly small vehicles. The surface layer of the asphalt pavement uses AC-13 type asphalt mixture. Central Avenue has lower traffic volume and no significant rutting on the road surface. Liangjiang North Road, due to its connection to construction sites, experiences a significantly higher frequency of dump truck traffic than Central Avenue, resulting in some edge damage caused by dump trucks. Central Avenue represents a typical design for low-traffic roads in ordinary cities, while Liangjiang North Road reflects common problems encountered by industrial zone side roads under occasional heavy loads, and has broad reference value.

[0077] During the data acquisition process, the short acquisition intervals necessitate avoiding image blurring, which requires a high camera frame rate and very short exposure times. To ensure better acquisition results, additional light sources were added during the acquisition process. This reduces the impact of subtle shadows and allows the camera more flexibility in setting the exposure time.

[0078] The above method dynamically acquires asphalt pavement texture images during vehicle movement, avoiding traffic flow interruptions caused by fixed-point measurements. However, the acquired images cannot be directly used for 3D texture reconstruction; digital image preprocessing is required to ensure effective identification of feature points and accurate reconstruction of the 3D model.

[0079] Step S2: Perform image cropping, Gaussian kernel convolution noise reduction, and low-pass filtering convolution enhancement on each image in the continuous digital image sequence to obtain a preprocessed image sequence. Based on the adjacent images in the preprocessed image sequence, use the scale-invariant feature transformation algorithm to detect feature points. In the second image, limit the neighborhood search range to match the feature points of the first image. The neighborhood search range is the strip-shaped region within a preset pixel range above and below the corresponding point in the second image to obtain a set of matching point pairs.

[0080] Due to boundary effects, most images captured during vehicle movement exhibit varying degrees of edge distortion, resulting in blurred textures. Furthermore, road markings are often included in the images during acquisition, and the markings and asphalt pavement exhibit significant grayscale differences. During subsequent reconstruction, image feature points tend to concentrate on the boundary between the markings and the asphalt pavement, leading to uneven feature point distribution and reducing the integrity of the 3D reconstruction. Therefore, it is necessary to crop and segment the acquired images to ensure high-quality reconstruction. The original image with 2448 x 2048 pixels is cropped to 2000 x 1646 pixels, maintaining the same aspect ratio as the original image.

[0081] like Figure 3 As shown, while the supplementary lighting can improve the noise of dark current to some extent, the sensor may still generate charge transfer noise when reading the signal. Therefore, Gaussian noise reduction is still needed for the image. Gaussian noise reduction is very effective in suppressing Gaussian noise that follows a normal distribution. Essentially, it uses a Gaussian kernel generated by a Gaussian function to perform a convolution operation on the image. That is, for each pixel, the average gray value of all pixels within a specified Gaussian radius is taken as the gray value of the center point. The texture image after Gaussian noise reduction shows that white noise points have been filtered out after Gaussian filtering.

[0082] Gaussian noise reduction can be achieved by calling a filter in MATLAB. Here, a Gaussian kernel of size 5×5 is selected, and filtering can be completed using imfilter.

[0083] like Figure 4 As shown, this embodiment uses Fourier transform to enhance the image. The specific process of low-pass filtering convolution enhancement includes: performing a Fourier transform on the denoised image to convert the image from the spatial domain to the frequency domain; centering the frequency domain image to obtain a centered frequency domain image; performing 8 to 12 convolution operations on the centered frequency domain image using a low-pass filter to obtain a filtered frequency domain image; and performing an inverse Fourier transform on the filtered frequency domain image to convert it back to the spatial domain to obtain the enhanced image.

[0084] Fourier transform, a mathematical method for converting signals from the time or spatial domain to the frequency domain, is widely used in image processing. By filtering an image in the frequency domain, certain frequency components can be enhanced or suppressed. Finally, the image is converted back to the spatial domain through inverse Fourier transform, achieving image enhancement. Images acquired during dynamic processes may exhibit motion blur due to insufficient lighting, or feature points may be difficult to identify due to limited grayscale range. Therefore, this embodiment performs Fourier transform on the acquired original image, and performs low-pass filtering and multiple convolution processes in the frequency domain to maximize the reduction of motion blur and enhance image contrast.

[0085] The spectral information of an image primarily characterizes the rate of grayscale change. Bright spots in the spectrogram reflect the low-frequency information of the image, i.e., the smooth parts of the image. After the original image undergoes a Fourier transform, the low-frequency centers are scattered in the four corners. For subsequent frequency domain processing, the spectrum needs to be centered. By frequency shifting, the bright spots are moved to the center of the spectrogram, at which point the DC component is concentrated in the center of the spectrogram.

[0086] Low-pass filtering allows low-frequency information to pass through while blocking high-frequency information such as noise. Therefore, low-pass filtering is necessary to enhance an image. Figure 5 As shown, the image undergoes multiple low-pass filtering convolutions. The spatial gradient metric is used to evaluate the image sharpness after each convolution.

[0087] Table 2-2 Spatial gradient indices of images under different convolution orders

[0088]

[0089] A larger spatial gradient indicates a clearer image. While multiple convolutions do improve image enhancement and reduce motion blur during dynamic shooting, more convolutions do not necessarily mean better results. When the number of convolutions exceeds 15, significant distortion occurs at the bottom of the texture image. Therefore, to ensure high texture image quality, the low-pass filter should be used for around 10 convolutions to achieve the best image enhancement effect.

[0090] Image feature matching consists of three main stages: image feature point detection, feature point description, and feature point matching. Image feature point detection locates salient key points in an image by detecting edge response functions or scale-space extrema. Feature point description uses mathematical vectors to systematically describe the extracted key points. Feature point matching calculates the distance between feature vectors; point pairs within a certain distance limit are considered a successful match.

[0091] SIFT corner detection, Harris corner detection, SURF corner detection, and energy-based corner detection were performed on the same image respectively. The results are shown in [the table below].

[0092] Table 2-3 Number of Feature Points Detected by Different Methods

[0093]

[0094] During feature matching of asphalt pavements, a large number of feature points fail to meet the matching threshold and are therefore discarded. When the initial feature point density is too low, the 3D point cloud coverage decreases, failing to meet the requirements for reconstruction integrity. Therefore, the SIFT corner detection method, which detects the most corner points, should be selected for texture feature detection to ensure smooth subsequent feature matching.

[0095] In this embodiment, the specific process of the scale-invariant feature transform algorithm for detecting feature points includes: constructing a scale space through Gaussian blur, wherein the scale space... Using Gaussian function Compared with the original image The convolution is obtained as follows:

[0096]

[0097] Equation 2-6, where and These are the x and y coordinates in the image coordinate system, respectively. This represents the convolution operation. The standard deviation of the Gaussian blur is given. The SIFT algorithm finds key points in different scale spaces. Obtaining the scale space relies on Gaussian blur. By convolving the Gaussian template matrix with the original image and then normalizing it, the Gaussian blurred image of the original image can be obtained. In the scale space, extreme points are detected using the difference of Gaussians operator, and the location, scale, and orientation information of the key points are extracted.

[0098] To ensure the descriptor's rotation invariance, a reference orientation is assigned to each keypoint using local image features. These keypoints are then acquired. Gradient magnitude of pixels within the neighborhood window and direction They are respectively:

[0099]

[0100] Through the above steps, each keypoint possesses three pieces of information: location, scale, and orientation. A vector descriptor is created for each keypoint; this vector is an abstraction of the region image information and is unique, indicating that the description of the keypoint is unique. A feature vector descriptor is then established based on the neighborhood gradient direction of the keypoint, and the feature point and its corresponding feature vector descriptor are output.

[0101] Feature point detection relies on the similarity of feature vectors between two images. For general buildings, the feature vectors of feature points differ significantly, so setting a low similarity threshold can achieve relatively accurate matching. However, for weak textures like asphalt, a low similarity threshold will result in large-scale mismatches. If there are too many incorrectly matched point pairs, it will cause chaos in the establishment of the basic reconstruction framework, leading to the failure of reconstruction due to the inability to find basic feature point pairs. Conversely, a high similarity threshold will result in too few matched point pairs, making it impossible to form an effective basic reconstruction framework.

[0102] To address the threshold setting issue, this embodiment employs a neighborhood cross-correlation method. The neighborhood cross-correlation algorithm constrains the feature point search range. Specifically, it selects the feature points to be matched in the first image and searches within a band of ±100 pixels of the corresponding points in the second image. This allows for the matching of as many feature points as possible while setting a relatively low similarity threshold, thereby improving the matching success rate.

[0103] In this embodiment, the preset pixel range is 80 to 120 pixels, preferably 100 pixels. The process of obtaining the matching point pair set includes: denoting the key point descriptor in the first image as... The candidate keypoint descriptors within the neighborhood search range in the second image are: Using Euclidean distance As a criterion for similarity:

[0104]

[0105] Equation 2-8, where For the first image 128-dimensional feature vectors of key points For the second image, the first The 128-dimensional feature vectors of candidate key points and The eigenvectors are respectively the th One portion, Number the key points. The feature vector dimension is indexed. Feature points are considered successfully paired when the following conditions are met:

[0106]

[0107] Equation 2-9, where To obtain the set of matching point pairs, candidate key points with the smallest Euclidean distance that is less than the preset similarity threshold are selected as matching points.

[0108] Without specifying the search range, many erroneous matches occur during feature point matching. However, by limiting the search range for matching points, the matching accuracy is significantly improved. Accurate feature point matching is a prerequisite for subsequent 3D reconstruction. This embodiment ensures the quantity and accuracy of matches through small-motion image acquisition and neighborhood cross-correlation, laying a solid foundation for subsequent 3D reconstruction.

[0109] Step S3: Solve the fundamental matrix based on the set of matching point pairs, derive the essential matrix in combination with the camera intrinsic parameter matrix, decompose the essential matrix to obtain the camera pose parameters, optimize the camera pose parameters and 3D point coordinates through bundle adjustment to obtain the initial 3D point cloud, and perform clustering denoising and statistical filtering denoising on the initial 3D point cloud in sequence to obtain the denoised 3D point cloud.

[0110] In the process of obtaining the fundamental matrix, the 8-point algorithm is usually used. However, in actual use, erroneous matching points have a significant impact on matrix estimation. Therefore, a random sampling consensus algorithm needs to be introduced during the calculation to eliminate erroneous matches and obtain the fundamental matrix that can successfully verify the maximum number of point pairs.

[0111] like Figure 9 As shown, the RANSAC algorithm, also known as the Random Sample Consensus algorithm, has the important function of correctly estimating mathematical model parameters from a set of data containing outliers. It uses the fundamental matrix obtained by the 8-point method to test other data points. If a point also applies the projection relationship of the matrix, it is considered an inlier. When there are enough inliers, a reasonable fundamental matrix can be obtained. The matrix is ​​then re-estimated using all the assumed inliers, and the matrix is ​​continuously optimized.

[0112] In this embodiment, step S3, the specific process of solving the essential matrix based on the set of matching point pairs includes: using a random sampling consensus algorithm to remove erroneous matching points from the set of matching point pairs to obtain a set of valid matching point pairs, and using an 8-point algorithm to solve the fundamental matrix based on the set of valid matching point pairs. Based on the fundamental matrix and the camera intrinsic parameter matrix Derivation of the essential matrix The derivation formula is as follows:

[0113]

[0114] Equation 2-1, where The camera intrinsic parameter matrix The transpose of .

[0115] In step S3, the essential matrix is ​​decomposed. The specific process of obtaining camera pose parameters includes: processing the essential matrix. Perform singular value decomposition, the essential matrix It can be written as:

[0116]

[0117] Equation 2-3, where and All orthogonal matrix, It is a diagonal matrix. The diagonal elements are respectively , , diagonal matrix, These are the singular values ​​of the diagonal matrix. For the diagonal matrix... Make corrections to make it conform. The requirements are as follows. Since the projection spaces differ by a scale factor, four camera poses can be decomposed from the essential matrix. Based on the singular value decomposition results, four candidate camera poses are obtained, and the projection matrices corresponding to the four candidate camera poses are as follows:

[0118]

[0119]

[0120] Equation 2-4, where and Given two candidate rotation matrices, and These are two candidate translation matrices.

[0121] like Figure 10 As shown, only one of the four candidate camera poses is correct. The essential matrix after RANSAC filtering is characterized by having the largest number of inliers. These inliers are characterized by their corresponding scene points being simultaneously in front of the focal planes of both cameras. Based on the constraint that the scene points are simultaneously in front of the focal planes of both cameras, the unique correct camera pose parameters are selected.

[0122] In step S3, the specific process of solving for the 3D point coordinates using the camera pose parameters includes: given the corresponding image points between two images and the camera poses of the two images, constructing the back projection equation:

[0123]

[0124] Equation 2-5, where , , The first The first row vector, second row vector, and third row vector of the camera projection matrix , , The first The first row vector, second row vector, and third row vector of the camera projection matrix , The first The x and y coordinates of the pixel points of a feature point in an image. , The first The x and y coordinates of the corresponding feature points in the image. The coordinates of a three-dimensional point in the world coordinate system. and The camera is numbered. Singular value decomposition is performed on the back projection equation to obtain the coordinates of the three-dimensional points in space. .

[0125] Bundle adjustment, also known as bundle adjustment, is commonly used to solve for camera pose and the coordinates of 3D points in space. Its core calculation is minimizing reprojection error. Bundle adjustment uses collinearity equations as its mathematical model, treating the observed image plane coordinates of image points as unknowns in a nonlinear function, and then solving for them using the least squares method.

[0126] In step S3, the specific process of optimizing the camera pose parameters and 3D point coordinates using the bundle adjustment method includes: setting the objective function. The least squares sum of reprojection errors, for including Images, each image includes In a scenario with n feature points, the objective function is expressed as:

[0127]

[0128] Equation 2-10, where , , , express A set of three-dimensional points Indicates the first A three-dimensional point, express The rotation matrix in the camera extrinsic parameters Indicates the first Rotation matrix of each camera, express Translation matrix in the extrinsic parameters of a camera Indicates the first Translation matrix of each camera, Indicates the first The intrinsic parameter matrix of each camera, For the first The three-dimensional point at the th t Predicted image point coordinates under each camera and The first The first image The x and y coordinates of the detected pixel coordinates of each feature point. For the first Zhang Image No. Reprojection error of each feature point Number the images. Number the feature points. Minimize the objective function. The camera pose parameters and the three-dimensional point coordinates are optimized.

[0129] In the incremental motion recovery framework, camera pose estimation errors are cumulative, so local bundle adjustment optimization must be performed after adding new images. Without bundle adjustment correction, the uncompensated errors will exceed a threshold, causing the optimization process to diverge and eventually fail to converge. After adding all images, global bundle adjustment correction is finally performed to optimize all cameras and 3D points, ensuring correct camera poses and 3D point coordinates.

[0130] The `bundleAdjustment` function in MATLAB can be used to quickly perform bundle adjustment correction. Enter the appropriate value in `xyzPoints`. The function takes the world coordinates of the three-dimensional points, as well as the coordinates of the image points (tracks), the world coordinates in the camera coordinate system (camPoses), and the camera parameters (cameraParams). The function finally outputs the reconstructed point cloud, camera pose, and reconstruction error.

[0131] Note that there is a small probability of errors during camera pose recovery using bundle adjustment. These errors can include initial pose estimation errors caused by cumulative motion estimation errors exceeding the bundle adjustment convergence radius, and feature matching drift caused by incorrect SIFT feature point matching. These errors can ultimately lead to significant separation in the recovered point cloud. In such cases, it is necessary to delete the camera's information or correct the error by selecting a Region of Interest (ROI) to remove the erroneous point cloud data.

[0132] Although the acquired texture images have been filtered and denoised before 3D reconstruction, the point cloud obtained from the camera sensor may still contain various noise points such as specks, outliers, and isolated points due to surface reflections and internal camera transmission. These noise points can significantly affect the subsequent calculation of texture feature parameters, so it is necessary to denoise the point cloud first.

[0133] In this embodiment, in step S3, the clustering denoising uses the kdtree clustering algorithm, setting the clustering distance threshold to 0.5 to 2, preferably 1. Points whose distance to each other is less than the clustering distance threshold are divided into the same cluster. The cluster with the most points is retained as the main point cloud, and the remaining discrete clusters are removed to obtain a first-stage denoised point cloud. The statistical filtering denoising includes calculating the average distance from each point in the first-stage denoised point cloud to its neighboring points. and standard deviation Remove those with an average distance greater than point.

[0134] Since outliers in the point cloud model reconstructed using small motion recovery structures tend to cluster and are far from the normal point cloud group, kdtree clustering can be used to separate the outliers. Kdtree clustering, also known as Euclidean clustering, uses a unique parameter during the clustering process. , specify When the distance between points is less than When the two points are internally connected, it indicates that they belong to the same class. Using this principle, kd-tree clustering is performed on the point cloud containing noisy points. After kd-tree clustering, the point cloud is displayed in different colors according to its class. Points usable for subsequent reconstruction are concentrated in the green area, while outliers are scattered in the red, yellow, and other areas. Removing the other colors and retaining only the green portion yields the point cloud after the first denoising step.

[0135] The point cloud segmented by kdtree clustering has effectively removed outliers that are far from the main point cloud group, but some noise points still remain. These noise points can easily form irregular sharp protrusions during the subsequent point cloud triangulation process, affecting the smoothness of the model surface. Therefore, further secondary noise reduction processing is needed to improve the point cloud quality.

[0136] After clustering and segmentation, the point clouds are relatively close together. Therefore, using Statistical Oriented Filtering (SOR) at this stage will not fail to filter out outliers due to excessive distance between point clouds. A second SOR denoising process is then applied to the point cloud. After these two denoising steps, a large number of outliers and noise points in the original point cloud are largely filtered out, resulting in a high-quality point cloud that lays the foundation for subsequent textured 3D model reconstruction. Following this second denoising process, noise points in the point cloud data are further removed, resulting in smoother surface undulations in the textured model and effectively improving the overall quality and visual realism of the model.

[0137] Step S4: Perform surface fitting to correct the bending deformation of the denoised 3D point cloud, extract the contour vertices of the corrected point cloud, solve the projection transformation matrix based on the contour vertices and apply it to the denoised 3D point cloud to obtain a tilted corrected point cloud, calculate the scaling adjustment coefficient based on the preset target side length and the contour side length of the tilted corrected point cloud, multiply the coordinates of each point in the tilted corrected point cloud by the scaling adjustment coefficient to obtain a standardized road surface texture 3D model.

[0138] During dynamic data acquisition, due to vehicle vibrations, the camera cannot maintain perfect parallelism with the ground, resulting in the reconstructed model often appearing flat or uneven. The plane has a certain angle of inclination, requiring tilt correction of the point cloud model to make it parallel to the ground plane. However, simply projecting the point cloud onto the horizontal plane would project the texture undulations onto the same horizontal plane, losing the original shape of the texture. Therefore, this embodiment combines the principle of tilted image projection correction to perform projection transformation on the point cloud.

[0139] The point cloud curved surface processing employs a combination of surface fitting and error correction. Surface fitting is performed on the curved point cloud set. There is an error between the fitted surface and the surface. This error is the distance from the point cloud to the surface. In MATLAB, the Curve Fitting Toolbox is used to import the point cloud data sequentially. , and coordinate.

[0140] Regarding the choice of fitting method, if interpolation fitting is used, the texture undulations will be completely integrated into the fitted surface. Since the value is always equal to 0, while local linear regression can achieve a good fit, processing a single point cloud takes more than 10 minutes, resulting in very low fitting efficiency. Therefore, this embodiment uses polynomial fitting to... and The degree of each step is set to 2. The surface is obtained by fitting. The surface fitted by the quadratic polynomial can reflect the bending trend of the texture well. Overall, it can accurately capture the main features of surface deformation and ensure that the texture undulation is preserved to the maximum extent while taking efficiency into account.

[0141] Starting with a known curved surface, MATLAB is used to fit the surface and fit a tilted plane representing its tilt angle. The actual coordinates of a point on the point cloud and the corresponding point on the curved surface are then compared. There is a difference in the axial coordinates. Find the point on the surface in the plane. The corresponding points are then used to obtain a point cloud after bending correction. Repeating the above operation for each point cloud yields a point cloud set that is basically parallel to the plane. The point cloud set after bending surface adjustment shows that the point cloud set that originally had obvious upward deformation has been significantly improved, and the overall point cloud distribution tends to be flat.

[0142] Point cloud boundary detection uses the `boundary` function in MATLAB. Partial code is shown in the table below.

[0143] Table 2-4 MATLAB boundary detection code examples

[0144]

[0145] The core of the code is the alpha shape algorithm, and the specific implementation process is based on the radius. The circle rolls around the point cloud, and the trajectory formed is its contour point. The four sides of the texture point cloud are obtained by fitting straight lines. Taking the blue boundary line on the right as an example, the four fitted boundary lines and the point cloud contour display four lines: red, green, purple, and blue. The four edges intersect each other to obtain the four vertices of the outer contour.

[0146] Let the coordinates of the four intersection points be respectively , , , Therefore, the corresponding rectangle can be obtained. , The extreme values ​​of the coordinates of the four points are used to form the coordinates of the four vertices of the target projection rectangle. , , , First, the transformation moment rot matrix between the two rectangles is calculated. The result of the projection transformation is a bright green point cloud. After the projection transformation, the boundary of the bright green point cloud is basically parallel to the coordinate axis. This phenomenon shows that the pose adjustment algorithm used can effectively correct the spatial orientation of the point cloud.

[0147] like Figure 6 As shown, referring to the transformation method of the boundary point cloud, the transformation matrix is ​​applied to every point in the point cloud set. The resulting projected point cloud is roughly parallel to the ground while still retaining the texture undulations. This phenomenon demonstrates that the adopted projection transformation method can effectively correct the spatial orientation of the point cloud, not only eliminating the rotational offset of the original point cloud but also maintaining the geometric consistency of the internal structure of the point cloud, thus laying a good foundation for subsequent accurate registration and 3D reconstruction.

[0148] In the process of small-motion recovery and structural reconstruction, only the relative geometry of the scene can usually be recovered, and absolute scale information cannot be directly obtained. Due to the lack of external references, the reconstructed point cloud model may have an incorrect scale factor estimation during coordinate transformation, resulting in a discrepancy between the model size and the actual scene, thus causing scale distortion. Therefore, it is necessary to adjust the scale of the asphalt pavement surface texture model.

[0149] The process of adjusting the scale of the asphalt pavement surface texture model is similar to the projection transformation process of a 3D point cloud. Given that the target side length of the model is fixed at 150 mm, the side length of the current point cloud model outline is obtained by the difference of the coordinates of the four vertices in the model. Then, the scale adjustment coefficient is calculated. By multiplying the points in the model by the scale adjustment coefficient, the scale-adjusted texture model can be obtained.

[0150] like Figure 7 As shown, before scaling, the texture model had a side length of only 70 mm. After scaling, the model was successfully enlarged to the expected side length of 150 mm, achieving precise size matching while preserving the original geometric features and texture details. This adjustment process effectively solved the problem of inconsistency between the original reconstructed model and the actual size, providing a standardized model for subsequent work.

[0151] Twenty measuring points at Liangjiang North Road and Central Avenue were reconstructed. The three-dimensional model of each measuring point can clearly reflect the texture undulation of the asphalt pavement. The average reprojection error of the model is 0.24 pixels, the average feature point matching success rate reaches 89%, the model scale error is less than 2%, and the error after tilt angle correction is less than 0.5 degrees. It can be directly used to calculate the macroscopic texture feature parameters of the pavement.

[0152] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for three-dimensional reconstruction of asphalt pavement texture, characterized in that, Includes the following steps: Step S1: Calibrate the industrial camera to obtain the camera intrinsic parameter matrix and distortion coefficients. Use the industrial camera to continuously acquire asphalt pavement images on a moving vehicle at a preset time interval. The preset time interval makes the shooting position distance between adjacent images 1 to 2 centimeters, thereby obtaining a continuous digital image sequence. Step S2: Perform image cropping, Gaussian kernel convolution noise reduction, and low-pass filtering convolution enhancement on each image in the continuous digital image sequence to obtain a preprocessed image sequence. Based on the adjacent images in the preprocessed image sequence, use the scale-invariant feature transformation algorithm to detect feature points. In the second image, limit the neighborhood search range to match the feature points of the first image. The neighborhood search range is the strip-shaped region within a preset pixel range above and below the corresponding point in the second image to obtain a set of matching point pairs. Step S3: Solve the fundamental matrix based on the set of matching point pairs, derive the essential matrix in combination with the camera intrinsic parameter matrix, decompose the essential matrix to obtain the camera pose parameters, optimize the camera pose parameters and 3D point coordinates through the bundle adjustment method to obtain the initial 3D point cloud, and perform clustering denoising and statistical filtering denoising on the initial 3D point cloud in sequence to obtain the denoised 3D point cloud. Step S4: Perform surface fitting to correct the bending deformation of the denoised 3D point cloud, extract the contour vertices of the corrected point cloud, solve the projection transformation matrix based on the contour vertices and apply it to the denoised 3D point cloud to obtain a tilted corrected point cloud, calculate the scaling adjustment coefficient based on the preset target side length and the contour side length of the tilted corrected point cloud, multiply the coordinates of each point in the tilted corrected point cloud by the scaling adjustment coefficient to obtain a standardized road surface texture 3D model.

2. The method for three-dimensional reconstruction of asphalt pavement texture according to claim 1, characterized in that, In step S3, the specific process of solving the essential matrix based on the set of matching point pairs includes: using a random sampling consensus algorithm to remove erroneous matching points from the set of matching point pairs to obtain a set of valid matching point pairs; and using an 8-point algorithm to solve the fundamental matrix based on the set of valid matching point pairs. Based on the aforementioned fundamental matrix and the camera intrinsic parameter matrix Derivation of the essential matrix The derivation formula is as follows: (Equation 2-1); in, The camera intrinsic parameter matrix The transpose of .

3. The method for three-dimensional reconstruction of asphalt pavement texture according to claim 1, characterized in that, In step S1, the camera intrinsic parameter matrix The properties that reflect the camera itself are in the following form: (Equation 2-2); in, The focal length of the industrial camera is in millimeters. express Number of pixels per millimeter in the direction. Use pixels to describe The length of the focal length along the axial direction, Use pixels to describe The length of the focal length in the direction, and This indicates the actual location of the principal point, and its unit is pixels.

4. The method for three-dimensional reconstruction of asphalt pavement texture according to claim 2, characterized in that, In step S3, the essential matrix is ​​decomposed. The specific process of obtaining camera pose parameters includes: processing the essential matrix. Perform singular value decomposition, the essential matrix It can be written as: (Equation 2-3); in, and All orthogonal matrix, It is a diagonal matrix. The diagonal elements are respectively , , diagonal matrix, The singular values ​​of the diagonal matrix; for the diagonal matrix Make corrections to make it conform. Requirements; Based on the singular value decomposition results, four candidate camera poses are obtained, and the projection matrices corresponding to the four candidate camera poses are as follows: (Equation 2-4); in, and Given two candidate rotation matrices, and There are two candidate translation matrices; among the four candidate camera poses, only one is correct. Based on the constraint that the scene point is simultaneously in front of the focal planes of the two cameras, the unique correct camera pose parameter is selected.

5. The method for three-dimensional reconstruction of asphalt pavement texture according to claim 4, characterized in that, In step S3, the specific process of solving for the 3D point coordinates using the camera pose parameters includes: given the corresponding image points between two images and the camera poses of the two images, constructing the back projection equation: (Equation 2-5); in, , , The first The first row vector, second row vector, and third row vector of the camera projection matrix , , The first The first row vector, second row vector, and third row vector of the camera projection matrix , The first The x and y coordinates of the pixel points of a feature point in an image. , The first The x and y coordinates of the corresponding feature points in the image. The coordinates of a three-dimensional point in the world coordinate system. and Number the camera; perform singular value decomposition on the back projection equation to obtain the coordinates of the three-dimensional points in space. .

6. The method for three-dimensional reconstruction of asphalt pavement texture according to claim 1, characterized in that, In step S2, the specific process of the scale-invariant feature transformation algorithm for detecting feature points includes: constructing a scale space through Gaussian blur, wherein the scale space... Using Gaussian function Compared with the original image The convolution is obtained as follows: (Equation 2-6) where, and These are the x and y coordinates in the image coordinate system, respectively. This represents the convolution operation. The standard deviation of the Gaussian blur is used; extreme points are detected in the scale space using the difference of Gaussians operator, and the location, scale, and orientation information of key points are extracted; the key points are collected. Gradient magnitude of pixels within the neighborhood window and direction They are respectively: (Equation 2-7a); (Equation 2-7b); A feature vector descriptor is established based on the neighborhood gradient direction of the key point, and the feature point and the corresponding feature vector descriptor are output.

7. The method for three-dimensional reconstruction of asphalt pavement texture according to claim 6, characterized in that, In step S2, the preset pixel range is 80 to 120 pixels; the process of obtaining the matching point pair set includes: denoting the key point descriptor in the first image as... The candidate keypoint descriptors within the neighborhood search range in the second image are: Using Euclidean distance As a criterion for similarity: (Equation 2-8); in, For the first image 128-dimensional feature vectors of key points For the second image, the first The 128-dimensional feature vectors of candidate key points and The eigenvectors are respectively the th One portion, Number the key points. The feature vector dimension index is used; feature points are considered successfully paired when the following conditions are met: (Equation 2-9); in, A preset similarity threshold is set; candidate key points with the smallest Euclidean distance that is less than the preset similarity threshold are selected as matching points to obtain the set of matching point pairs.

8. The method for three-dimensional reconstruction of asphalt pavement texture according to claim 1, characterized in that, In step S3, the specific process of optimizing the camera pose parameters and 3D point coordinates using the bundle adjustment method includes: setting the objective function. The least squares sum of reprojection errors, for including Images, each image includes In a scenario with n feature points, the objective function is expressed as: (Equation 2-10); in, , , , express A set of three-dimensional points Indicates the first A three-dimensional point, express The rotation matrix in the camera extrinsic parameters Indicates the first Rotation matrix of each camera, express Translation matrix in the extrinsic parameters of a camera Indicates the first Translation matrix of each camera, Indicates the first The intrinsic parameter matrix of each camera, For the first The three-dimensional point at the th t Predicted image point coordinates under each camera For the first The first image The detection pixel coordinates of each feature point and These are the x-axis and y-axis, respectively. For the first Zhang Image No. Reprojection error of each feature point Number the images. Number the feature points; minimize the objective function The camera pose parameters and the three-dimensional point coordinates are optimized.

9. The method for three-dimensional reconstruction of asphalt pavement texture according to claim 1, characterized in that, In step S2, the specific process of low-pass filtering convolution enhancement includes: performing a Fourier transform on the denoised image to convert the image from the spatial domain to the frequency domain; centering the frequency domain image to obtain a centered frequency domain image; performing 8 to 12 convolution operations on the centered frequency domain image using a low-pass filter to obtain a filtered frequency domain image; and performing an inverse Fourier transform on the filtered frequency domain image to convert it back to the spatial domain to obtain the enhanced image.

10. The method for three-dimensional reconstruction of asphalt pavement texture according to claim 1, characterized in that, In step S3, the clustering denoising uses the kdtree clustering algorithm, setting the clustering distance threshold to 0.5 to 2. Points whose distance to each other is less than the clustering distance threshold are grouped into the same cluster. The cluster with the most points is retained as the main point cloud, and the remaining discrete clusters are removed to obtain a first-stage denoised point cloud. The statistical filtering denoising includes calculating the average distance from each point in the first-stage denoised point cloud to its neighboring points. and standard deviation Remove those with an average distance greater than point.