Road disease area estimation method based on image transformation
By using an image transformation-based method, combining perspective transformation and anchor frame positioning with the YOLO framework and SAM model, the error problem in calculating the area of road defects was solved, achieving higher accuracy in the estimation of defect area.
Patent Information
- Application Number
- CN202510760824.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies suffer from several drawbacks in calculating the area of road defects. These include large errors when multiple types of defects coexist, calculation results for defects with elongated shapes that far exceed actual values, and image segmentation area errors caused by fixed-angle shooting, all of which affect the accuracy of the calculation.
An image transformation-based method, including perspective transformation and anchor frame localization, is adopted. Combining the YOLO framework and the SAM model, the background interference is reduced and the accuracy of disease area calculation is improved by calibrating the optimal division points within the anchor frame and segmenting using the SAM model.
By using image transformation and model segmentation, the mapping deviation between pixel area and actual area is reduced, thereby improving the accuracy and precision of disease area calculation.
Smart Images

Figure CN120876581A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of road defect detection technology, and in particular to a method for estimating the area of road defects based on image transformation. Background Technology
[0002] Road defects refer to structural or functional damage to road surfaces caused by a combination of factors, including natural elements, human loads, material aging, and construction defects. These defects directly affect the road's load-bearing capacity, driving safety, and service life, and require timely treatment through scientific maintenance methods.
[0003] The current method for calculating the area of road defects mainly uses image instance segmentation algorithms. By learning the feature correspondence between the original image and the mask image through the model, different types of defects are color-coded, and then the area is estimated by counting the number of pixels of each color.
[0004] However, existing methods are based on an accumulation calculation mode of 0.1*0.1 unit grid, which is prone to significant errors in images with multiple types of diseases coexisting: First, when multiple types of diseases coexist in the same grid, the grid classification of small-area diseases will be mistakenly included in the category of large-area diseases; Second, for diseases with slender shapes, a small amount of coverage at the grid edge will trigger the area statistics of the entire grid, resulting in calculation results that far exceed the actual values.
[0005] Furthermore, the fixed-angle image acquisition method used in road inspections causes a mapping deviation between the segmented pixel area and the actual physical area due to the shooting tilt angle. This results in geometric distortion between the pixel area and the actual area after SAM image segmentation, further affecting the accuracy of disease area calculation.
[0006] To address this problem, the present invention proposes a method for estimating the area of road defects based on image transformation. Summary of the Invention
[0007] The purpose of this invention is to provide a method for estimating the area of road defects based on image transformation, so as to solve the technical problems mentioned in the background art.
[0008] To achieve the above objectives, the present invention provides the following technical solution: a method for estimating the area of road defects based on image transformation, comprising the following steps: S1. Obtain the original image of the road; S2. Determine whether the original image meets the standard; if it does not meet the standard, return to S1. S3. Calculate and crop the actual area of the effective region to obtain the effective region area and the effective region image; S4. Perform perspective transformation on the effective area image to obtain an orthophoto image; S5. Use the YOLO framework to detect orthophotos and locate the defects using anchor frames. S6. Mark the optimal division points within the anchor frame and use the SAM model to segment the lesions in the image; S7. Based on the relationship between the effective area and the area of the disease's pixels, estimate the area value of the disease and store it in the database.
[0009] Preferably, the image judgment criteria in S2 include sharpness and resolution, exposure and contrast, color accuracy, and image integrity; The judgment methods include automated screening and manual review.
[0010] Preferably, the method for calculating the effective area in S3 includes the following steps: S301. Establish a vertical coordinate system based on the vertical height L1 from the camera to the road surface and the farthest shooting distance L0 in the road sampling equipment. S302. Calculate the length L of the effective area using the equipment parameters. 有效 ; S303, Based on the actual widths W and L 有效 The effective region area is calculated as W*L. 有效 .
[0011] Preferably, the perspective transformation method in S4 includes the following steps: S401. Select the coordinates of four points A, B, C, and D in the effective area, and measure the actual length of any two adjacent line segments; S402. Suppose that after being converted into an orthophoto, the pixel width of the image is a fixed value Wdown. Then, based on the relationship between Wdown and the length of the line segment, the coordinates of points A, B, C, and D in the orthophoto can be obtained. S403. Establish the transformation matrix formula between the coordinates of the forward-looking road image and the orthophoto image, and substitute the coordinate values of the four points into the matrix relationship to obtain the proportional relationship between the two.
[0012] Preferably, the forward-looking road image in S403 represents an effective area image that has not undergone perspective transformation.
[0013] Preferably, the method for calibrating the optimal division point in S6 includes the following steps: S601. Let the intersection of the diagonals of the rectangle and the two sides be halved be the first valid selected point P1. S602. Draw a line at a one-quarter scale and select the intersection point closest to P1; S603. Draw lines proportionally to one-eighth of the length and length of the rectangle, and select the intersection point closest to P1.
[0014] Preferably, the estimation of the diseased area in S7 includes the following steps: S701. Obtain the actual area corresponding to each pixel based on the ratio of the effective area to the sum of the pixels of the orthophoto image. S702. The estimated area of the disease is obtained by multiplying the total number of pixels of the disease in the image by the actual area of the pixels.
[0015] Preferably, the YOLO framework in S5 has been trained using a disease detection model, and the training method includes the following steps: Step 1: Split your dataset according to general rules to serve as training, testing, and validation sets. Step 2: Train the new classification weights of the initial YOLO model using the training set; Step 3: Evaluate the trained model using the validation set and determine whether the model meets the standards. If it does not meet the standards, return to Step 2. Step 4: Test the compliant model using the test set.
[0016] Preferably, the general rule in step one is a general rule for training machine learning models, and the proportion of the training set split is greater than that of the test set and the validation set.
[0017] Preferably, in step S6, after performing multiple instance segmentations, the SAM model filters out segmentation maps whose mask score values exceed a threshold, and the threshold value ranges from 0.5 to 0.8.
[0018] The beneficial effects of this invention are: This invention uses 27 equally divided points (P1-P27) generated by the diagonals, quarter lines, and eighth lines within the anchor frame to systematically cover the geometric center and key transition areas (such as the 1 / 8 mark of the long side) of the defect anchor frame. This ensures that the SAM (Signal Analysis and Detection) prompts point to the core morphological features of the defect (such as the center point of potholes and depressions, the boundary between the repair area and the undamaged pavement), reducing missegmentation of the SAM due to background interference. Furthermore, by converting the forward-looking road image into an orthophoto image, the mapping deviation between the image pixel area and the actual physical area is reduced, ensuring that the mapping relationship between the pixel area and the actual area remains linear and consistent, thus improving the accuracy of defect area calculation. Attached Figure Description
[0019] Figure 1 This is a flowchart of the road defect detection method of the present invention.
[0020] Figure 2 This is a detailed flowchart of the road defect detection method of the present invention.
[0021] Figure 3 This is a schematic diagram of the image acquired by the road sampling equipment of the present invention.
[0022] Figure 4 This is a schematic diagram illustrating the division of the image regions according to the present invention.
[0023] Figure 5 This is a schematic diagram illustrating the conversion of a forward-looking road image into an orthophoto image according to the present invention.
[0024] Figure 6 This is a diagram illustrating the effective area selection process in this invention.
[0025] Figure 7 This is a flowchart of the disease detection model training process in this invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] When detecting road defects, the road sampling equipment uses a fixed tilt angle to capture images. Due to the tilt angle, there is a mapping deviation between the segmented pixel area and the actual physical area. This causes geometric distortion between the pixel area and the actual area after image segmentation by SAM, further affecting the accuracy of defect area calculation.
[0028] Furthermore, existing methods, based on a cumulative calculation mode of 0.1*0.1 unit grid, are prone to significant errors in images with multiple types of diseases coexisting: First, when multiple types of diseases coexist in the same grid, the grid classification of small-area diseases may be mistakenly included in the category of large-area diseases; second, for diseases with slender shapes, even a slight overlap at the grid edge triggers the area statistics of the entire grid, resulting in calculation results far exceeding the actual values. This embodiment is invented to solve the above problems.
[0029] Please see Figures 1 to 7 As shown, an embodiment of the present invention provides a method for estimating the area of road defects based on image transformation, comprising the following steps: S1. Obtain the original image of the road.
[0030] S2. Determine whether the original image meets the standard. If it does, proceed to S3; otherwise, return to S1.
[0031] S3. Calculate the actual area of the effective region using road sampling equipment parameters and trigonometric functions, and then crop the effective region to obtain the effective region area and the effective region image.
[0032] S4. Perform perspective transformation on the effective area image to obtain an orthophoto image.
[0033] S5. Use the YOLOv8 framework to detect diseases in orthophotos and locate the diseases using anchor frames.
[0034] S6. After marking the optimal division points within the anchor frame, use the SAM model to perform multiple instance divisions of the disease.
[0035] S7. Based on the relationship between the effective area and the area of the disease's pixels, estimate the area value of the disease and store it in the database.
[0036] In this embodiment, YOLOv8 first divides an image into three feature scales of different sizes. The head module performs a category prediction branch and a regression prediction branch. The category prediction branch performs sigmoid calculation, and the regression prediction branch decodes and restores the image to xyxy format. The predicted confidence value is used for threshold filtering and comparison. The remaining detection boxes are restored to the original image size to complete the final calculation.
[0037] SAM is a model that has a significant impact on the field of image segmentation. It is a hybrid model that combines episodic memory and semantic memory. By combining the characteristics of episodic memory and semantic memory, it improves the accuracy and efficiency of image segmentation. The core of this model is that it can associate complete and clear images from incomplete or blurry images, thereby achieving accurate segmentation of objects in the image.
[0038] YOLOv8 and the SAM model are existing technologies and will not be discussed further.
[0039] It should be added that the image judgment criteria in S2 include sharpness and resolution, exposure and contrast, color accuracy, and image integrity; The judgment methods include automated screening and manual review.
[0040] In addition, the SAM model in S6 filters out the segmentation maps whose mask score values exceed the threshold after performing multiple instance segmentations. The threshold value ranges from 0.5 to 0.8. The higher the threshold, the higher the segmentation accuracy, but the slower the speed. In this embodiment, the threshold is set to 0.7.
[0041] Please see Figure 3 and Figure 4 As shown, the method for calculating the effective area in S3 includes the following steps: S301. Establish a vertical coordinate system based on the vertical height L1 of the camera from the road surface and the farthest shooting distance L0 in the road survey equipment.
[0042] S302. Divide the image into invalid region, near-distance blurred region, valid region, far-distance blurred region and full camera capture region, and let the angles between each region and the camera be α, β-α, χ-β, δ-χ and δ-α, respectively.
[0043] S303. After obtaining the values of α, β, χ, and δ based on the camera parameters, calculate the effective length L of the region corresponding to β and χ using trigonometric functions. β and L χ L χ -L β The total length L of the effective region 有效 .
[0044] S304, Based on the actual widths W and L 有效 The effective region area is calculated as W*L. 有效 .
[0045] In this embodiment, the actual width W is the width of a single lane of the highway.
[0046] Please see Figure 5 As shown, perspective transformation is a transformation that utilizes the collinearity of the perspective center, image point, and target point, and rotates the projection plane (perspective plane) around the trace line (perspective axis) by a certain angle according to the law of perspective rotation, thereby disrupting the original projection ray beam while maintaining the unchanged geometric shape of the projection on the projection plane.
[0047] In this embodiment, the perspective transformation method in S4 specifically includes the following steps: S401. Select the coordinates of four points A, B, C, and D of the effective area in the forward-looking road image, and measure the lengths of line segments AD and AB in the real world. The forward-looking road image represents the effective area image without perspective transformation.
[0048] S402. Suppose that after being converted into an orthophoto, the pixel width of the image is a fixed value Wdown. Then, based on the relationship between Wdown and the length of the line segment, the coordinates of points A, B, C, and D in the orthophoto can be obtained. Where AD and AB are the lengths of line segments AD and AB in the real world.
[0049] S403. Establish the matrix formula for the coordinates of the forward-looking road image and the orthophoto image, and substitute the coordinate values of A, B, C, and D in S402 into the matrix relationship to obtain the correspondence between the two. The matrix formula is: Where (u, v) are the coordinates of a point in the forward-looking road image, and (x, y) are the corresponding coordinates in the orthophoto image.
[0050] Since the transformation matrix contains 8 unknowns, the ratio of the coordinates of the forward-looking road image to the coordinates of the orthophoto image can be obtained by substituting the coordinates of the four points A, B, C, and D into the matrix and solving for the unknowns.
[0051] Furthermore, when performing batch top-down transformations on multiple images, since the camera's tilt angle does not change during the same image acquisition process, the perspective transformation matrix obtained from one image can be applied to all images acquired in that acquisition.
[0052] In this embodiment, please refer to Figure 6 As shown, the method for calibrating the optimal division point in S6 includes the following steps: S601. Let the first valid selected point P1 be the intersection of the diagonals of the rectangle and the lines that halve the horizontal and vertical sides.
[0053] S602. Based on the above, draw lines at a quarter scale and select the intersection point closest to P1 to obtain 8 selected points from P2 to P9.
[0054] S603. Draw lines proportional to one-eighth of the long side of the rectangle, and select the intersection point closest to P1 to obtain 10 selected points from P10 to P20.
[0055] S604. Draw lines proportional to one-eighth of the shorter side of the rectangle, and select the intersection point closest to P1 to obtain 6 selected points P21-P27.
[0056] The above process iterates by drawing equal dividing lines in a loop according to the proportion of the disease in the overall actual acquired image. The result of selecting the effective area of the whole process provides the optimal selection point for SAM, and then the effective area of the disease is extracted by the SAM instance segmentation algorithm.
[0057] In this embodiment, the estimation of the diseased area in S7 includes the following steps: S701. The actual area corresponding to each pixel is obtained based on the ratio of the effective area to the sum of the pixels in the orthophoto image, i.e.: Let the sum of the pixels of the image to be normalized be SUM. 像素 The actual total area is SUM 面积 Then the actual area of each pixel is P. 面积 The correspondence is as follows: .
[0058] S702. The estimated area of the disease is obtained by multiplying the total number of pixels of the disease in the image by the actual area of each pixel, that is: Among them, S 面积 NUM represents the actual area affected by the disease. 像素 This represents the total number of pixels in the image where the disease is present.
[0059] Please note that you should refer to [link / reference]. Figure 7 As shown, the YOLOv8 framework in S5 has been trained with a disease detection model, and the training method includes the following steps: Step 1: According to the general rules for training machine learning models, split your own dataset into training, test and validation sets in a ratio of 6:2:2, with the training set accounting for 60%, and the test and validation sets each accounting for 20%.
[0060] Step 2: Train the new classification weights of the initial YOLOv8 model using the training set.
[0061] Step 3: Evaluate the trained model using the validation set, determine whether the model meets the standards, and adjust or deploy the model based on the results.
[0062] Step 4: Test the compliant model using the test set.
[0063] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for estimating the area of road defects based on image transformation, characterized in that, Includes the following steps: S1. Obtain the original road image; S2. Determine whether the original image meets the standard. If it does not meet the standard, return to S1. S3. Calculate and crop the actual area of the effective region to obtain the effective region area and the effective region image; S4. Perform perspective transformation on the effective area image to obtain an orthophoto image; S5. Use the YOLO framework to detect orthophotos and locate the defects using anchor frames. S6. Mark the optimal division points within the anchor frame and use the SAM model to segment the lesions in the image; S7. Based on the relationship between the effective area and the area of the disease's pixels, estimate the area value of the disease and store it in the database.
2. The method for estimating the area of road defects based on image transformation according to claim 1, characterized in that, The image judgment criteria in S2 include sharpness and resolution, exposure and contrast, color accuracy, and image integrity. The judgment methods include automated screening and manual review.
3. The method for estimating the area of road defects based on image transformation according to claim 2, characterized in that, The method for calculating the effective area in S3 includes the following steps: S301. Establish a vertical coordinate system based on the vertical height L1 from the camera to the road surface and the farthest shooting distance L0 in the road sampling equipment. S302. Calculate the length L of the effective area using the equipment parameters. 有效 ; S303, Based on the actual widths W and L 有效 The effective region area is calculated as W*L. 有效 .
4. The method for estimating the area of road defects based on image transformation according to claim 3, characterized in that, The perspective transformation method in S4 includes the following steps: S401. Select the coordinates of four points A, B, C, and D in the effective area, and measure the actual length of any two adjacent line segments; S402. Suppose that after being converted into an orthophoto, the pixel width of the image is a fixed value Wdown. Then, based on the relationship between Wdown and the length of the line segment, the coordinates of points A, B, C, and D in the orthophoto can be obtained. S403. Establish the transformation matrix formula between the coordinates of the forward-looking road image and the orthophoto image, and substitute the coordinate values of the four points into the matrix relationship to obtain the proportional relationship between the two.
5. The method for estimating the area of road defects based on image transformation according to claim 4, characterized in that, The forward-looking road image in S403 represents the effective area image without perspective transformation.
6. The method for estimating the area of road defects based on image transformation according to claim 5, characterized in that, The method for determining the optimal division point in S6 includes the following steps: S601. Let the intersection of the diagonals of the rectangle and the two sides be halved be the first valid selected point P1. S602. Draw a line at a one-quarter scale and select the intersection point closest to P1; S603. Draw lines proportionally to one-eighth of the length and length of the rectangle, and select the intersection point closest to P1.
7. The method for estimating the area of road defects based on image transformation according to claim 6, characterized in that, The estimation of diseased area in S7 includes the following steps: S701. Obtain the actual area corresponding to each pixel based on the ratio of the effective area to the sum of the pixels of the orthophoto image. S702. The estimated area of the disease is obtained by multiplying the total number of pixels of the disease in the image by the actual area of the pixels.
8. The method for estimating the area of road defects based on image transformation according to claim 7, characterized in that, The YOLO framework in S5 has been trained using a disease detection model, and the training method includes the following steps: Step 1: Split your dataset according to general rules to serve as training, testing, and validation sets. Step 2: Train the new classification weights of the initial YOLO model using the training set; Step 3: Evaluate the trained model using the validation set and determine whether the model meets the standards. If it does not meet the standards, return to Step 2. Step 4: Test the compliant model using the test set.
9. The method for estimating the area of road defects based on image transformation according to claim 8, characterized in that, The general rule in step one is a general rule for training machine learning models, and the proportion of the training set split is greater than that of the test set and the validation set.
10. The method for estimating the area of road defects based on image transformation according to claim 9, characterized in that, In S6, after performing multiple instance segmentations, the SAM model filters out segmentation maps whose mask score values exceed a threshold, and the threshold value ranges from 0.5 to 0.8.