Foresight borehole video image crack intelligent identification and quantitative characterization method and system
By improving the YOLOv8 and Mask R-CNN models combined with SURF and optical flow method, accurate crack detection and segmentation in complex backgrounds is achieved, the accuracy and efficiency problems of crack identification and quantitative characterization in geotechnical engineering are solved, and high-precision crack geometric information support is provided.
Patent Information
- Application Number
- CN202510424519.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing crack identification and quantitative characterization methods have problems with insufficient accuracy and efficiency in the field of geotechnical engineering, especially in the complex geological background, and traditional methods are prone to false detection and missed detection.
The improved YOLOv8 model is used for crack detection, combined with the optimized Mask R-CNN model for segmentation, and a boundary-weighted binary cross-entropy loss function is introduced. The cross-frame matching and tracking of cracks is achieved through feature matching SURF and optical flow method Lucas-Kanade, and the crack length and pose changes are calculated.
It significantly improves the accuracy and robustness of crack identification, enhances the dynamic tracking capability of cracks, provides high-precision quantitative characterization of cracks, and can better support rock mass stability assessment and geological disaster prevention and control.
Smart Images

Figure CN120279464A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of geotechnical engineering analysis, resource extraction design, disaster warning, etc., and particularly relates to a method and system for intelligent identification and quantitative characterization of fractures in forward-looking borehole video images. Background Art
[0002] In the fields of mining engineering, geotechnical engineering, and geological disaster prevention and control, the forward-looking borehole camera technology can accurately record the characteristics and structures of fractures by obtaining high-resolution images inside rock mass boreholes, and has become an important tool for fracture identification and analysis in underground projects such as mines and tunnels. By analyzing the forward-looking borehole images, the fracture characteristics of underground rock masses can be deeply understood, providing accurate basis for rock mass stability assessment and support design, and then effectively guiding the safety and optimization design of underground projects.
[0003] However, traditional fracture identification methods mainly rely on manual observation, with low efficiency and being easily affected by subjective factors. With the progress of technology, automated methods such as edge detection, texture analysis, and morphological processing have gradually improved the identification efficiency, but still prone to false detection and missed detection in complex backgrounds. In recent years, convolutional neural network (CNN) models have made certain progress by automatically extracting fracture image features, but there are still limitations in processing high-resolution and complex background images. For this reason, advanced CNN models such as ResNet and Inception have significantly improved the accuracy and robustness of fracture identification through residual connections and multi-scale feature extraction mechanisms, but false detection and missed detection may still occur in complex backgrounds. At the same time, object detection models also play an important role in fracture identification. YOLO series models (such as YOLOv5, YOLOv6, YOLOv7) have high real-time detection capabilities, but there are still deficiencies in fine fracture segmentation. Although Faster R-CNN performs excellently in terms of accuracy, its high computational complexity limits its wide use in real-time applications. Single detection models have limitations in dealing with complex fracture geometries, especially difficult to meet the requirements in detail processing.
[0004] Therefore, in the field of geotechnical engineering, the accuracy and efficiency of existing fracture identification and quantitative characterization methods need to be improved. Summary of the Invention
[0005] In order to improve the accuracy and efficiency of fracture identification and quantitative characterization in forward-looking borehole images, especially in complex geological backgrounds, and provide more accurate and comprehensive fracture information, the present invention provides a method and system for intelligent identification and quantitative characterization of fractures in forward-looking borehole video images, which can provide more accurate fracture geometric information and provide reliable data support for rock mass stability assessment, support design, and geological disaster prevention and control.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] An intelligent recognition and quantitative characterization method for cracks in forward-looking borehole video images, the method comprising:
[0008] Step 1: Use an improved YOLOv8 model to detect cracks in the forward-looking borehole video and generate crack region bounding boxes;
[0009] Step 2: Use an optimized Mask R-CNN model to segment the region within the crack region bounding box to obtain a crack mask;
[0010] Step 3: Extract the crack morphological features in the crack mask;
[0011] Step 4: Based on the crack region and crack morphological feature parameters, use feature matching SURF and optical flow method Lucas-Kanade to achieve cross-frame matching and tracking of cracks, and calculate the crack length and attitude change.
[0012] Preferably, the method of using an improved YOLOv8 model to detect cracks in the forward-looking borehole video includes:
[0013] Add a Capsule layer after the C2f layer in the Neck part of the YOLOv8 model, transmit and aggregate feature information through a dynamic routing mechanism, and capture the spatial relationship of cracks.
[0014] Preferably, the method of using an optimized Mask R-CNN model to segment the region within the crack region bounding box to obtain a crack mask includes:
[0015] Introduce a boundary-weighted binary cross-entropy loss function, and assign weights to the crack boundary region by defining a weight matrix W; the expression is:
[0016]
[0017] In the formula, L WBCE is the weighted segmentation loss function; y i is the true label, 0 represents the background, and 1 represents the crack; p i is the probability that the predicted pixel i belongs to the crack, and W i is the weight of pixel i in the weight matrix;
[0018] The weight matrix is calculated based on the crack gradient information and the edge region, guiding the Mask R-CNN model to focus on the crack boundary region.
[0019] Preferably, the method of calculating the weight matrix based on the crack gradient information and the edge region to guide the Mask R-CNN model to focus on the crack boundary region includes:
[0020] Calculate the gradients G(x, y) in the horizontal and vertical directions of the image, i.e., the drilling video frame, using the Sobel operator and generate an edge intensity map. Analyze the edge intensity at each position in the image. The gradient calculation is as follows:
[0021]
[0022] In the formula, G x (x, y) and G y (x, y) are the gradients of the image in the horizontal and vertical directions respectively;
[0023] Adopt the Canny edge detection method to extract the edge information C(x, y) in the image, and combine the distance transformation to calculate the distance from each pixel to the nearest boundary; combine the gradient magnitude with the Canny edge detection result, and at the same time introduce the weight based on the distance transformation. Map the distance value to the [0, 1] interval through normalization operation to assign weights to the crack edge region. Among them, the distance change is as follows:
[0024]
[0025] In the formula, dist(x, y) is the distance from the current pixel to the nearest boundary; max(x, y) is the maximum distance from all pixels in the image to the nearest boundary;
[0026] Define the weight matrix W(x, y) by the product of the gradient magnitude and the edge detection result, and combine the result of the distance transformation to guide the Mask R-CNN model to focus on the boundary region of the crack. The calculation formula is as follows:
[0027]
[0028] In the formula, α and β are adjustment parameters to control the relative importance of gradient information, edge detection, and distance transformation; min(W) and max(W) are the minimum and maximum values of all pixels in the weight matrix respectively.
[0029] Preferably, the method for extracting the crack morphological features in the crack mask includes:
[0030] Statistically calculate the area of the crack by counting the number of non-zero pixels in the binary image, and calculate the perimeter using the foreground pixel boundary point set; adopt the Zhang-Suen thinning algorithm to extract the skeleton of the crack, and identify the bifurcation points in the skeleton through neighbor analysis; calculate the crack width through distance transformation, and process the width of the bifurcation points to obtain the average width of the crack; at the same time, use the image moment method to extract the centroid position of the crack, and fit the crack ellipse to determine the crack direction.
[0031] Preferably, the method for statistically calculating the area of the crack by counting the number of non-zero pixels in the binary image includes:
[0032] Assume that the input binary image is I(x, y), and the size of the image is H×D, where H is the height and D is the width. Traverse each pixel point of the crack mask image, count the number of non-zero pixels, and convert it into the actual area unit. The calculation formula is as follows:
[0033]
[0034] In the formula, countMonZero(I) is the area of crack pixels; 1 I(i,j) ≠0 is the indicator function, which returns 1 if I(i, j) is not 0, otherwise it returns 0; I(i, j) is the pixel value of image I at (i, j);
[0035] The method for calculating the perimeter using the set of foreground pixel boundary points includes:
[0036] Assume that the input binary image is I(x, y), where I(x, y)=255 represents the foreground pixel and I(x, y)=0 represents the background pixel. The boundary of the foreground region forms a contour, and the contour is composed of a set of points {P0, P1, …, P n-1}, and the perimeter L is expressed as the total length of the point set. The calculation formula is as follows:
[0037]
[0038] In the formula, L is the total length of the crack contour point set; dist(P i , P i+1 ) is the Euclidean distance between two adjacent contour points;
[0039] The method for extracting the skeleton of the crack using the Zhang-Suen thinning algorithm includes:
[0040] Gradually thin the foreground region in the binary image I(x, y) to generate a skeleton with a width of one pixel, while maintaining the topological structure of the crack. For each pixel point P i , when it meets the preset conditions, it is deleted, and finally a skeleton structure is generated. The expression of the preset conditions is as follows:
[0041]
[0042] In the formula, the pixel value of 255 represents the crack and 0 represents the background; A(P i ) is the number of transitions from background to foreground in the 8-neighborhood of pixel point P i ; B(P1) is the number of foreground pixels in the 8-neighborhood of pixel point P i ; D i (1) is the removal condition check function D i (1)= P2·P4·P6 = 0 and P4·P6·P8 = 0, D i (1) = P2·P4·P8 = 0 and P2·P6·P8 = 0;
[0043] The method for identifying the bifurcation points in the skeleton through neighbor analysis includes:
[0044] For each pixel point P in the skeleton Li perform neighbor analysis and calculate the number N(P Li ) of neighbor pixels. If N(P Li ) > 2, then mark this pixel point as a bifurcation point. Assume that the neighborhood pixel distribution of a certain point P Li is as follows. Calculate N(P Li ) = 4 > 2, so P Li is a bifurcation point;
[0045]
[0046] In the formula, P Li is the skeleton pixel point; 1 is the neighbor pixel;
[0047] The method for calculating the crack width through distance transformation includes:
[0048] Perform distance transformation on the binary mask image and calculate the Euclidean distance dist(P i ) from each foreground pixel point to the nearest background pixel point. The calculation formula is:
[0049]
[0050] In the formula, β is the set of all background pixel points; P i (x i ,y i ) is the coordinate of the foreground pixel point; P j (x j ,y j ) is the coordinate of the background pixel point;
[0051] For the contour point P i , the crack width w i is obtained by multiplying the distance value by 2. The calculation formula is:
[0052] w i = 2·dist(P i )
[0053] For the bifurcation point P Li , calculate the width in each direction respectively. Assume there are K directions, and the width is:
[0054]
[0055] where dist j (P Li ) is the nearest background distance of the bifurcation point in the j-th direction;
[0056] The method for extracting the centroid position of the crack using the image moment method includes:
[0057] Calculate the coordinates of the crack centroid using the image moment, and obtain the center position of the crack through geometric analysis technology. The calculation formula is:
[0058] Zero-order moment: M 00 = ∑ x ∑ y I(x, y), first-order moment:
[0059] Centroid coordinates:
[0060] where M ij is the (i + j)-th order moment of the image; i and j are the orders of the moment; C x and C y are the coordinates of the crack centroid;
[0061] The method for fitting a crack ellipse to determine the crack direction includes:
[0062] Based on the contour point set {(x i , y i )}, fit an ellipse in the sense of least squares, and adjust the parameters of the ellipse to minimize the sum of the squares of the distances from the contour points to the ellipse boundary to determine the crack direction. The calculation formula is:
[0063]
[0064] where (x i , y i ) are the coordinates of the image contour points; (cx, cy) are the coordinates of the ellipse center; (a, b) are the major and minor axes of the ellipse; θ is the rotation angle of the ellipse, that is, the required crack direction angle.
[0065] Preferably, the method for implementing cross-frame matching and tracking of cracks using feature matching SURF and optical flow method Lucas-Kanade includes:
[0066] In the crack mask obtained by two-stage network detection, use the SURF algorithm to extract the key feature points in the crack area and calculate the feature descriptor v for inter-frame matching;
[0067] Use the KNN algorithm based on the Euclidean distance d(v i , v j)Find the nearest neighbor feature points for each feature point, thus initially establishing the matching relationship of fracture feature points between multiple frames;
[0068] Use the Lucas-Kanade optical flow method to further track the position changes of feature points in the image sequence, especially in scenarios where the position changes of fractures are small in consecutive frames. For the already matched feature points, calculate their motion vectors through the optical flow method, obtain the motion trajectories of fractures in consecutive frames, and track the changes of fractures;
[0069] Combine the motion vectors tracked by optical flow with the results of feature matching, and frame by frame track the position and morphological changes of fracture feature points. For the feature points matched in adjacent frames, calculate the length of the fracture by accumulating the distances between feature points along the fracture trajectory.
[0070] The present invention provides a forward-looking borehole video image fracture intelligent recognition and quantitative characterization system, which is used to implement the foregoing method, including a detection module, a segmentation module, an extraction module, and a calculation module;
[0071] The detection module is used to detect fractures in the forward-looking borehole video by using an improved YOLOv8 model, and generate fracture region bounding boxes;
[0072] The segmentation module is used to segment the region within the fracture region bounding box by using an optimized Mask R-CNN model to obtain a fracture mask;
[0073] The extraction module is used to extract the fracture morphological features in the fracture mask;
[0074] The calculation module is used to realize cross-frame matching and tracking of fractures based on the fracture region and fracture morphological feature parameters by using feature matching SURF and the Lucas-Kanade optical flow method, and calculate the fracture length and attitude changes.
[0075] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0076] 1. Improve the accuracy of fracture recognition and segmentation. Traditional fracture recognition methods (such as edge detection and threshold segmentation) usually have problems of false detection, missed detection, and inefficiency, and it is difficult to accurately identify fractures especially in complex backgrounds. However, the present invention combines the YOLOv8 and Mask R-CNN models, and uses deep learning technology to achieve fast detection and fine segmentation, significantly improving the accuracy and robustness of fracture recognition. YOLOv8 can quickly detect the position of fractures in a short time, and Mask R-CNN improves the accuracy of fracture boundaries through refined segmentation, especially showing excellent performance in complex backgrounds and the recognition of small fractures. This enables fracture recognition to be more accurately applied to the risk assessment and monitoring of underground projects such as mines and tunnels.
[0077] 2. Enhance the dynamic tracking ability of fractures. This invention adopts the cross-frame matching and dynamic tracking technology that combines the SURF algorithm and the optical flow method (Lucas-Kanade), and can track the development process of fractures in real time and accurately in the video. This technology can solve the problem that traditional methods cannot effectively track the evolution of fractures dynamically. By tracking the position changes of fractures in multiple video frames, the invention can accurately calculate the length, attitude and evolution process of fractures, provide dynamic data on the development trend of fractures, and thus better support the safety monitoring and optimal design of underground engineering.
[0078] 3. Provide high-precision quantitative characterization of fractures. This invention extracts morphological features such as the area, width, direction, and centroid of fractures through image processing and mathematical calculation techniques, and conducts precise quantitative analysis. Especially in terms of details such as fracture width, perimeter, and skeleton extraction, this invention provides more accurate and refined quantitative methods. These data can help engineers comprehensively understand the geometric morphology of fractures and their variation laws, and provide a scientific basis for rock mass stability analysis, support design, and geological disaster prevention and control. Compared with traditional qualitative or rough quantification methods, this invention greatly improves the accuracy and usability of fracture characterization. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the technical solutions of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0080] Figure 1 Schematic diagram of the rapid identification and fine segmentation results of fractures in the forward-looking borehole video in the embodiment of the present invention;
[0081] Figure 2 Schematic diagram of the fracture feature extraction and matching results in the forward-looking borehole video in the embodiment of the present invention;
[0082] Figure 3 Schematic diagram of the fracture parameter characterization results in the forward-looking borehole video in the embodiment of the present invention: (a) Fracture length characterization; (b) Fracture width characterization; (c) Fracture angle characterization;
[0083] Figure 4 Schematic diagram of the video correlation dynamic matching method in the embodiment of the present invention;
[0084] Figure 5 Visualization result diagram of the fracture parameter characterization in the forward-looking borehole video in the embodiment of the present invention;
[0085] Figure 6 Schematic diagram of the architecture of the dual-strategy identification and characterization method in the embodiment of the present invention;
[0086] Figure 7 This is the optimized diagram of the YOLOv8 network structure in the embodiment of the present invention. Detailed implementation manners
[0087] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0088] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0089] Embodiment 1
[0090] The present invention provides a method for intelligent recognition and quantitative characterization of cracks in forward-looking drilling video images, and the specific steps are as follows:
[0091] Step 1: Use the improved YOLOv8 model to quickly detect cracks in the forward-looking drilling video and generate crack region bounding boxes;
[0092] Step 2: Based on the crack bounding boxes in Step 1, use the Mask R-CNN model to finely segment the region within the crack bounding boxes. Introduce a boundary-weighted binary cross-entropy loss function and design a weight matrix W to give higher weights to the crack boundary regions to improve the model segmentation accuracy, thereby obtaining crack masks;
[0093] Step 3: Extract the crack morphological features in the crack masks in Step 2, count the number of non-zero pixels in the binary image to calculate the area of the cracks, and calculate the perimeter using the set of foreground pixel boundary points; Use the Zhang-Suen thinning algorithm to extract the skeletons of the cracks, and identify the bifurcation points in the skeletons through neighbor analysis; Calculate the crack width through distance transformation and process the widths of the bifurcation points to obtain the average width of the cracks; At the same time, use the image moment method to extract the centroid position of the cracks, and fit the crack ellipse to determine the crack direction;
[0094] Step 4: Based on the crack regions and crack morphological feature parameters in Step 1 and Step 3, use feature matching (SURF) and optical flow method (Lucas-Kanade) to achieve cross-frame matching and tracking of the cracks, and calculate the crack length and its attitude change.
[0095] Furthermore, the dual-strategy architecture is as Figure 6 shown.
[0096] Furthermore, the YOLOv8 model described in step 1 is an improved YOLOv8 model for rapid crack detection. Specifically:
[0097] A Capsule layer is newly added after the C2f layer in the Neck part of the YOLOv8 model. Through the dynamic routing mechanism, it accurately transmits and aggregates feature information, enhances the multi-scale feature fusion ability, effectively improves the detection performance of small cracks and complex crack morphologies, and captures the spatial relationship of cracks, thereby improving the generalization ability of the model. As Figure 7 shown.
[0098] Furthermore, the Mask R-CNN model described in step 2 is a Mask R-CNN model that optimizes the segmentation loss for fine segmentation of the crack area. Specifically:
[0099] The boundary-weighted binary cross-entropy loss function is introduced. By defining the weight matrix W, a higher weight is assigned to the crack boundary area, improving the segmentation accuracy of the model for the boundary area.
[0100]
[0101] In Equation 3, L WBCE is the weighted segmentation loss function; y i is the ground truth label, 0 represents the background, and 1 represents the crack; p i is the probability that the predicted pixel i belongs to the crack, W i is the weight of pixel i in the weight matrix. By calculating the distance of each pixel to the crack boundary, a higher weight is assigned to the pixel closer to the boundary: w i = 1 + exp(-d i 2 / δ 2 ), d i is the distance from pixel i to the nearest boundary, and δ is a parameter to adjust the weight decay rate.
[0102] The weight matrix is calculated based on the crack gradient information and the edge region, guiding the segmentation model to pay more attention to the crack boundary area. First, the horizontal and vertical gradients G(x, y) of the image, i.e., the drilling video frame, are calculated using the Sobel operator and an edge intensity map is generated. The edge intensity at each position in the image is analyzed, and the gradient calculation is as follows:
[0103]
[0104] In Equation 4, G x (x, y) and G y (x, y) are the gradients of the image in the horizontal and vertical directions respectively.
[0105] Subsequently, the Canny edge detection method is used to extract the edge information C(x, y) in the image, and the distance transform is combined to calculate the distance from each pixel to the nearest boundary. To further enhance the weight of the edge region, the gradient magnitude is combined with the Canny edge detection result, and at the same time, a weight based on the distance transform is introduced. Through normalization operation, the distance value is mapped to the [0, 1] interval, and a higher weight is assigned to the crack edge region. The distance change is as follows:
[0106]
[0107] In Equation 5, dist(x, y) is the distance from the current pixel to the nearest boundary; max(x, y) is the maximum distance from all pixels in the image to the nearest boundary.
[0108] Finally, the weight matrix W(x, y) is defined by the product of the gradient magnitude and the edge detection result, and combined with the result of the distance transform to optimize the segmentation accuracy of the model in the crack boundary region. This weighting strategy enables the model to pay more attention to the edge details of the cracks during the training process, thereby improving the overall segmentation performance.
[0109]
[0110] In Equation 6, α and β are adjustment parameters that control the relative importance of gradient information, edge detection, and distance transform; min(W) and max(W) are the minimum and maximum values of all pixels in the weight matrix, respectively.
[0111] Furthermore, the crack feature extraction method described in step three is specifically as follows:
[0112] (1) Crack area calculation: Assume that the input binary image is I(x, y), and the size of the image is H×D (height is H, width is D). Traverse each pixel point of the crack mask image and count the number of non-zero pixels, which is then converted to the actual area unit (such as square centimeters).
[0113]
[0114] In Equation 7, countMonZero(I) is the crack pixel area; 1 I(i,j) ≠0 is the indicator function, which returns 1 if I(i, j) is not 0, otherwise returns 0; I(i, j) is the pixel value of image I at (i, j).
[0115] (2) Crack perimeter calculation: Assume that the input binary image is I(x, y), where I(x, y) = 255 represents the foreground pixel and I(x, y) = 0 represents the background pixel. The boundary of the foreground region forms a contour, and the contour consists of a set of points {P0, P1, …, P n-1}, and the perimeter L is expressed as the total length of the point set.
[0116]
[0117] In Equation 8, L is the total length of the crack contour point set; dist(P i , P i+1 ) is the Euclidean distance between two adjacent contour points.
[0118] (3) Crack skeleton extraction: Gradually thin the foreground area (pixel value is 255) in the binary image I(x, y) to generate a one-pixel-wide skeleton while maintaining the topological structure of the crack. For each pixel point P i , it is deleted when the following conditions are met, and finally a skeleton structure is generated, as shown in Figure 4 .
[0119]
[0120] In Equation 9, the pixel value of 255 represents the crack, and 0 represents the background; A(P i ) is the number of transitions from background to foreground in the 8-neighborhood of the pixel point P i ; B(P1) is the number of foreground pixels in the 8-neighborhood of the pixel point P i ; D i (1) is the removal condition check function D i (1) = P2·P4·P6 = 0 and P4·P6·P8 = 0, D i (1) = P2·P4·P8 = 0 and P2·P6·P8 = 0.
[0121] (4) Node analysis: Perform neighbor analysis on each pixel point P Li in the skeleton, calculate the number of neighbor pixels N(P Li ), if N(P Li ) > 2, then mark this pixel point as a bifurcation point. Assume that the neighborhood pixel distribution of a certain point P Li is as follows, and it is calculated that N(P Li ) = 4 > 2, so P Li is a bifurcation point.
[0122]
[0123] In Equation 10, P Li is the skeleton pixel point; 1 is the neighbor pixel.
[0124] (5) Crack width calculation: Perform distance transformation on the binary mask image, and calculate the Euclidean distance dist(P i)。
[0125]
[0126] In Equation 11, β is the set of all background pixel points; P i (x i ,y i ) is the coordinate of the foreground pixel point; P j (x j ,y j ) is the coordinate of the background pixel point.
[0127] For the contour point P i , the crack width w i can be obtained by multiplying the distance value by 2.
[0128] w i =2·dist(P i )(12)
[0129] For the bifurcation point P Li , calculate the widths in each direction respectively. Suppose there are K directions, and its widths are:
[0130]
[0131] In Equation 13, dist j (P Li ) is the nearest background distance of the bifurcation point in the j-th direction.
[0132] (6) Crack centroid extraction: Calculate the coordinates of the crack centroid using image moments, and obtain the center position of the crack through geometric analysis techniques.
[0133] Zero-order moment (area): M 00 =∑ x ∑ y I(x,y); First-order moment:
[0134] Centroid coordinates:
[0135] In Equation 15, M ij is the (i + j)-th order moment of the image; i and j are the orders of the moments, C x and C y are the crack centroid coordinates.
[0136] (7) Crack direction calculation: Based on the contour point set {(x i ,y i )}, fit an ellipse in the least squares sense, and adjust the parameters of the ellipse to minimize the sum of the squares of the distances from the contour points to the ellipse boundary to determine the crack direction.
[0137]
[0138] In Equation 16, (x i , y i ) are the coordinate points of the image contour; (cx, cy) are the coordinates of the ellipse center; (a, b) are the major and minor axes of the ellipse; θ is the rotation angle of the ellipse, that is, the required crack direction angle.
[0139] Furthermore, in the crack feature extraction method described in Step 3, to avoid interference from feature points at the crack edge and the peripheral area, the centroid of the first-frame crack mask image is used as the center point for crack tracking, and the range initially identified by YOLOv8 is used as the feature point detection area, so as to extract the initial feature points of the first-frame crack.
[0140] Furthermore, the crack dynamic tracking and cross-frame matching method specifically includes the following implementation steps:
[0141] Step 1: In the crack mask obtained by two-stage network detection, the SURF algorithm is used to extract the key feature points in the crack area, and the feature descriptor v is calculated for inter-frame matching.
[0142] v = [∑dx, ∑|dx|, ∑dy, ∑|dy|] (17)
[0143] In Equation 17: dx is the Haar wavelet response in the horizontal direction; dy is the Haar wavelet response in the vertical direction; ∑dx is the sum of the Haar wavelet responses in the horizontal direction within the neighborhood of the feature point; ∑|dx| is the sum of the absolute values of the Haar wavelet responses in the horizontal direction within the neighborhood of the feature point; ∑dy is the sum of the Haar wavelet responses in the vertical direction within the neighborhood of the feature point; ∑|dy| is the sum of the absolute values of the Haar wavelet responses in the vertical direction within the neighborhood of the feature point.
[0144] Step 2: The KNN algorithm is used to find the nearest neighbor feature point for each feature point based on the Euclidean distance d(v i , v j ) to initially establish the matching relationship of crack feature points between multiple frames and ensure that these feature points belong to the same crack area. The Euclidean distance matching calculation is as follows:
[0145]
[0146] In Equation 18: v i , v j are feature vectors; v i k , v j k are the k-th components of the feature vector; n is the dimension of the feature vector.
[0147] Step 3: Use the Lucas-Kanade optical flow method to further track the position changes of feature points in the image sequence, especially in scenarios where the position changes of the fissure are small in consecutive frames. For the already matched feature points, calculate their motion vectors using the optical flow method to obtain the motion trajectories of the fissure in consecutive frames, and further accurately track the changes of the fissure.
[0148] I x u+I y v+I t =0 (19)
[0149] Solve the optical flow equation using the least squares method:
[0150]
[0151] In Equations 19 and 20: I x is the gradient of the image in the x direction; I y is the gradient of the image in the y direction; I t is the gradient of the image at time t (time derivative); u is the optical flow (motion vector component) of the pixel point in the x direction; v is the optical flow (motion vector component) of the pixel point in the y direction.
[0152] Step 4: Combine the motion vectors obtained by optical flow tracking with the results of feature matching to track the position and morphological changes of the fissure feature points frame by frame. For the matched feature points in adjacent frames, calculate the length of the fissure by accumulating the distances between feature points along the fissure trajectory. To avoid problems of repeated matching and tracking interruption when summarizing fissures in multiple frames, use the RANSAC algorithm to eliminate mis-matched feature points to ensure the accuracy of the fissure tracking results.
[0153] Det(H)=L xx L yy -(L xy ) 2 (21)
[0154] In Equation 21: Det(H) is the determinant of the Hessian matrix; L xx is the second derivative in the x direction of the image at point (x, y) and scale σ; L yy is the second derivative in the y direction of the image at point (x, y) and scale σ; L xy is the mixed derivative (xy direction) of the image at point (x, y) and scale σ.
[0155] Furthermore, in the above quantitative characterization, the maximum trajectory length is used as a reasonable substitute for the fissure length.
[0156] The present invention improves the feature fusion method of the YOLOv8 model, optimizes the segmentation loss function of Mask R-CNN, and coordinates the real-time processing performance and computational overhead of the two-stage model in a cascaded form, significantly improving the crack identification speed while ensuring high detection accuracy. The present invention mainly includes the generation of crack bounding boxes, the extraction of crack masks, the analysis of crack morphological features (such as area, width, centroid, etc.), as well as the dynamic tracking and cross-frame matching of cracks, and can effectively extract information such as the trace length, aperture, and occurrence of cracks in borehole video images. The present invention dynamically tracks cracks through feature matching and optical flow method, directly quantifies crack features from borehole video images, and avoids the influence of adverse factors such as borehole video acquisition, stitching, and environment on identification and characterization. The present invention can provide more accurate crack geometric information, providing reliable data support for rock mass stability assessment, support design, and geological disaster prevention and control.
[0157] Example Two
[0158] The forward-looking borehole images in the examples are from the roof of a mining roadway in a coal mine in the Jiaozuo mining area. The borehole depth is 15m, the aperture is 50mm, the resolution is 800×600mm, and the imaging speed is 0.5m / min.
[0159] The specific implementation steps are as Figures 1 to 5 shown. A method for intelligent identification and quantitative characterization of cracks in forward-looking borehole video images according to the present invention is as follows in detail:
[0160] Step 1: Use the YOLOv8 model to quickly detect cracks in the forward-looking borehole video and generate crack bounding boxes, as shown in (a) and (b) in Figure 1 ;
[0161] Step 2: Crop the crack bounding box area as the input of the Mask R-CNN model, finely segment the crack area within the bounding box, and obtain the crack mask, as shown in (c), (d), and (e) in Figure 1 ;
[0162] Step 3: Use image processing techniques and mathematical calculations to extract the morphological features in the crack mask, including parameters such as area, width, direction, centroid, etc. in each frame of video, as shown in (f), (g), and (h) in the figure;
[0163] Step 4: Use feature matching (SURF) and optical flow method (Lucas-Kanade) to achieve cross-frame matching and tracking of cracks (the process is as shown in Figure 4 ), dynamically track the development process of cracks, and calculate the changes in crack length and its attitude, as shown in Figure 2 ;
[0164] The results of crack quantitative characterization are as shown in Figure 3As shown, a total of 4 cracks were detected in this video segment. According to Figure 3 (a) and the visualization results (such as Figure 5 shown), the characteristic point trajectory length of crack 1 is the largest, reaching 672.03 mm, and the data is concentrated between 300 mm and 500 mm, indicating that crack 1 is large in scale and has strong ductility; a small number of data points are less than 200 mm, indicating that the local variability of crack 1 is small, but the overall trend is a relatively long crack. The maximum length of crack 2 is 437.40 mm, slightly shorter than crack 1, and its data points are distributed in multiple length intervals, showing a relatively uniform expansion of local cracks. The maximum length of crack 3 is 557.46 mm, close to crack 2, and the data is sparsely distributed, concentrated between 390 mm and 400 mm, indicating that the characteristic point trajectory length of crack 3 is relatively average and the change is small, and it may be a relatively regular crack. The maximum length of crack 4 is 210.91 mm, significantly shorter than the other cracks.
[0165] Figure 3 (b) and the visualization results (such as Figure 5 shown) show that the average width of crack 1 is 13.77 mm, and the width distribution range is wide, ranging from 1.20 mm to 19.60 mm, indicating that the width change of crack 1 is large, and the data at the wider end (>10 mm) is relatively concentrated, showing a higher degree of cracking. The average width of crack 2 is 8.86 mm, significantly less than that of crack 1, and the width variation is small, mainly concentrated between 7 mm and 18 mm. The average width of crack 3 is 13.17 mm, close to that of crack 1, but its width distribution is relatively uniform, concentrated between 11.20 mm and 19.70 mm, and the extreme width fluctuation is less. The average width of crack 4 is 8.10 mm, slightly less than that of crack 2, and the width distribution is extremely concentrated, with almost no obvious fluctuation.
[0166] Figure 3 (c) and the visualization results (such as Figure 5 shown) show that the angle ranges of crack 1 and crack 2 are relatively wide and the directionality is diverse, where the angle distribution of crack 1 is from 13.45° to 156.80°, and that of crack 2 is from 50.60° to 110.24°. On the contrary, the angle ranges of crack 3 and crack 4 are relatively narrow, showing strong direction consistency. The angle range of crack 3 is from 76.45° to 107.49°, and the angle of crack 4 is 102.06°.
[0167] Basis for reliability: By combining advanced deep learning models and dynamic feature tracking technology, the characterization results of crack parameters are not only close to manual annotation in terms of accuracy (see Table 1), especially in the application of crack dynamic evolution analysis and large-scale crack identification, and even exceed the capabilities of traditional image processing in aspects such as crack boundary processing.
[0168] Table 1 Quantitative characterization results of fracture length
[0169]
[0170] Example 3
[0171] The present invention provides a method and system for intelligent identification and quantitative characterization of fractures in forward-looking borehole video images. The system is used to implement the foregoing method, including a detection module, a segmentation module, an extraction module, and a calculation module;
[0172] The detection module is used to detect fractures in the forward-looking borehole video by using an improved YOLOv8 model and generate fracture region bounding boxes;
[0173] The segmentation module is used to segment the region within the fracture region bounding box by using a Mask R-CNN model, introduce a boundary-weighted binary cross-entropy loss function and design a weight matrix W to assign weights to the fracture boundary region, and obtain a fracture mask;
[0174] The extraction module is used to extract the fracture morphological features in the fracture mask;
[0175] The calculation module is used to implement cross-frame matching and tracking of fractures by using feature matching SURF and the Lucas-Kanade optical flow method based on the fracture region and fracture morphological feature parameters, and calculate the fracture length and attitude change.
[0176] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for intelligent recognition and quantitative characterization of fissures in forward-looking drilling video images, characterized in that, The method includes: Step 1: Use the improved YOLOv8 model to detect fractures in the forward-looking borehole video and generate fracture region bounding boxes; Step 2: Use the optimized Mask R-CNN model to segment the region within the fracture region bounding box to obtain a fracture mask; Step 3: Extract the fracture morphological features from the fracture mask; Step 4: Based on the fracture region and fracture morphological feature parameters, use feature matching SURF and the Lucas-Kanade optical flow method to achieve cross-frame matching and tracking of fractures, and calculate the fracture length and attitude changes.
2. The method according to claim 1, characterized in that The method of using the improved YOLOv8 model to detect fractures in the forward-looking borehole video includes: Add a Capsule layer after the C2f layer in the Neck part of the YOLOv8 model, and transmit and aggregate feature information through a dynamic routing mechanism to capture the spatial relationship of fractures.
3. The method according to claim 1, wherein The method of using the optimized Mask R-CNN model to segment the region within the fracture region bounding box to obtain a fracture mask includes: Introduce a boundary-weighted binary cross-entropy loss function, and assign weights to the fracture boundary region by defining a weight matrix W; the expression is: where L WBCE is the weighted segmentation loss function; y i is the ground truth label, 0 represents the background, and 1 represents the crack; p i is the probability that the predicted pixel i belongs to the crack, and W i is the weight of pixel i in the weight matrix; The weight matrix is calculated based on fracture gradient information and the edge region, guiding the Mask R-CNN model to focus on the fracture boundary region to obtain a fracture mask.
4. The method according to claim 3, wherein The method of calculating the weight matrix based on fracture gradient information and the edge region to guide the Mask R-CNN model to focus on the fracture boundary region includes: Calculate the horizontal and vertical gradients G(x,y) in the image, i.e., the borehole video frame, using the Sobel operator and generate an edge intensity map, and analyze the edge intensity at each position in the image. The gradient calculation is: where G x (x, y) and G y (x, y) are the gradients of the image in the horizontal and vertical directions, respectively; Use the Canny edge detection method to extract the edge information C(x,y) in the image, and combine it with distance transformation to calculate the distance from each pixel to the nearest boundary; combine the gradient magnitude with the Canny edge detection result, and introduce a weight based on distance transformation. Map the distance value to the [0,1] interval through normalization operation to assign weights to the fracture edge region. Among them, the distance transformation is: In the formula, dist(x,y) is the distance from the current pixel to the nearest boundary; max(x,y) is the maximum distance from all pixels in the image to the nearest boundary; Define the weight matrix W(x,y) as the product of the gradient magnitude and the edge detection result, and combine the result of distance transformation to guide the Mask R-CNN model to focus on the fracture boundary region. The calculation formula is: In the formula, α and β are adjustment parameters that control the relative importance of gradient information, edge detection, and distance transformation; min(W) and max(W) are the minimum and maximum values of all pixels in the weight matrix respectively.
5. The method according to claim 1, wherein The method of extracting the fracture morphological features from the fracture mask includes: Count the number of non-zero pixels in the binary image to calculate the area of the crack, and calculate the perimeter using the set of foreground pixel boundary points; use the Zhang-Suen thinning algorithm to extract the skeleton of the crack, and identify the bifurcation points in the skeleton through neighbor analysis; calculate the crack width through distance transformation, and process the width of the bifurcation points to obtain the average width of the crack; at the same time, use the image moment method to extract the centroid position of the crack, and fit the crack ellipse to determine the crack direction.
6. The method according to claim 5, wherein The method for counting the number of non-zero pixels in the binary image to calculate the area of the crack includes: Assume that the input binary image is I(x, y), and the size of the image is H×D, where H is the height and D is the width. Traverse each pixel point of the crack mask image, count the number of non-zero pixels, and convert it to the actual area unit. The calculation formula is: where countMonZero(I) is the area of the crack pixels; 1 I(i,j) ≠0 is the indicator function that returns 1 if I(i,j) is not 0, otherwise returns 0; I(i,j) is the pixel value of the image I at (i,j); The method for calculating the perimeter using the set of foreground pixel boundary points includes: Assume that the input binary image is I(x, y), where I(x, y) = 255 represents foreground pixels and I(x, y) = 0 represents background pixels. The boundary of the foreground region forms a contour, which is composed of a set of points {P0, P1, …, P n-1}. The perimeter L is expressed as the total length of the point set, and the calculation formula is: where L is the total length of the set of crack contour points; dist(P i , P i+1 ) is the Euclidean distance between two adjacent contour points; The method for using the Zhang-Suen thinning algorithm to extract the skeleton of the crack includes: Gradually thin the foreground region in the binary image I(x, y) to generate a skeleton with a width of one pixel while maintaining the topological structure of the fissures. For each pixel point P i , is deleted when the preset conditions are met, and finally a skeleton structure is generated. The expression of the preset conditions is: In the formula, a pixel value of 255 represents a crack, and 0 represents the background; A(P i ) is the number of transitions from the background to the foreground in the 8-neighborhood of pixel point P i ; B(P1) is the number of foreground pixels in the 8-neighborhood of pixel point P i ; D i (1) is the removal condition check function D i (1) = P2·P4·P6 = 0 and P4·P6·P8 = 0, D i (1) = P2·P4·P8 = 0 and P2·P6·P8 = 0; The method for identifying the bifurcation points in the skeleton through neighbor analysis includes: For each pixel point P in the skeleton Li perform neighbor analysis to calculate the number N(P Li ) of neighbor pixels. If N(P Li ) > 2, then mark this pixel point as a bifurcation point. Assume that the neighborhood pixel distribution of a certain point P Li is as follows. Calculate N(P Li ) = 4 > 2. Therefore, P Li is a bifurcation point; Where, P Li is the skeleton pixel point; 1 is the neighbor pixel; The method for calculating the crack width through distance transformation includes: Perform a distance transformation on the binarized mask image to calculate the Euclidean distance dist(P i ) from each foreground pixel to the nearest background pixel. The calculation formula is as follows: Where β is the set of all background pixel points; P i (x i , y i ) is the coordinate of the foreground pixel point; P j (x j , y j ) is the coordinate of the background pixel point; For the contour point P i , the crack width w i is obtained by multiplying the distance value by 2, and the calculation formula is: w i = 2·dist(P i ) For the bifurcation point P Li , calculate the width in each direction respectively. Assuming there are K directions, the width is: where dist j (P Li ) is the nearest background distance of the bifurcation point in the j-th direction; The method for using the image moment method to extract the centroid position of the crack includes: Calculate the coordinates of the crack centroid using the image moment, and obtain the center position of the crack through geometric analysis techniques. The calculation formula is: Zero-order moment: M 00 = ∑ x ∑ y I(x, y), first-order moment: Centroid coordinates: Where M ij is the (i + j)-th moment of the image; i and j are the orders of the moment; C x and C y are the coordinates of the centroid of the crack; The method for fitting the crack ellipse to determine the crack direction includes: Based on the set of contour points \(\{(x i , y i )\}, fit an ellipse in the sense of least squares, and adjust the parameters of the ellipse to minimize the sum of the squares of the distances from the contour points to the ellipse boundary, and determine the fracture direction. The calculation formula is: where (x i , y i ) are the coordinate points of the image contour; (cx, cy) are the coordinates of the ellipse center; (a, b) are the major and minor axes of the ellipse; θ is the rotation angle of the ellipse, i.e., the required crack direction angle.
7. The method according to claim 1, wherein The method for implementing cross-frame matching and tracking of the crack using feature matching SURF and optical flow method Lucas-Kanade includes: In the crack mask obtained by two-stage network detection, use the SURF algorithm to extract the key feature points of the crack area, and calculate the feature descriptor v for inter-frame matching; Using the KNN algorithm, based on the Euclidean distance d(v i , v j ), find the nearest neighbor feature points of each feature point, thereby initially establishing the matching relationship of crack feature points between multiple frames; Use the optical flow method Lucas-Kanade to further track the position change of the feature points in the image sequence, especially in the scenario where the position change of the crack in consecutive frames is small. For the already matched feature points, calculate their motion vectors through the optical flow method to obtain the motion trajectory of the crack in consecutive frames and track the change of the crack; Combine the motion vectors of optical flow tracking with the results of feature matching, and frame by frame track the position and morphological changes of the crack feature points. For the feature points matched in adjacent frames, calculate the length of the crack by accumulating the distance between the feature points along the crack trajectory.
8. A forward-looking drilling video image crack intelligent recognition and quantitative characterization system, which is used to implement the method described in any one of claims 1-7, and is characterized in that, Detection module, segmentation module, extraction module and calculation module; The detection module is used to detect the cracks in the forward-looking borehole video using the improved YOLOv8 model and generate the crack area bounding box; The segmentation module is used to segment the area within the crack area bounding box using the optimized Mask R-CNN model to obtain the crack mask; The extraction module is used to extract the crack morphological features in the crack mask; The calculation module is used to implement cross-frame matching and tracking of the crack using feature matching SURF and optical flow method Lucas-Kanade based on the crack area and crack morphological feature parameters, and calculate the crack length and attitude change.
Citation Information
Patent Citations
Concrete bridge crack detection method based on computer vision
CN119130986A
Rock mass fracture identification method based on boundary detection neural network model
CN119130994A
Methods for the Segmentation of Lungs, Lung Vasculature and Lung Lobes from CT Data and Clinical Applications
US20220092791A1
Method and system for determining rock objects in petrographic data using machine learning
US20250076273A1
Image spectrum technology-based geological sketching method and system for tunnel face
WO2024083261A1
Cited By
Railway track crack defect detection method and device based on machine vision
CN121366160A