Rock mass three-dimensional structural surface intelligent identification method based on AI visual large model

By combining AI vision big model and motion recovery structure technology, the precise conversion from two-dimensional images to three-dimensional point clouds is achieved, and the efficiency and accuracy of rock mass structure surface recognition in complex geological environments is solved, providing efficient tools for geological surveys and engineering practices.

CN120279318AActive Publication Date: 2025-07-08ZHEJIANG UNIV +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510354089.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately identify rock mass structural surfaces in complex geological environments. Traditional methods have a trade-off between sample number and labor efficiency, and are limited by environmental and data requirements.

Method used

Combining AI visual big model and motion recovery structure technology, precise conversion from a two-dimensional image to a three-dimensional point cloud is achieved through image selection and synthesis, interactive structural surface detection, pixel point relationship generation and three-dimensional structural surface index.

Benefits of technology

It realizes high-precision and high-efficiency three-dimensional structural surface recognition of rock mass, provides efficient geological data annotation tools, suitable for engineering practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279318A_ABST
    Figure CN120279318A_ABST
Patent Text Reader

Abstract

The invention relates to the field of rock mass structural surface recognition, in particular to an intelligent recognition method for a rock mass three-dimensional structural surface based on an AI visual large model. An artificial intelligence visual large model is combined with a motion recovery structure technology, and three-dimensional structural surface identification is realized by using data of different photogrammetry devices. According to the method, an identification task is converted from point cloud to original image data through the incidence relation (often neglected by previous research) between image pixels and point cloud in a motion recovery structure, and the implementation difficulty is remarkably reduced. A benchmark test shows that the method realizes the average precision of 0.91 in a three-dimensional database, and 87 structural surfaces are identified. The result shows that the method has high precision and high efficiency, and provides a powerful tool for geological data labeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rock mass structural plane identification, and particularly to an intelligent identification method for three-dimensional rock mass structural planes based on an AI vision large model. Background Art

[0002] A rock mass is a complex fracture system, mainly composed of two components: structural planes and rock blocks. Structural planes play an important role in the mechanical properties of rock masses because they define the weak planes in the rock mass, and rock blocks separate and break along the structural planes. Therefore, characterizing the structural planes in a rock outcrop is an important step that requires collecting input information for further rock mechanics analysis and rock engineering design.

[0003] To characterize the structural planes visible on the surface of a rock outcrop, a typical set of parameters that are usually recommended to measure includes attitude, spacing, persistence, and roughness. These key information about rock structural planes help predict groundwater flow, evaluate rock mass quality, and provide better stability conditions for rock engineering. Traditionally, these parameters were measured manually in the field of rock outcrops, such as using a compass and inclinometer to measure the attitude of structural planes. However, these methods involve a trade-off between the number of samples and labor efficiency, and are usually difficult to access and dangerous. To overcome these limitations, three-dimensional digital outcrop models provide high-resolution digital twins of rock outcrops.

[0004] A digital outcrop model is a three-dimensional model of the outcrop surface represented by a point cloud, which can be further generated into a polygon mesh. Non-contact measurement, lidar, and photogrammetry techniques acquire the digital data required to create a digital outcrop model. Lidar directly obtains a digital outcrop model, while photogrammetry first captures multiple images and then uses three-dimensional reconstruction techniques, such as structure from motion, to generate a digital outcrop model. The main challenge in characterizing structural planes from digital outcrop models is first to identify them. However, the characteristics of digital outcrop model data make this task difficult. Digital outcrop models are non-intuitive, that is, they do not contain any geological features or geometries. In addition, they are full of interferences, such as weathering, vegetation interference, and outliers. These characteristics usually lead to lengthy solutions, including the need for manual preprocessing, continuous parameter adjustment, or model training across different sites.

[0005] In addition to the complexity of the digital outcrop model itself, the complexity of geological and environmental conditions further limits the universality of existing methods. Exogenous and endogenous geological processes (such as seismic activity, hydrodynamic action, and long-term erosion) affect the formation and evolution of rock masses, resulting in significant changes in the joint texture and geometric morphology of rock masses. Environmental factors (such as light, shadow, and vegetation noise) further increase the complexity of joint identification. Existing methods can be roughly divided into two categories: (1) identification methods based on digital outcrop models, which directly use digital outcrop models for joint identification; (2) vision-based identification methods, which rely on pattern information in the original images. These two methods usually follow two strategies: rule-based strategies and data-driven strategies. Rule-based strategies use specific rules to distinguish joints from other features; in the segmentation based on digital outcrop models, methods such as plane fitting, region growing, and clustering are included, while in vision-based segmentation, image processing techniques such as edge detection are usually adopted. Data-driven strategies use labeled datasets and deep learning models to learn joint features and identify them. However, both of these strategies have limitations: rule-based methods rely on thresholds, while data-driven methods rely on data, and each model or each algorithm parameter set is customized for a specific rock outcrop. Therefore, there is an urgent need for a general-purpose and low-cost method that can span multiple outcrops to overcome the limitations of the above methods.

[0006] The emergence of large AI vision models has revolutionized the field of computer vision because of their excellent generalization ability in a wide range of recognition tasks. Although limited data volume has traditionally restricted the robustness of data-driven models, large AI vision models can perform well in various recognition tasks even with little additional training. This is mainly attributed to the huge parameter space of large AI vision models and their pre-training on a wide range of datasets, enabling a single model to handle a wide range of recognition challenges. In addition, as the training data increases, the generalization ability of large AI vision models also increases. The above scenario is similar to the situation encountered in joint identification, which requires extracting detailed features from data with high environmental complexity and demanding high precision.

[0007] In view of the above situation, it is an urgent problem to be solved in the technical field to propose an intelligent recognition method for three-dimensional joints of rock masses based on large AI vision models. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to overcome the deficiencies in the prior art and provide an intelligent recognition method for three-dimensional joints of rock masses based on large AI vision models.

[0009] To solve the technical problem, the solution of the present invention is:

[0010] An intelligent recognition method for three-dimensional rock mass structural planes based on the large AI vision model. This method combines the large AI vision model with the structure from motion technology, and uses data from different photogrammetry devices to achieve three-dimensional structural plane recognition. By leveraging the correlation between image pixels and point clouds in structure from motion (which has often been overlooked in previous studies), this method transfers the recognition task from point clouds to the original image data, significantly reducing the implementation difficulty. Benchmark tests show that this method achieves an average accuracy of 0.91 in the three-dimensional database and identifies 87 structural planes. The results indicate that this method has high precision and efficiency, providing a powerful tool for geological data annotation. This invention lays the foundation for efficient geological surveys and opens up the way for the development of dedicated large-scale models in engineering practice.

[0011] The proposed method consists of two parallel processes: two-dimensional structural plane detection and three-dimensional reconstruction. This method is divided into four parts, which are introduced in sequence as follows:

[0012] (1) Image selection and synthesis

[0013] Due to the overlap between the captured images, image selection or synthesis is required to simplify two-dimensional detection. Photogrammetry devices are classified into mobile devices and fixed devices according to the shooting position.

[0014] For mobile devices, a subset of images will be selected for detection. For example, drones or handheld cameras. Using the pixel point relationships generated during the three-dimensional reconstruction process, the automated image selection program selects images or their regions based on the descending order of the number of valid pixels and the principle of non-overlapping three-dimensional points.

[0015] For fixed devices, such as cameras mounted on tripods, images are usually taken from more than two positions for three-dimensional reconstruction. The image synthesis editor developed by Microsoft is used to stitch the overlapping photos taken at a single camera position into a seamless high-resolution panoramic image, which is called a synthetic image. Therefore, a synthetic image can be constructed from a single position. Similarly, using the pixel point relationships generated during the three-dimensional reconstruction process, the automated image selection program selects the synthetic image or its region according to the above principles for two-dimensional structural plane detection.

[0016] (2) Interactive structural plane detection

[0017] For two-dimensional structural plane detection, different from traditional data-driven methods (such as deep learning) that require a large amount of labeled data for dataset construction and model training, the pre-trained large AI vision model reduces the need for retraining and lowers the threshold of geological surveys. Therefore, this method provides a unified solution for structural plane detection in various geological environments by leveraging the generalization ability of the large AI vision model.

[0018] The AI vision large model enables real-time interactive use. The model architecture consists of three core components: an image encoder, a prompt encoder, and a mask decoder. The image encoder calculates image embeddings, the prompt encoder provides prompt information, and the lightweight mask decoder combines this information from both parts to predict segmentation results. By decomposing the AI vision large model into these components, the same image embeddings can be reused with different prompt information, thus reducing redundant computational overhead. Given an image embedding and a prompt input, the mask decoder can predict a mask in only 55 milliseconds. Real-time and accurate two-dimensional structural plane recognition is achieved.

[0019] The prompt module of the AI vision large model enables interactive structural plane detection. This invention uses two types of prompts: points (foreground points are used to add regions; background points are used to remove regions) and boxes.

[0020] (3) Pixel-point relationship generation

[0021] The input is unordered overlapping images, and the output includes three parts: a colored sparse point cloud model, camera poses, and pixel-point relationships. The structure from motion technology generates a sparse three-dimensional point cloud by matching pixel features between images and calculating the pose of each image, and preserves the relationship between pixels and points without adding additional computational costs.

[0022] (4) Three-dimensional structural plane indexing

[0023] The pixel-point relationship is a bridge between image segmentation and point cloud segmentation. Therefore, three-dimensional structural planes can be obtained through automated indexing. Using the pixel-point geometric relationship between two dimensions and three dimensions, the corresponding three-dimensional point cloud is extracted from the two-dimensional image mask. After carefully selecting parameters to obtain satisfactory structure from motion modeling results, the pixel-point relationship is used to convert the two-dimensional structural plane segmentation result into three-dimensional point cloud data. This conversion is achieved through automated indexing, which can spatially locate the structural plane based on the established two-dimensional to three-dimensional relationship. Since each observed pixel corresponds to a uniquely determined three-dimensional structural point, the conversion process is simple and accurate.

[0024] For mobile devices, a single index connects the image with its point cloud. For fixed devices, a dual index is used to connect the image, its composite image, and the corresponding point cloud. The segmentation result is recognized on the composite image, so that it can be uniquely determined which pixel in the original image corresponds to a given pixel in the composite image. In addition, the image observations are indexed to their corresponding three-dimensional points, which are uniquely determined by the pixel-point relationship.

[0025] The working principle of this invention is as follows:

[0026] The recognition of three-dimensional structural planes in photogrammetry follows three different approaches. Different from lidar that directly provides three-dimensional point clouds, photogrammetry relies on two-dimensional images as the original data. For lidar, the recognition of structural planes is usually achieved through point segmentation. However, due to the lack of three-dimensional annotation data, existing point segmentation methods are limited, which restricts their performance and generalization ability. Therefore, the present invention chooses to use a large-scale model instead of point segmentation for image segmentation.

[0027] The requirements for the recognition of three-dimensional structural planes include accuracy, feasibility, generality, preferably low cost, and timely response. Different from autonomous vehicles that usually use RGB-D cameras to obtain depth information, photogrammetry technology requires three-dimensional reconstruction to obtain spatial data. Current three-dimensional reconstruction software usually uses structure from motion and multi-view stereo to generate depth maps. However, the present invention finds that multi-view stereo may not be suitable for rock mass reconstruction due to its inherent limitations. In addition, the depth fusion process in multi-view stereo complicates the establishment of two-dimensional to three-dimensional relationships. Therefore, the present invention does not recommend using multi-view stereo for reconstruction.

[0028] In view of these considerations, the proposed method adopts a third approach, combining the robust recognition ability and generalization of a large-scale model with the accuracy of pixel point relationships generated by structure from motion. This method effectively meets the technical requirements for the recognition of three-dimensional structural planes.

[0029] Compared with images, the inherent complexity of point clouds hinders the success of point segmentation in achieving image segmentation. Challenges such as inconsistent data formats, poor model scalability, and insufficient annotation masks further limit the performance of three-dimensional models. These problems have prevented previous researchers from obtaining results comparable to those of large AI vision models. Therefore, the present invention avoids direct point segmentation and instead adopts image segmentation. As mentioned before, this method not only performs the recognition task but also serves as a data engine for collecting large-scale annotation data from discontinuous point clouds.

[0030] Compared with the prior art, the beneficial effects of the present invention are:

[0031] 1. The method achieves an average precision of 0.91 in a three-dimensional database and identifies 87 structural planes. The results show that the method has high precision and high efficiency, providing a powerful tool for geological data annotation.

[0032] 2. For the first time, a large AI vision model is used for the analysis of geological structural planes.

[0033] 3. Pixel-to-point cloud mapping realizes three-dimensional segmentation from two-dimensional data.

[0034] 4. To verify the generalization in two-dimensional structural plane detection, the present invention introduces a two-dimensional database (a newly constructed dataset with different image distributions for evaluation), thereby obtaining data from different rock outcrops. The two-dimensional database contains 2,342 identified two-dimensional structural planes in 170 images, including 90 images from the Internet and 80 self-captured images, ensuring a wide representation of geological and environmental conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is the flowchart of the method of the present invention;

[0036] Figure 2 is the architecture diagram of the AI vision large model for interactive structural plane detection;

[0037] Figure 3 is the step-by-step demonstration diagram of the interactive structural plane detection of complex structural planes;

[0038] Figure 4 is the schematic diagram of the input, key steps, and output of the structure from motion for 3D rock mass reconstruction;

[0039] Figure 5a is the schematic diagram of the pixel point relationship in the structure from motion;

[0040] Figure 5b is the schematic diagram of the pixel point relationship data format in the structure from motion;

[0041] Figure 6 is the schematic diagram of the process of automatically indexing and extracting 3D structural planes using a fixed device;

[0042] Figure 7 is an example image with an overlaid mask from the two-dimensional database;

[0043] Figure 8a is the box plot of the intersection over union distribution of cross-point prompts (1–20); where the dots represent the mean values;

[0044] Figure 8b is the comparison of the mean intersection over union between the original database and the two-dimensional database;

[0045] Figure 9 is the segmentation result under the standard interactive segmentation protocol;

[0046] Figure 10a is the location map of the study area;

[0047] Figure 10b is the fixed-position device and image acquisition diagram;

[0048] Figure 10c is the 3D reconstruction result diagram;

[0049] Figure 10d It is a three-dimensional structural plane diagram with annotations; different structural planes are represented by colors;

[0050] Figure 11a It is the result of three-dimensional recognition of structural planes presented in the form of point clouds, and the structural planes are colored to match the corresponding annotated three-dimensional structural planes;

[0051] Figure 11b It is a semi-box plot and semi-violin plot of precision;

[0052] Figure 11c It is a semi-box plot and semi-violin plot of precision;

[0053] Figure 11d It is a schematic diagram of false positive points shown on the annotation results, and there is no overlap between them;

[0054] Figure 12 is for analyzing the results of three-dimensional structural plane recognition;

[0055] (a) Results of two-dimensional recognized structural planes;

[0056] (b) Scatter plot of precision and pixel area;

[0057] (c) Scatter plot of precision / pixel area and pixel area;

[0058] Figure 13 It is the frequency distribution and box plot of the attitude error (angle difference of the upper unit normal vector) between the recognized structural plane and the manually annotated structural plane. Specific implementation method

[0060] The following further elaborates on the present invention in conjunction with the attached drawings. The following specific implementation steps can enable those skilled in the art to understand the present invention more comprehensively, but do not limit the present invention in any form.

[0061] First of all, it should be noted that the technical solution of the present invention will involve a large amount of existing technologies, and their definitions or concepts are already well-known or mastered by those skilled in the art. Therefore, unless the present invention makes a special explanation for a specific meaning, for those with the same meaning as the existing well-known ones, the present invention will not be described one by one.

[0062] The intelligent recognition method for three-dimensional structural planes of rock masses based on the AI vision large model described in the present invention, as Figure 1 shown, includes the following steps:

[0063] (1) Image selection and synthesis

[0064] Since there is overlap between the captured images, image selection or synthesis is required to simplify two-dimensional detection. Photogrammetry equipment is divided into mobile equipment and fixed equipment according to the shooting position.

[0065] For mobile devices, a subset of images will be selected for detection. For example, for drones or hand-held cameras, using the pixel point relationships generated during the 3D reconstruction process, the automated image selection program selects images or their regions based on the descending order of the number of valid pixels and the principle of non-overlapping 3D points.

[0066] For fixed devices, such as cameras mounted on tripods, images are usually taken from more than two positions for 3D reconstruction. The image synthesis editor developed by Microsoft is used to stitch overlapping photos taken at a single camera position into a seamless high-resolution panorama, which is called a synthetic image. Therefore, a synthetic image can be constructed from a single position. Similarly, using the pixel point relationships generated during the 3D reconstruction process, the automated image selection program selects the synthetic image or its region according to the above principles for 2D structural plane detection.

[0067] (2) Interactive structural plane detection

[0068] For 2D structural plane detection, different from traditional data-driven methods (such as deep learning) that require a large amount of labeled data for dataset construction and model training, the pre-trained large AI vision model reduces the need for re-training and lowers the threshold of geological surveys. Therefore, this method provides a unified solution for structural plane detection in various geological environments by leveraging the generalization ability of the large AI vision model.

[0069] The large AI vision model enables real-time interactive use. As Figure 2 shown, the model architecture consists of three core components: an image encoder, a prompt encoder, and a mask decoder. The image encoder computes the image embedding, the prompt encoder provides the prompt information, and the lightweight mask decoder combines these two parts of information to predict the segmentation result. By decomposing the large AI vision model into these components, the same image embedding can be reused with different prompt information, thus reducing the redundant computational overhead. Given the image embedding and prompt input, the mask decoder can predict the mask within only 55 milliseconds. Real-time and accurate 2D structural plane recognition is achieved.

[0070] The prompt module of the large AI vision model enables interactive structural plane detection. Two types of prompts are used in the present invention: point selection (foreground points are used to add regions, marked in green; background points are used to remove regions, marked in red) and box selection. Figure 3Shows the step - by - step process of interactive structural plane detection on complex structural planes using only point prompts. The goal is to identify the structural plane marked as number 6 in the image. Due to its complex boundaries, blurred contours, partial invisibility, and location at the intersection of light and shadow, it is very difficult to clearly outline this structural plane. Finally, by using 5 foreground points and 5 background points, this structural plane was successfully detected.

[0071] (3) Pixel - point relationship generation

[0072] The structure - from - motion process is as Figure 4 shown. The input is unordered overlapping images, and the output includes three parts: a colored sparse point - cloud model, camera poses, and pixel - point relationships. The structure - from - motion technology generates a sparse three - dimensional point cloud by matching pixel features between images and calculating the pose of each image, and preserves the relationship between pixels and points without adding additional computational costs, as Figure 5a and Figure 5b shown.

[0073] (4) Three - dimensional structural plane indexing

[0074] Pixel - point relationships are the bridge between image segmentation and point - cloud segmentation. Therefore, three - dimensional structural planes can be obtained through automated indexing. Using the pixel - point geometric relationship between two - dimensional and three - dimensional, the corresponding three - dimensional point cloud is extracted from the two - dimensional image mask. After carefully selecting parameters to obtain a satisfactory structure - from - motion modeling result, the pixel - point relationship is used to convert the two - dimensional structural plane segmentation result into three - dimensional point - cloud data. This conversion is achieved through automated indexing, which can spatially locate the structural plane based on the established two - dimensional - three - dimensional relationship. Since each observed pixel corresponds to a uniquely determined three - dimensional structural point, the conversion process is simple and accurate.

[0075] For mobile devices, a single index connects the image with its point cloud. For stationary devices, a dual index is used to connect the image, its composite image, and the corresponding point cloud. The segmentation result is identified on the composite image, so that it can be uniquely determined which pixel in the original image corresponds to a given pixel in the composite image. In addition, image observations are indexed to their corresponding three - dimensional points, which are uniquely determined by the pixel - point relationship. Figure 6 Shows the process of obtaining three - dimensional structural planes from a fixed - position device using automated indexing.

[0076] Specific application examples:

[0077] Two examples are involved here.

[0078] 1. Two - dimensional database

[0079] To verify the generalization ability of two - dimensional structural plane detection, the present invention introduces a two - dimensional database (see Figure 7), This is a newly constructed diverse image distribution dataset for evaluation. The 2D database contains 2,342 identified 2D structural planes in 170 images, where 90 images are from the Internet and 80 are self-taken images, ensuring a wide representation of geological and environmental conditions.

[0080] The masks in the 2D database were manually annotated using the LabelMe tool as the ground truth. LabelMe is an open-source image annotation tool widely used in object detection, semantic segmentation, and panoramic segmentation tasks. As Figure 8a and Figure 8b shown, the 2D database captures a wide range of structural plane textures and geometries, as well as varying environmental factors (such as lighting, shadows, and vegetation interference), reflecting the complexity of the structure and the environment.

[0081] The present invention uses all 170 image samples and 2,342 identified 2D structural planes in the 2D database for evaluation. The large AI vision model achieved an average intersection over union (IoU) of 78% at 11-point prompts, and its performance then stabilized despite an increase in the number of prompts, as shown in the distribution box plot in Figure 8a . Figure 8b The comparison results in

[0082] show the zero-shot ability of the large AI vision model in structural plane recognition. This confirms that the large AI vision model, although not specifically trained for structural planes, still has strong generalization ability and can recognize unseen structural plane images. Figure 9 As shown in

[0083] 2. 3D Database

[0084] To verify the accuracy of the proposed method in 3D structural plane recognition, data was collected using a fixed device to form a 3D database. The 3D structural planes in the point cloud were manually annotated using CloudCompare as the ground truth. Compared with mobile devices, fixed devices introduce greater complexity in the workflow and indexing process. This complexity is the main reason for choosing the dataset of fixed devices to evaluate the recognition accuracy of the method.

[0085] The 3D database (asFigures 10a to 10d ) was obtained through fixed equipment, specifically using a Nikon D300s camera with a resolution of 4256×2832 pixels. The images were taken at the Sunbury Rock Bench beside Highway 15, a provincial road near Kingston, Ontario, Canada, with a focus on the granite rock mass exposed at the roadside. A total of 471 images were taken from three fixed positions, with 146, 168, and 157 images taken at each position respectively. Using the method of three-dimensional discontinuity identification, a sparse point cloud containing 392,229 points was reconstructed, and the pixel-point relationship was established. The point cloud resolution Resolution is usually defined as the average distance between adjacent points and is calculated using formula (1), where M represents the number of point pairs, and d i represents the distance between the i-th pair of nearest points. Due to the lack of GPS data or known object sizes, no real-world scale calibration was applied during the reconstruction process. Therefore, the resolution is expressed in local coordinate units, and the calculated result is 0.0013. Although the real-world scale cannot be obtained, this does not affect the verification of the identification accuracy. Manually annotating identified 87 discontinuities of different sizes and attitudes, establishing a benchmark for the method. The unannotated areas were classified as interference due to excessive fragmentation or lack of sufficient three-dimensional information.

[0086]

[0087] The present invention evaluates the accuracy of discontinuity segmentation by measuring the intersection over union (IoU) and precision between the predicted point cloud and the ground truth point cloud. IoU is a commonly used metric in segmentation tasks to quantify the overlap between the predicted point cloud and the ground truth point cloud. Its value ranges from 0 to 1, where 1 represents perfect alignment and 0 represents no overlap. On the other hand, precision evaluates the reliability of the model prediction by measuring the proportion of correctly identified points among all predicted points. While IoU reflects the overall accuracy of the segmentation, precision reflects the method's ability to minimize false positives.

[0088] To further evaluate the segmentation performance, the present invention also calculates the mean intersection over union (mIoU) and mean precision. The mIoU is obtained by averaging the IoU scores of all discontinuities in the dataset, providing a global assessment of the segmentation quality. The mean precision is derived from the precision-recall curve and is used to measure the method's performance at different confidence thresholds, providing a comprehensive assessment of its prediction ability.

[0089] The method successfully identified 87 discontinuities (see Figure 11a ), achieving a mean precision of 0.91 and a mean intersection over union of 0.70 in three-dimensional discontinuity identification. Figure 11b and Figure 11cThe distribution of intersection over union and precision is shown using half box plots and half violin plots, highlighting the extreme values, interquartile range, median, and mean. Approximately 70% of the precision in identifying 3D structural planes is higher than 0.90, and approximately 75% of the intersection over union exceeds 0.60. These results indicate that the method maintains high accuracy in 3D structural plane identification, providing a practical and computationally efficient tool for large-scale geological annotation.

[0090] Figure 11d The false detection points (marked in black) and the complete annotated point cloud are shown. The false detection points mainly cluster near the annotated structural planes (marked in yellow), indicating a strong consistency between the segmentation results and the ground truth. The main sources of these errors include inaccuracies in 2D interactive structural plane detection and subjectivity in manual 3D annotation. Directly segmenting structural planes in the point cloud usually requires frequent adjustment of the viewing angle, making manual annotation both time-consuming and prone to blurred boundaries. Fundamentally, the method operates based on an image segmentation method, while the ground truth annotation is generated through direct point cloud segmentation. Both methods utilize the 3D reconstruction results, but the method combines 3D reconstruction with image segmentation, while the ground truth segmentation is performed sequentially after 3D reconstruction. This difference essentially explains the similarity between the two segmentation results.

[0091] By analyzing the structural plane with relatively low recognition accuracy (number 83), the present invention further explores the 3D structural plane recognition results. The structural plane numbered 83 has the largest pixel area among all low-precision structural planes. Figure 12a The results of 2D interactive structural plane detection are shown.

[0092] Further analyze the relationship between pixel area and precision, Figure 12b and Figure 12c respectively show the scatter plot of precision vs. pixel area and the scatter plot of precision / pixel area vs. pixel area. The fitting function y = 0.6332x -0.962 The results show that there is a strong negative correlation between precision / pixel area and pixel area (the fitted R 2 value is 0.9565). This indicates that there is no direct linear relationship between precision and pixel area, thus confirming that the method can maintain stable precision in identifying structural planes of different sizes. This robustness in structural plane identification highlights the reliability of the method in identifying geological structures at different scales.

[0093] To further verify the accuracy of structural plane attitude recognition, the present invention uses CloudCompare software to automatically calculate the attitude of the structural planes. The attitude of each structural plane is converted into a unit normal vector. Subsequently, the angular difference, i.e., the attitude error γ, between the identified structural plane and the manually annotated structural plane is calculated through formula (2). Where n proposed_frameworkis the upper unit normal vector of the identified structural plane, n proposed_framework is the upper unit normal vector of the manually marked structural plane.

[0094]

[0095] Figure 13 shows the frequency distribution and box plot of the attitude error (unit normal vector difference). From Figure 13 it can be seen that the error range is between 0° and 3.14°, indicating that the deviation between the identified structural plane and the manually marked structural plane is relatively small, which is consistent with the results discussed above. This shows that the method can accurately capture the attitude of the structural plane. The fundamental reason for the difference is that the point clouds of the structural planes are not exactly the same. This result further verifies the accuracy of the proposed method.

[0096] Note: The actual scope of the present invention not only includes the specific embodiments disclosed above, but also includes all equivalent solutions that implement or execute the present invention under the claims.

Claims

1. An intelligent recognition method for the three-dimensional rock mass structural plane based on the AI vision large model, characterized in that Combine the AI vision large model with the structure from motion technology to achieve 3D structural plane recognition using data from different photogrammetry devices; through the correlation between image pixels and point clouds in the structure from motion, transfer the recognition task from point clouds to the original image data.

2. The intelligent recognition method for three-dimensional rock mass structural planes based on an AI vision large model according to claim 1, wherein It includes two parallel processes: 2D structural plane detection and 3D reconstruction, which are divided into four parts: (1) Image selection and synthesis For the overlap existing between the captured images, use image selection or synthesis to simplify the 2D detection; (2) Interactive structural plane detection For 2D structural plane detection, use a pre-trained AI vision large model; (3) Pixel-point relationship generation The input is unordered overlapping images, and the output includes three parts: a colored sparse point cloud model, camera poses, and pixel-point relationships; use the structure from motion technology to generate a sparse 3D point cloud by matching pixel features between images and calculating the pose of each image; (4) 3D structural plane indexing Obtain the 3D structural plane through automated indexing; use the pixel-point geometric relationship between 2D and 3D to extract the corresponding 3D point cloud from the 2D image mask; After selecting parameters to obtain a satisfactory structure from motion modeling result, the pixel-point relationship is used to convert the 2D structural plane segmentation result into 3D point cloud data; This conversion is achieved through automated indexing, which can spatially locate the structural plane based on the established 2D-3D relationship.

3. An intelligent recognition method for three-dimensional structural planes of rock masses based on an AI vision large model according to claim 2, characterized in that (1) In image selection and synthesis, photogrammetry devices are divided into mobile devices and fixed devices according to the shooting location; For mobile devices, select a subset of images for detection; using the pixel-point relationships generated during the 3D reconstruction process, the automated image selection program selects images or their regions based on the descending order of the number of valid pixels and the principle of non-overlapping 3D points; For fixed devices, capture images from more than two positions for 3D reconstruction; stitch the overlapping photos taken at a single camera position into a seamless high-resolution panoramic image, which is called a synthetic image; using the pixel-point relationships generated during the 3D reconstruction process, the automated image selection program selects the synthetic image or its region according to the above principle for 2D structural plane detection.

4. An intelligent recognition method for three-dimensional rock mass structural planes based on an AI vision large model according to claim 2, characterized in that, (2) In interactive structural plane detection, the AI vision large model architecture consists of three core components: an image encoder, a prompt encoder, and a mask decoder; the image encoder calculates the image embedding, the prompt encoder provides prompt information, and the lightweight mask decoder combines these two parts of information to predict the segmentation result; By decomposing the AI vision large model into these components, the same image embedding and different prompt information are reused to achieve real-time and accurate 2D structural plane recognition; The prompt module of the AI vision large model enables interactive structural plane detection; use two types of prompts: points and boxes; points include: foreground points for adding regions and background points for removing regions.

5. The intelligent recognition method for three-dimensional rock mass structural planes based on an AI vision large model according to claim 2, characterized in that, (4) In 3D structural plane indexing, for mobile devices, a single index connects the image with its point cloud; for fixed devices, a dual index is used to connect the image, its synthetic image, and the corresponding point cloud; The segmentation results are recognized on the composite image, enabling the unique determination of which pixel in the original image corresponds to a given pixel in the composite image; additionally, the image observations are indexed to their corresponding 3D points, which are uniquely determined by the pixel-point relationship.

Citation Information

Patent Citations

  • Rapid logging method for high cut slope geology based on three-dimensional point cloud reconstruction technology

    CN109961510A

  • Landslide mass three-dimensional reconstruction and volume calculation method based on mobile photographic image

    CN114596347A

  • Rock mass structural surface identification method and terminal thereof, and storage medium

    CN116824571A

  • Rock mass structural surface recognition method, device and equipment and readable storage medium

    CN117456280A

  • Rock surface fracture characterization method based on data fusion

    CN118298097A