Rock mass three-dimensional structural plane intelligent identification method based on AI vision large model

By combining AI visual large models with motion recovery structure technology, high-precision conversion from two-dimensional images to three-dimensional point clouds is achieved, solving the problems of high efficiency and universality in rock structural surface identification, and providing an efficient tool for geological surveys and engineering practices.

CN120279318BActive Publication Date: 2025-10-10ZHEJIANG UNIV +1

Patent Information

Application Number
CN202510354089.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-10-10
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently identify rock structural surfaces in complex environments. Traditional methods face a trade-off between sample quantity and labor efficiency, and are limited by the complexity of geological and environmental conditions. Existing methods lack universality and high precision.

Method used

Combining AI visual large models with motion recovery structure technology, high-precision conversion from two-dimensional images to three-dimensional point clouds is achieved through image selection and synthesis, interactive structural surface detection, pixel point relationship generation and three-dimensional structural surface indexing, reducing implementation difficulty and computational cost.

Benefits of technology

It achieved an average accuracy of 0.91, identified 87 structural planes, and provided an efficient geological data annotation tool suitable for the development of large-scale models in engineering practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279318B_ABST
    Figure CN120279318B_ABST
Patent Text Reader

Abstract

The present application relates to the field of rock mass structure plane identification, and in particular to a rock mass three-dimensional structure plane intelligent identification method based on an AI vision large model. The method combines an artificial intelligence vision large model with a motion recovery structure technology and uses data of different photogrammetry devices to realize three-dimensional structure plane identification. The method converts the identification task from point cloud to original image data by using the correlation between image pixels and point cloud in the motion recovery structure (which is often ignored by previous research), thereby significantly reducing the implementation difficulty. Benchmark tests show that the method achieves an average accuracy of 0.91 in a three-dimensional database and identifies 87 structure planes. The results show that the method has high precision and high efficiency and provides a powerful tool for geological data labeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rock mass structural surface recognition, and in particular to a method for intelligently recognizing three-dimensional rock mass structural surfaces based on an AI visual large model. Background Art

[0002] A rock mass is a complex fracture system composed primarily of two components: structural planes and structural bodies. Structural planes play an important role in the mechanical properties of a rock mass because they define the planes of weakness along which rock blocks separate and fail. Therefore, characterizing structural planes in rock mass outcrops is an important step in gathering input information for further rock mechanical analysis and rock engineering design.

[0003] To characterize the structural surfaces visible on rock outcrop surfaces, a typical set of parameters that are often recommended for measurement include occurrence, spacing, persistence, and roughness. This critical information about rock structural surfaces helps predict groundwater flow, assess rock mass quality, and provide better stability conditions for rock engineering. Traditionally, these parameters are measured manually on-site at rock outcrops, for example using a compass and inclinometer to measure the occurrence of structural surfaces. However, these methods suffer from a trade-off between sample quantity and labor efficiency and are often inaccessible and dangerous. To overcome these limitations, three-dimensional digital outcrop models provide a high-resolution digital twin of the rock outcrop.

[0004] A digital outcrop model is a three-dimensional representation of the outcrop surface represented as a point cloud, which can be further converted into a polygonal mesh. Non-contact surveying, LiDAR, and photogrammetry techniques acquire the digital data necessary to create digital outcrop models. LiDAR directly acquires the digital outcrop model, while photogrammetry first captures multiple images and then uses 3D reconstruction techniques, such as structure from motion, to generate the digital outcrop model. The primary challenge in characterizing structural surfaces from digital outcrop models is first identifying them. However, the characteristics of digital outcrop model data make this task difficult. Digital outcrop models are non-intuitive, meaning they do not contain any geological features or geometry. Furthermore, they are full of artifacts such as weathering, vegetation interference, and outliers. These characteristics often lead to lengthy solutions, requiring manual preprocessing, continuous parameter adjustment, or model training across different sites.

[0005] In addition to the complexity of the digital outcrop model itself, the complexity of geological and environmental conditions further limits the universality of existing methods. Exogenous and endogenous geological processes (such as seismic activity, hydrodynamic action, and long-term erosion) affect the formation and evolution of rock masses, resulting in significant changes in the texture and geometry of the structural surfaces of the rock masses. Environmental factors (such as illumination, shadows, and vegetation noise) further increase the complexity of structural surface recognition. Existing methods can be roughly divided into two categories: (1) recognition methods based on digital outcrop models, which directly use digital outcrop models for structural surface recognition; (2) vision-based recognition methods, which rely on pattern information in the original image. These two methods generally follow two strategies: rule-based strategies and data-driven strategies. Rule-based strategies use specific rules to distinguish structural surfaces from other features; in segmentation based on digital outcrop models, methods such as plane fitting, region growing, and clustering are included, while in vision-based segmentation, image processing techniques such as edge detection are usually used. Data-driven strategies use labeled datasets and deep learning models to learn structural surface features and recognize them. However, both strategies have limitations: rule-based methods rely on thresholds, while data-driven methods rely on data, and each model or each algorithm parameter set is customized for a specific rock outcrop. Therefore, a universal, low-cost method that can be applied across multiple outcrops is urgently needed to overcome the limitations of the above methods.

[0006] The emergence of AI vision big models has revolutionized the field of computer vision by demonstrating remarkable generalizability across a wide range of recognition tasks. While limited data has traditionally limited the robustness of data-driven models, AI vision big models can now excel across a wide range of recognition tasks, even with little additional training. This is largely due to the vast parameter space of AI vision big models and their pre-training on a wide range of datasets, enabling a single model to address a wide range of recognition challenges. Furthermore, the generalizability of AI vision big models increases with the amount of training data. The aforementioned scenarios are similar to those encountered in structural surface recognition, which requires the extraction of detailed features from data with high environmental complexity and demands high precision.

[0007] In view of the above situation, proposing an intelligent identification method for three-dimensional structural surfaces of rock masses based on AI visual large models is an urgent problem to be solved in this technical field. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to overcome the deficiencies in the prior art and provide a method for intelligently identifying three-dimensional structural surfaces of rock masses based on an AI visual large model.

[0009] To solve the technical problem, the solution of the present invention is:

[0010] A method for intelligent identification of three-dimensional structural surfaces of rock masses based on an AI visual large model. This method combines the AI ​​visual large model with structure-from-motion technology and uses data from different photogrammetric equipment to achieve three-dimensional structural surface identification. This method converts the identification task from point clouds to raw image data by exploiting the correlation between image pixels and point clouds in structure-from-motion (often overlooked in previous studies), significantly reducing the difficulty of implementation. Benchmark tests show that this method achieved an average accuracy of 0.91 in a three-dimensional database and identified 87 structural surfaces. The results demonstrate that this method has high precision and efficiency, providing a powerful tool for geological data annotation. This invention lays the foundation for efficient geological surveys and paves the way for the development of specialized large-scale models in engineering practice.

[0011] The proposed method consists of two parallel processes: 2D structural surface detection and 3D reconstruction. The method is divided into four parts, which are introduced as follows:

[0012] (1) Image selection and synthesis

[0013] Since there is overlap between the captured images, image selection or synthesis is required to simplify 2D inspection. Photogrammetry equipment is divided into mobile equipment and fixed equipment according to the shooting location.

[0014] For mobile devices, such as drones or handheld cameras, a subset of images is selected for inspection. Leveraging the pixel relationships generated during 3D reconstruction, an automated image selection process selects images or regions based on descending order of valid pixel count and non-overlapping 3D points.

[0015] For fixed equipment, such as a camera mounted on a tripod, images are typically captured from more than two positions for 3D reconstruction. An image composite editor developed by Microsoft is used to stitch overlapping photos taken from a single camera position into a seamless, high-resolution panorama, known as a composite image. Thus, a composite image can be constructed from a single position. Similarly, using the pixel relationships generated during 3D reconstruction, an automated image selection process selects a composite image or region thereof for 2D structural surface detection based on the aforementioned principles.

[0016] (2) Interactive structural surface detection

[0017] For 2D structural surface detection, unlike traditional data-driven methods (such as deep learning), which require large amounts of labeled data for dataset construction and model training, pre-trained AI visual large models reduce the need for retraining and lower the threshold for geological surveys. Therefore, this approach, by leveraging the generalization capabilities of AI visual large models, provides a unified solution for structural surface detection in various geological environments.

[0018] The AI ​​Vision Big Model enables real-time interactive use. The model architecture consists of three core components: an image encoder, a hint encoder, and a mask decoder. The image encoder computes image embeddings, the hint encoder provides hint information, and the lightweight mask decoder combines these two pieces of information to predict segmentation results. By breaking down the AI ​​Vision Big Model into these components, the same image embedding can be reused with different hint information, reducing redundant computational overhead. Given an image embedding and hint input, the mask decoder can predict a mask in just 55 milliseconds. This enables real-time and accurate recognition of two-dimensional structural surfaces.

[0019] The prompt module of the AI ​​visual large model can realize interactive structural surface detection. The present invention uses two types of prompts: points (foreground points are used to add areas; background points are used to remove areas) and boxes.

[0020] (3) Pixel point relationship generation

[0021] The input is unordered, overlapping images, and the output consists of three components: a colored sparse point cloud model, the camera pose, and the pixel-point relationships. Structure from Motion (SM) generates a sparse 3D point cloud by matching pixel features between images and calculating the pose of each image. This preserves the relationships between pixels and points without incurring additional computational cost.

[0022] (4) 3D structural surface index

[0023] The pixel-point relationship bridges the gap between image segmentation and point cloud segmentation. Therefore, automated indexing allows for the acquisition of 3D structural surfaces. Utilizing the pixel-point geometric relationship between 2D and 3D, the corresponding 3D point cloud is extracted from the 2D image mask. After carefully selecting parameters to achieve satisfactory structure-from-motion modeling results, the pixel-point relationship is used to convert the 2D structural surface segmentation results into 3D point cloud data. This conversion is achieved through automated indexing, enabling the spatial localization of structural surfaces based on the established 2D-3D relationship. Because each observed pixel corresponds to a uniquely identified 3D structural point, the conversion process is simple and precise.

[0024] For mobile devices, a single index links the image to its point cloud. For stationary devices, a dual index connects the image, its composite image, and the corresponding point cloud. Segmentation results are identified on the composite image, allowing unique identification of which pixel in the original image corresponds to a given pixel in the composite image. Furthermore, image observations are indexed to their corresponding 3D points, which are uniquely identified through pixel-to-point relationships.

[0025] The working principle of the present invention is:

[0026] Three-dimensional structural surface recognition in photogrammetry follows three distinct approaches. Unlike LiDAR, which directly provides 3D point clouds, photogrammetry relies on 2D images as raw data. For LiDAR, structural surface recognition is typically achieved through point segmentation. However, existing point segmentation methods are limited by the lack of 3D annotated data, which restricts their performance and generalization capabilities. Therefore, this paper chooses to use large-scale models rather than point segmentation for image segmentation.

[0027] Requirements for three-dimensional structural surface recognition include accuracy, feasibility, versatility, preferably low cost, and timely response. Unlike self-driving cars, which typically use RGB-D cameras to obtain depth information, photogrammetry technology requires three-dimensional reconstruction to obtain spatial data. Current three-dimensional reconstruction software typically uses motion recovery structure and multi-view stereo to generate depth maps. However, the present invention finds that multi-view stereo may not be suitable for rock mass reconstruction due to its inherent limitations. In addition, the depth fusion process in multi-view stereo complicates the establishment of two-dimensional-three-dimensional relationships. Therefore, the present invention does not recommend the use of multi-view stereo for reconstruction.

[0028] In light of these considerations, the proposed method adopts a third approach, combining the robust recognition and generalization of large-scale models with the precision of pixel-wise relationships generated by structure-from-motion. This method effectively meets the technical requirements for 3D structural surface recognition.

[0029] Compared to images, the inherent complexity of point clouds hinders the successful implementation of image segmentation using point segmentation. Challenges such as inconsistent data formats, poor model scalability, and insufficient annotation masks further limit the performance of 3D models. These issues have prevented previous approaches from achieving results comparable to large AI vision models. Therefore, this paper avoids direct point segmentation and instead employs image segmentation. As previously described, this method not only performs the recognition task but also serves as a data engine for collecting large-scale annotated data from discontinuous point clouds.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] 1. This method achieved an average accuracy of 0.91 in the 3D database and identified 87 structural planes. The results show that this method has high accuracy and efficiency, providing a powerful tool for geological data annotation.

[0032] 2. The first use of AI visual large models for geological structural surface analysis.

[0033] 3. Pixel to point cloud mapping enables 3D segmentation from 2D data.

[0034] 4. To verify the generalizability of 2D structural surface detection, we introduced a 2D database (a newly constructed dataset with a diverse image distribution for evaluation), capturing data from diverse rock outcrops. The 2D database contains 2,342 identified 2D structural surfaces across 170 images, including 90 from the internet and 80 self-captured images, ensuring a broad representation of geological and environmental conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a flow chart of the method of the present invention;

[0036] Figure 2 This is the architecture diagram of the AI ​​vision model for interactive structural surface detection;

[0037] Figure 3 A step-by-step demonstration of interactive structural surface detection for complex structural surfaces;

[0038] Figure 4 Schematic diagram of the input, key steps, and output of structure from motion recovery for 3D rock mass reconstruction;

[0039] Figure 5a The relationship between pixels in motion recovery structure is shown as a schematic diagram;

[0040] Figure 5b This is a schematic diagram of the pixel point relationship data format in motion recovery structure;

[0041] Figure 6 Schematic diagram of the process of extracting three-dimensional structural surfaces using automatic indexing of stationary equipment;

[0042] Figure 7 is an example image with superimposed masks from a 2D database;

[0043] Figure 8a Box plots of the intersection-over-union ratio distribution across the cue points (1–20); the points represent the mean;

[0044] Figure 8b is the average intersection-over-union comparison between the original database and the two-dimensional database;

[0045] Figure 9 is the segmentation result under the standard interactive segmentation protocol;

[0046] Figure 10a This is the location map of the study area;

[0047] Figure 10b Acquire maps for fixed-position equipment and images;

[0048] Figure 10c This is the three-dimensional reconstruction result image;

[0049] Figure 10d It is a three-dimensional structural surface diagram with annotations; different structural surfaces are represented by colors;

[0050] Figure 11a The 3D identified structural surface results are presented in the form of point clouds. The structural surfaces are colored to match the corresponding annotated 3D structural surfaces.

[0051] Figure 11b Half box plot and half violin plot for precision efficiency;

[0052] Figure 11c Half box plot and half violin plot for precision;

[0053] Figure 11d The figure shows the false positive points in the annotation results. There is no overlap between them.

[0054] Figure 12 shows the results of analyzing the three-dimensional structural surface recognition;

[0055] (a) Result of two-dimensional identification of structural surface;

[0056] (b) Scatter plot of accuracy and pixel area;

[0057] (c) Scatter plot of accuracy / pixel area versus pixel area;

[0058] Figure 13 Frequency distribution and boxplot of the orientation error (angle difference of the upper unit normal vector) between the identified structural surfaces and the manually annotated structural surfaces. Specific implementation methods

[0059] The present invention will be further described in detail below with reference to the accompanying drawings. The following specific implementation steps may enable those skilled in the art to more fully understand the present invention, but are not intended to limit the present invention in any form.

[0060] First, it should be noted that the technical solutions of the present invention involve a large amount of prior art, the definitions and concepts of which are already well known or familiar to those skilled in the art. Therefore, unless otherwise specifically explained, the present invention will not further elaborate on other concepts that are consistent with the prior art.

[0061] The present invention provides an intelligent identification method for three-dimensional rock mass structural surface based on AI visual large model, such as Figure 1 As shown, the following steps are included:

[0062] (1) Image selection and synthesis

[0063] Since there is overlap between the captured images, image selection or synthesis is required to simplify 2D inspection. Photogrammetry equipment is divided into mobile equipment and fixed equipment according to the shooting location.

[0064] For mobile devices, such as drones or handheld cameras, a subset of images is selected for inspection. Leveraging the pixel relationships generated during 3D reconstruction, an automated image selection process selects images or regions based on descending order of valid pixel count and non-overlapping 3D points.

[0065] For fixed equipment, such as a camera mounted on a tripod, images are typically captured from more than two positions for 3D reconstruction. An image composite editor developed by Microsoft is used to stitch overlapping photos taken from a single camera position into a seamless, high-resolution panorama, known as a composite image. Thus, a composite image can be constructed from a single position. Similarly, using the pixel relationships generated during 3D reconstruction, an automated image selection process selects a composite image or region thereof for 2D structural surface detection based on the aforementioned principles.

[0066] (2) Interactive structural surface detection

[0067] For 2D structural surface detection, unlike traditional data-driven methods (such as deep learning), which require large amounts of labeled data for dataset construction and model training, pre-trained AI visual large models reduce the need for retraining and lower the threshold for geological surveys. Therefore, this approach, by leveraging the generalization capabilities of AI visual large models, provides a unified solution for structural surface detection in various geological environments.

[0068] AI visual big model can realize real-time interactive use. Figure 2 As shown in Figure 2, the model architecture consists of three core components: an image encoder, a hint encoder, and a mask decoder. The image encoder computes image embeddings, the hint encoder provides hint information, and the lightweight mask decoder combines these two components to predict segmentation results. By decomposing large AI vision models into these components, the same image embeddings can be reused with different hint information, reducing redundant computational overhead. Given the image embedding and hint input, the mask decoder can predict the mask in just 55 milliseconds, achieving real-time and accurate 2D structured surface recognition.

[0069] The prompt module of the AI ​​visual large model can realize interactive structural surface detection. The present invention uses two types of prompts: point selection (foreground points are used to add areas and are marked in green; background points are used to remove areas and are marked in red) and box selection. Figure 3The step-by-step process of interactive structure plane detection using only point cues on a complex structure plane is demonstrated. The goal is to identify the structure plane labeled as No. 6 in the image, which is difficult to be clearly outlined due to its complex boundary, fuzzy contour, partial invisibility, and location at the light-shadow junction. Finally, the structure plane is successfully detected by using 5 foreground points and 5 background points.

[0070] (3) Pixel-point relationship generation

[0071] The motion recovery structure process is shown in Figure 4 . The input is unordered overlapping images, and the output includes three parts: color sparse point cloud model, camera pose, and pixel-point relationship. The motion recovery structure technique generates sparse three-dimensional point clouds by matching pixel features between images and calculating the pose of each image, and preserves the relationship between pixels and points without additional computational cost, as shown in Figure 5a and Figure 5b .

[0072] (4) Three-dimensional structure plane indexing

[0073] Pixel-point relationship is a bridge between image segmentation and point cloud segmentation. Therefore, three-dimensional structure planes can be obtained by automatic indexing. Using the pixel-point geometric relationship between two and three dimensions, the corresponding three-dimensional point cloud is extracted from the two-dimensional image mask. After carefully selecting the parameters to obtain satisfactory motion recovery structure modeling results, the pixel-point relationship is used to convert the two-dimensional structure plane segmentation results into three-dimensional point cloud data. This conversion is achieved by automatic indexing, which can spatially locate the structure plane based on the established two-dimensional-three-dimensional relationship. Since each observed pixel corresponds to a uniquely determined three-dimensional structure point, the conversion process is simple and accurate.

[0074] For mobile devices, a single index connects the image with its point cloud. For fixed devices, a double index is used to connect the image, its composite image, and the corresponding point cloud. The segmentation result is identified on the composite image, so that it can uniquely determine which pixel in the original image corresponds to a given pixel in the composite image. In addition, image observations are indexed to their corresponding three-dimensional points, which are uniquely determined by the pixel-point relationship. Figure 6 The process of obtaining three-dimensional structure planes from fixed position devices using automatic indexing is demonstrated.

[0075] Specific application examples:

[0076] Two examples are involved here.

[0077] 1. Two-dimensional database

[0078] To verify the generalization ability of two-dimensional structure plane detection, the invention introduces a two-dimensional database (see Figure 7), a newly constructed dataset of diverse image distributions for evaluation. The 2D database contains 2,342 identified 2D structural surfaces in 170 images, 90 of which were sourced from the internet and 80 were self-taken, ensuring a broad representation of geological and environmental conditions.

[0079] The masks in the 2D database are manually annotated using the LabelMe tool as ground truth. LabelMe is an open source image annotation tool that is widely used in object detection, semantic segmentation, and panoptic segmentation tasks. Figure 8a and Figure 8b As shown, the 2D database captures a wide range of structural surface textures and geometric forms, as well as varying environmental factors (such as lighting, shadows, and vegetation disturbance), reflecting the complexity of the structures and environments.

[0080] The present invention uses all 170 image samples and 2,342 identified 2D structural surfaces in the 2D database for evaluation. The AI ​​vision large model achieves an average intersection-over-union ratio of 78% when prompted with 11 points, and the performance stabilizes thereafter, even though the number of prompts increases, such as Figure 8a The distribution is shown in the box plot. Figure 8b The comparison results in

[15] demonstrate the zero-shot capability of the AI ​​Vision Big Model in structural surface recognition. This confirms that, despite not being specifically trained on structural surfaces, the AI ​​Vision Big Model still has strong generalization capabilities and can recognize unseen structural surface images.

[0081] like Figure 9 As shown in the figure, the AI ​​Vision Big Model outperforms manual annotation in two aspects: (1) it achieves strong noise suppression by distinguishing structural surfaces from interfering elements such as vegetation and cables; and (2) it accurately outlines the boundaries of structural surfaces with complex geometric shapes. This capability stems from the AI ​​Vision Big Model's inherent target recognition and semantic understanding capabilities, enabling it to achieve robust structural surface recognition through interactive prompts while maintaining computational efficiency. In addition, the AI ​​Vision Big Model can run effectively on input images of arbitrary resolution without the need for preprocessing, further enhancing its practicality in field applications.

[0082] 2. 3D database

[0083] To validate the proposed method's accuracy in 3D structural surface recognition, data was collected using stationary equipment to form a 3D database. The 3D structural surfaces in the point cloud were manually annotated using CloudCompare and used as ground truth. Compared to mobile equipment, stationary equipment introduces greater complexity in the workflow and indexing process. This complexity was the primary reason for selecting a stationary equipment dataset to evaluate the method's recognition accuracy.

[0084] 3D database (such as Figures 10a to 10d ) were obtained from a fixed setup, specifically using a Nikon D300s camera with a resolution of 4256 x 2832 pixels. The images were taken at the Sunbury Rock Bench, located along Highway 15 near Kingston, Ontario, Canada, focusing on the granite outcrop exposed at the roadside. A total of 471 images were taken from three fixed positions, with 146, 168, and 157 images taken from each position, respectively. Using a 3D structural surface recognition method, a sparse point cloud containing 392,229 points was reconstructed, and a pixel-point relationship was established. The point cloud resolution, typically defined as the average distance between neighboring points, was calculated using Equation (1), where M represents the number of point pairs, di represents the distance between the i-th pair of nearest points. i Due to the lack of global positioning system data or known object dimensions, no real-world scale calibration was applied during the reconstruction process. Therefore, the resolution is expressed in local coordinate units, with a calculated result of 0.0013. Although real-world scales cannot be obtained, this does not affect the verification of recognition accuracy. Manual annotation identified 87 structural surfaces of different sizes and attitudes, establishing a benchmark for the method. Unannotated areas were classified as interference due to excessive fragmentation or lack of sufficient three-dimensional information.

[0085]

[0086] The present invention evaluates the accuracy of structural surface segmentation by measuring the intersection over union and precision between the predicted and true point clouds. Intersection over union is a commonly used metric in segmentation tasks, quantifying the degree of overlap between the predicted and true point clouds. Its value ranges from 0 to 1, where 1 indicates perfect alignment and 0 indicates no overlap. On the other hand, precision assesses the reliability of the model's predictions by measuring the proportion of correctly identified points among all predicted points. While intersection over union reflects the overall accuracy of segmentation, precision reflects the method's ability to minimize false positives to the greatest extent.

[0087] To further evaluate segmentation performance, the present invention also calculates the mean intersection over union and mean precision. The mean intersection over union is obtained by averaging the intersection over union scores for all structural surfaces in the dataset, providing a global assessment of segmentation quality. The mean precision is derived from the precision-recall curve, which measures the method's performance at different confidence thresholds, providing a comprehensive assessment of its predictive ability.

[0088] The method successfully identified 87 structural surfaces (see Figure 11a ), achieving a mean precision of 0.91 and a mean intersection over union of 0.70 in 3D structural surface recognition. Figure 11b and Figure 11cHalf boxplots and half violin plots are used to display the distribution of IoU and precision, highlighting extreme values, interquartile ranges, medians, and means. Approximately 70% of identified 3D structural surfaces achieved a precision exceeding 0.90, and approximately 75% achieved an IoU exceeding 0.60. These results demonstrate that our method maintains high accuracy in 3D structural surface identification, providing a practical and computationally efficient tool for large-scale geological annotation.

[0089] Figure 11d The false positive points (marked in black) are shown along with the complete annotated point cloud. The false positive points are mainly clustered near the annotated structural surfaces (marked in yellow), indicating a strong consistency between the segmentation results and the true values. The main sources of these errors include inaccuracies in 2D interactive structural surface detection and the subjectivity of manual 3D annotation. Directly segmenting structural surfaces in point clouds usually requires frequent adjustments to the viewing angle, which makes manual annotation time-consuming and prone to blurred boundaries. Fundamentally, the method operates based on image segmentation methods, while the true value annotations are generated by direct point cloud segmentation. Both methods utilize 3D reconstruction results, but the method combines 3D reconstruction with image segmentation, while the true value segmentation is performed sequentially after 3D reconstruction. This difference essentially explains the similarity between the two segmentation results.

[0090] By analyzing the structural surface with low recognition accuracy (No. 83), the present invention further explores the recognition results of three-dimensional structural surfaces. The structural surface No. 83 has the largest pixel area among all low-precision structural surfaces. Figure 12a The results of 2D interactive structural surface detection are presented.

[0091] Further analysis of the relationship between pixel area and accuracy, Figure 12b and Figure 12c The scatter plots of accuracy and pixel area and the scatter plots of accuracy / pixel area and pixel area are shown respectively. Fitting function y = 0.6332x -0.962 The results show that there is a strong negative correlation between precision / pixel area and pixel area (fitted R 2 The value is 0.9565). This indicates that there is no direct linear relationship between accuracy and pixel area, confirming that the method can maintain stable accuracy in identifying structural surfaces of different sizes. This robustness in structural surface identification highlights the reliability of the method in identifying geological structures at different scales.

[0092] In order to further verify the accuracy of structural surface orientation recognition, the present invention uses CloudCompare software to automatically calculate the orientation of the structural surface. The orientation of each structural surface is converted into an upper unit normal vector. Subsequently, the angular difference between the identified structural surface and the manually marked structural surface, i.e., the orientation error γ, is calculated using formula (2). proposed_frameworkis the upper unit normal vector of the identified structural surface, n proposed_framework is the upper unit normal vector of the manually annotated structural surface.

[0093]

[0094] Figure 13 The frequency distribution and box plot of the occurrence error (unit normal vector difference) are shown. Figure 13 As can be seen, the error range is between 0° and 3.14°, indicating that the deviation between the identified structural surfaces and the manually annotated ones is relatively small, comparable to the results discussed previously. This demonstrates that the proposed method accurately captures the occurrence of the structural surfaces. The root cause of the discrepancy is that the point clouds of the structural surfaces are not completely consistent. This result further validates the accuracy of the proposed method.

[0095] Note: The actual scope of the present invention includes not only the specific embodiments disclosed above, but also all equivalent solutions that implement or execute the present invention under the claims.

Claims

1. A method for intelligent identification of three-dimensional structural surfaces of rock masses based on AI visual large models, characterized in that: Combining AI visual large models with structure-from-motion technology, the system uses data from different photogrammetry devices to achieve 3D structural surface recognition. By using the correlation between image pixels and point clouds in structure-from-motion, the recognition task is converted from point clouds to raw image data. It includes two parallel processes: 2D structural surface detection and 3D reconstruction, which is divided into four parts: (1) Image selection and synthesis When there is overlap between captured images, use image selection or synthesis to simplify 2D detection; (2) Interactive structural surface detection For 2D structural surface detection, a pre-trained AI vision large model is used; (3) Pixel point relationship generation The input is unordered overlapping images, and the output consists of three parts: a colored sparse point cloud model, the camera pose, and the pixel-point relationship. The motion recovery structure technology is used to generate a sparse 3D point cloud by matching pixel features between images and calculating the pose of each image. (4) Three-dimensional structural surface index Acquire 3D structural surfaces through automated indexing; extract the corresponding 3D point cloud from the 2D image mask using the pixel-point geometric relationship between 2D and 3D; After selecting parameters to obtain satisfactory motion-recovered structure modeling results, the pixel-point relationship is used to convert the 2D structural surface segmentation results into 3D point cloud data; This conversion is achieved through automated indexing, which allows the spatial positioning of structural surfaces based on established 2D-3D relationships.

2. The method for intelligently identifying three-dimensional structural surfaces of rock masses based on an AI visual large model according to claim 1, characterized in that: (1) In image selection and synthesis, photogrammetry equipment is divided into mobile equipment and fixed equipment according to the shooting location; For mobile devices, a subset of images is selected for inspection. Using the pixel relationships generated during 3D reconstruction, an automated image selection process selects images or regions thereof based on descending order of the number of valid pixels and non-overlapping 3D points. For fixed equipment, images are taken from two or more positions for 3D reconstruction; overlapping photos taken from a single camera position are stitched together into a seamless high-resolution panorama, which is called a composite image; using the pixel point relationships generated during the 3D reconstruction process, an automated image selection program selects the composite image or its area according to the above principles for 2D structural surface detection.

3. The method for intelligently identifying three-dimensional structural surfaces of rock masses based on an AI visual large model according to claim 1, characterized in that: (2) In interactive structural surface detection, the AI ​​vision large model architecture consists of three core components: image encoder, hint encoder, and mask decoder; the image encoder calculates the image embedding, the hint encoder provides hint information, and the lightweight mask decoder combines these two parts of information to predict the segmentation result; By breaking down the AI ​​vision model into these components, the same image embedding can be reused with different prompt information to achieve real-time and accurate 2D structural surface recognition; The prompt module of the AI ​​vision large model realizes interactive structural surface detection; it uses two types of prompts: points and boxes; points include: foreground points for adding areas and background points for removing areas.

4. The method for intelligently identifying three-dimensional structural surfaces of rock masses based on an AI visual large model according to claim 1, characterized in that: (4) In the 3D structural surface index, for mobile devices, a single index connects the image to its point cloud; for fixed devices, a dual index is used to connect the image, its composite image, and the corresponding point cloud; The segmentation results are identified on the synthesized image, making it possible to uniquely determine which pixel in the original image corresponds to a given pixel in the synthesized image; furthermore, image observations are indexed to their corresponding 3D points, which are uniquely identified via pixel-point relationships.

Citation Information

Patent Citations

  • Rock mass structural surface recognition method, device and equipment and readable storage medium

    CN117456280A

Cited By

  • Deep ground surrounding rock structural plane in-situ image segmentation method based on dual-network model

    CN122066715A