In-situ rock recognition and volume estimation method based on AI visual large model
Interactive block detection and automated segmentation are performed through AI vision big models, combined with convex hull algorithm to calculate volume, the problem of block recognition and volume estimation in complex rock bodies is solved, and efficient and accurate block recognition and volume calculation are achieved.
Patent Information
- Application Number
- CN202510347845.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to quickly and accurately identify and calculate the volume of rock mass from point cloud data, especially under complex structural surface conditions, which are time-consuming and labor-intensive and error-free.
The AI vision big model is used for interactive block detection and automated block segmentation, combined with the convex hull algorithm to calculate the block volume, and the heavy image encoder, flexible prompt encoder and fast mask decoder are used for block recognition. Unknown parameters are recovered through motion recovery structure technology, and block information is directly extracted from point cloud data.
It realizes efficient and accurate block recognition and volume estimation under complex rock mass conditions, avoids the assumption of structural surface characteristics, is suitable for diversified geological scenarios, and improves recognition efficiency and accuracy.
Smart Images

Figure CN120298709A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of rock mass structural plane identification, and particularly to an in-situ rock block identification and volume estimation method based on an AI vision large model. Background Art
[0002] The formation of in-situ rock blocks in rock masses is controlled by the intersection of structural planes, and this intersection results in the formation of rock blocks with diverse three-dimensional geometric shapes. The block size is recognized by the International Society for Rock Mechanics and Rock Engineering as a key parameter characterizing rock mass structural planes and is also the basis of rock mass classification systems (such as the Geological Strength Index). In addition, the characteristics of in-situ blocks on outcrops significantly affect the probability of rock failure. Therefore, the block size also plays a key role in predicting the volume of potential rockfalls and their associated risks. This is because the rocks falling from cliffs may occur in the form of individual blocks, and the volume of these blocks directly affects the intensity of fragmented rockfalls.
[0003] When measuring the block size on outcrops, traditional methods based on field geological survey data usually adopt the spacing equation method, statistical modeling method, and block identification method. However, the methods based on field geological surveys are time-consuming, laborious, and may be dangerous. Currently, remote sensing technologies (such as photogrammetry conducted on the surface or even underground) have been widely applied in engineering geological practices as an alternative to field geological surveys for generating three-dimensional point cloud data of outcrops. The research mainly focuses on how to digitally measure the block size from point cloud data or how to define the surface from data containing noise. However, an efficient algorithm for directly calculating the block size from point clouds is still lacking because point clouds are disordered and not intuitive enough for block identification, making it extremely challenging to identify blocks from point clouds.
[0004] There are two main methods for measuring block sizes - the structure - plane - based method or the block - based method, and their applicability depends on the type of structure - plane pattern on the outcrop, such as trace - type structure planes or planar - type structure planes. For example, methods based on structure - plane surveys, such as grid mapping or scan - line surveys, rely on the structure - plane characteristics obtained from the structure - plane survey to simulate block formation. The methods used include the spacing equation, which requires measuring the spacing between structure planes; the structure - plane intersection method, which divides blocks based on the information of the exposed structure planes; and the statistical modeling method, which involves structure - plane characteristics such as orientation and spacing distribution, persistence, and other properties. Similarly, the block - based survey method directly measures the blocks from relevant dimensions. Therefore, it is obvious that this method is not suitable for structure - plane patterns exposed in the form of traces. However, compared with the structure - plane - based survey methods, the block - based survey method usually has advantages because it is much easier to directly measure the block sizes than to consider all the structure planes, especially in the case of irregular structure planes. However, due to the complexity of interpreting blocks from (point - cloud) digital outcrop models, there is currently no relevant research on directly conducting block surveys from digital models.
[0005] AI large models have achieved unparalleled efficiency in visual perception tasks (such as object classification, detection, and segmentation). This achievement is mainly attributed to the increase in the scale and complexity of AI models, as well as their ability to learn features from massive amounts of data. AI visual large models can generalize to multiple perception tasks without explicit training (i.e., zero - shot generalization), which marks a major breakthrough in the field of computer vision. In addition, when equipped with a prompt - engineering module, AI visual large models can solve a wide range of perception tasks in complex real - world environments and identify objects beyond predefined categories. Therefore, it is possible for AI large models to quickly detect and segment in - situ rock blocks in the field, thus enabling digital measurement of directly calculating the block volume from point - cloud data. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to overcome the deficiencies in the prior art and provide a method for identifying and estimating the volume of in - situ rock blocks based on an AI visual large model.
[0007] To solve the technical problem, the solution of the present invention is:
[0008] (1) Interactive block detection
[0009] The AI visual large model is used to intelligently and interactively detect blocks from only one or more selected original images. The "interactive" here refers to its ability to optimize the block mask through mouse - interaction prompts or filtering, thus helping to achieve accurate results.
[0010] Detecting blocks in complex rock outcrops is a major challenge due to the diverse field conditions and geological features. Since large AI vision models have been pre-trained on large-scale datasets and can respond appropriately to any prompt almost in real-time, the block detection problem can be solved by designing appropriate prompts. Notably, in this detection step, the large AI vision models do not require any additional model training, meaning it is a zero-shot learning. Therefore, the large AI vision models can be applied to a wide range of practical block detection tasks without modification, tasks that go beyond the scope of their original training data.
[0011] The large AI vision models demonstrate powerful visual perception capabilities through three core components in their model architecture, namely, a heavy image encoder, a flexible prompt encoder, and a fast mask decoder. The image encoder generates image embeddings, enabling the real-time creation of block masks through various prompts. The prompt encoder enables interactive prompts by supporting four input types (points, boxes, masks, and text), each with a tailored encoding method for optimal representation. In this invention, only points and boxes are used for interactive block detection. The mask decoder combines the information from the image embedding and the prompt decoder to generate image masks, each with a confidence score.
[0012] When applying the large AI vision models to detect blocks, first, representative images are selected, and the number can vary from one to several, depending on the complexity of the block shape and the extent of the outcrop area. When choosing the viewing angle, it is recommended to select images that can clearly capture the appearance of the blocks and expose multiple surfaces, as this will help the large AI vision models detect more accurately. Additionally, since the three-dimensional block segments identified from a single viewing angle only represent partial views of the blocks, combining images from multiple different viewing angles is crucial for a complete outcrop characterization. For larger outcrop areas, they can be divided into smaller regions, and image regions can be selected from each region for block identification.
[0013] The large AI vision models are able to recognize these possibilities and provide a confidence score for each detected block. This score indicates the model's confidence in the accuracy of the predicted segmentation, reflecting the likelihood that the blocks or objects identified in the image correspond to the real blocks or objects in the real world, thus helping to evaluate the quality of the segmentation and determine whether the results are acceptable or need further refinement.
[0014] In complex rock masses with numerous discontinuities, the formation of in-situ blocks is affected by the intersections of these discontinuities, resulting in multiple potential block configurations. To address the diversity of block configurations, the proposed method adopts two interaction modes to adapt to specific site conditions and provide in-situ block identification results that reflect the above factors. The first interaction mode is called the "prompt mode", which uses the prompt decoder of the model to process various input prompts (such as points, boxes) through mouse interaction. This enables the creation of block masks through interactive sampling to better conform to expert expectations. The second interaction mode is called the "all mode", which generates a large number of masks through automatic grid sampling and manual screening to eliminate misidentified blocks.
[0015] (2) Automated block segmentation
[0016] Using the intermediate file of the structure from motion technique (structure from motion data file), the block fragments obtained in the previous step are used as input, and their corresponding block points are automatically indexed. This limits the applicability of the proposed method, which is applicable to digital outcrop model data based on photogrammetry and requires the preservation of images taken by cameras or drones. The three-dimensional point cloud of the blocks is extracted and stored as a three-dimensional point cloud for direct block identification, making it applicable to complex rock masses with four or more discontinuity sets, or conditions with random and dense discontinuities.
[0017] AI large models and vision large models can automatically recover all unknown parameters, such as camera poses and three-dimensional points, which are derived from overlapping images. Different from traditional aerial triangulation, structure from motion does not require prior knowledge of these parameters, making it highly adaptable to drone-based photogrammetry.
[0018] The incremental structure from motion process is synchronized with block detection in the first step, reducing the total time required for the proposed method. Since all the original images collected by the drone need to be processed in this step, the process is computationally intensive and usually takes several hours to process hundreds of images, which is much longer than the block detection stage. However, the structure from motion data file only needs to be calculated once, and subsequent steps rely on simple indexing of the structure from motion data file, thus simplifying the automated block segmentation process. Image fragments can be imported in batches, followed by an indexing algorithm to efficiently segment blocks in the three-dimensional point cloud. This enables a seamless automated processing process from image masks to block segmentation.
[0019] (3) Block volume measurement
[0020] Then the point cloud is input into an automated convex hull algorithm, which calculates the volume of the block according to the shape and geometric features of the block. The size of the block can be described by two dimensions: the block volume ([L3]) and the block size ([L]). The latter can be defined in various ways, such as the nominal diameter of an equi - volume cube or the maximum length between two vertices of the block. In the present invention, the block volume is used to represent the size of the block because this is a commonly used measure in volume estimation.
[0021] The present invention introduces a convex hull algorithm. This assumption provides a more accurate representation of the in - situ block surface and is more flexible because it is not restricted by the characteristics of structural planes. Based on this, the proposed method directly constructs the convex hull of the block from the point cloud data and calculates its volume, ensuring a one - to - one correspondence with the original outcrop position. In addition, this method eliminates the need to identify structural plane information, thus simplifying the volume estimation process. Compared with the method of assuming convex polyhedra, this method only estimates the exposed part of the block, but avoids assumptions about structural planes, such as the complete penetration of structural planes or the planar geometry of structural planes, making it more suitable for practical scenarios. This process is automated and can be applied efficiently in batches. The convex hull is the smallest convex shape that completely encloses these points and is a fundamental concept in mathematics and computational geometry.
[0022] Principle description of the present invention:
[0023] Extracting block information from an unordered point cloud containing millions of points (or more) is a complex task. However, the in - situ rock block intelligent recognition and volume estimation method based on the AI large - model vision large - model proposed in the present invention establishes an accurate spatial mapping relationship between each point in the digital outcrop model and the pixels of the original photo. Therefore, it transforms the block recognition task from point cloud segmentation into image segmentation, which is a more intuitive and easier - to - implement task.
[0024] Therefore, the present invention proposes a new in - situ rock block intelligent recognition and volume estimation method based on the AI vision large - model. Based on the block geometric data obtained from three - dimensional block recognition,
[0025] Compared with the prior art, the beneficial effects of the present invention are:
[0026] 1) The in - situ rock block intelligent recognition and volume estimation method based on the AI vision large - model achieves efficient block recognition in digital geological surveys by utilizing large - models that perform well in visual perception tasks. This method can effectively handle segmentation tasks without additional training and shows strong performance when dealing with new data, which highlights its flexibility and engineering practicality. These advantages make this method a powerful tool for advancing geological analysis.
[0027] 2) Different from the methods based on discontinuity surveys, which involve various assumptions (such as combining discontinuities into blocks), the proposed method directly targets rock blocks. This method avoids unnecessary assumptions, including the parallelism assumption of discontinuity sets in the spacing equation method, the complete penetration assumption of discontinuities in the intersection line method of discontinuities, and the specific parameter distribution assumptions of discontinuity orientations and centers in statistical modeling methods. This method can bypass these assumptions and represent in-situ blocks more accurately, being particularly suitable for field applications that require reliable block volume estimation.
[0028] 3) The direct identification of blocks enables this method to address the challenges in complex rock mass analysis, and the present invention has successfully solved the problem of quickly identifying complex blocks under random and densely discontinuous conditions. This method is applicable to various rock outcrops, can directly identify blocks without relying on the identification of discontinuities, thereby expanding its application scope in diverse geological scenarios. Description of the Drawings
[0029] Figure 1 Overview of the in-situ rock block identification and volume estimation method based on the large AI vision model; (a) The parallel workflow of the large AI vision model and structure from motion enables 3D block identification using images; (b) The detected block fragments through interactive prompts are highlighted in randomly assigned colors;
[0030] Figure 2 Flowchart of the proposed in-situ rock block identification and volume estimation method based on the large AI vision model;
[0031] Figure 3 Input, architecture, and output mask of the large AI vision model for interactive block detection;
[0032] Figure 4 The workflow of interactive block detection is divided into "prompt mode" and "all mode";
[0033] Figure 5 Step-by-step demonstration of interactive detection of blocks with complex boundaries in prompt mode;
[0034] Figure 6 Cardboard experiment for method verification;
[0035] Figure 7a Select an image as the input;
[0036] Figure 7b In prompt mode, the masks of all blocks in the image are intelligently detected, and the blocks marked with red numbers are fully exposed blocks;
[0037] Figure 7cThe block mask is then used as input and the structural data file index for motion recovery is used to obtain the point cloud corresponding to all blocks.
[0038] Figure 7d The obtained block point cloud is used as input for calculating the convex hull volume.
[0039] Figure 8 For Specific Application Example 2, an overview of the research site; a reservoir slope is covered with a large number of fragmented blocks and boulders.
[0040] Figure 9 For Specific Application Example 2, a schematic diagram of the research area division; the research area is divided into 20 sub-regions, and each sub-region corresponds to an image area.
[0041] Figure 10a For Specific Application Example 2, a histogram of the measured in-situ block volumes in the research area and the best-fit distribution function (the horizontal axis uses a logarithmic scale).
[0042] Figure 10b For Specific Application Example 2, a cumulative distribution diagram of the measured in-situ block volumes in the research area and the best-fit distribution function (the horizontal axis uses a logarithmic scale). Specific implementation method
[0044] The following further details the present invention in conjunction with the accompanying drawings. The following specific implementation steps can enable those skilled in the art to more comprehensively understand the present invention, but do not limit the present invention in any form.
[0045] It should be noted first that the technical solution of the present invention involves a large amount of prior art, and its definitions or concepts are already well-known or mastered by those skilled in the art. Therefore, unless the present invention makes a special explanation for a specific meaning, for those with the same meaning as the existing well-known ones, the present invention will not be described one by one.
[0046] The method for in-situ rock block recognition and volume estimation based on an AI vision large model described in the present invention ( Figure 1 ) includes the following steps ( Figure 2 ):
[0047] (1) Interactive block detection
[0048] The AI vision large model is an out-of-the-box and user-friendly large model for intelligently and interactively detecting blocks from only one or multiple selected original images. The "interactive" mentioned here refers to its ability to optimize the block mask through mouse interaction prompts or filtering, thereby helping to achieve accurate results.
[0049] Detecting blocks in complex rock outcrops is a major challenge due to the diverse field conditions and geological features. Since large AI vision models have been pre-trained on large-scale datasets and can respond appropriately to any prompt almost in real-time, the block detection problem can be solved by designing appropriate prompts. Notably, in this detection step, the large AI vision model does not require any additional model training, which means it is a zero-shot learning. Therefore, the large AI vision model can be applied, without modification, to a wide range of practical block detection tasks that go beyond the scope of its original training data.
[0050] The large AI vision model demonstrates powerful visual perception capabilities through three core components in its model architecture, namely a heavy image encoder, a flexible prompt encoder, and a fast mask decoder, as Figure 3 shown. The image encoder generates image embeddings, enabling the real-time creation of block masks through various prompts. The prompt encoder enables interactive prompts by supporting four input types (points, boxes, masks, and text), each with a tailored encoding method for optimal representation. In this invention, only points and boxes are used for interactive block detection. The mask decoder combines the information from the image embedding and the prompt decoder to generate image masks, each with a confidence score.
[0051] When applying the large AI vision model to detect blocks, the first step is to select representative images, the number of which can vary from one to several, depending on the complexity of the block shape and the extent of the outcrop area. When choosing the viewing angle, it is recommended to select images that can clearly capture the appearance of the block and expose multiple surfaces, as this will help the large AI vision model to detect more accurately. Additionally, since the three-dimensional block segments identified from a single viewing angle only represent partial views of the block, combining images from multiple different viewing angles is crucial for a complete outcrop characterization. For larger outcrop areas, they can be divided into smaller regions, and image regions can be selected from each region for block identification.
[0052] The large AI vision model is able to recognize these possibilities and provides a confidence score for each detected block, which indicates the model's confidence in the accuracy of the predicted segmentation, reflecting the likelihood that the block or object identified in the image corresponds to the real block or object in the real world, thus helping to evaluate the quality of the segmentation and determine whether the result is acceptable or requires further refinement.
[0053] In complex rock masses with numerous structural planes, the formation of in-situ blocks is influenced by the intersections of these structural planes, resulting in multiple potential block configurations. To address the diversity of block configurations, the proposed method adopts two interaction modes to adapt to specific site conditions and provide in-situ block recognition results that reflect the above factors. The first interaction mode is called the "prompt mode", which uses the prompt decoder of the model to process various input prompts (such as points, boxes) through mouse interaction. The second interaction mode is called the "all mode", which generates a large number of masks through automatic grid sampling and manual screening to eliminate misidentified blocks. This enables the creation of block masks through interactive sampling to better conform to expert expectations. Figure 5 Shows the step-by-step process of interactive block detection in the prompt mode. The block is recognized through four simple mouse clicks, and its boundary is defined by the surrounding structural planes, forming a complex boundary. Compared with manual polygon annotation, this interactive method provides a more efficient and intelligent way to identify blocks. The second interaction mode is called the "all mode", which generates a large number of masks through automatic grid sampling and manual screening to eliminate misidentified blocks (such as Figure 4 shown).
[0054] (2) Automated block segmentation
[0055] Using the intermediate file of the structure from motion technology (structure from motion data file), the block fragments obtained in the previous step are used as input, and their corresponding block points are automatically indexed. This limits the applicability of the proposed method, which is applicable to digital outcrop model data based on photogrammetry and requires the retention of images taken by cameras or drones. The three-dimensional point cloud of the block is extracted and stored as a three-dimensional point cloud for direct block identification, making it applicable to complex rock masses with four or more structural plane groups, or random and structurally dense conditions.
[0056] The AI vision large model can automatically recover all unknown parameters, such as camera pose and three-dimensional points, which are derived from overlapping images. Different from traditional aerial triangulation, structure from motion does not require prior knowledge of these parameters, making it highly adaptable to drone-based photogrammetry.
[0057] The incremental structure from motion process is synchronized with the block detection in the first step, reducing the total time required for the proposed method. Since all the original images collected by the drone need to be processed in this step, the process is computationally intensive and usually takes several hours to process hundreds of images, which is much longer than the block detection stage. However, the structure from motion data file only needs to be calculated once, and subsequent steps rely on simple indexing of the structure from motion data file, thus simplifying the automated block segmentation process. Image fragments can be imported in batches, followed by an indexing algorithm to efficiently segment blocks in the three-dimensional point cloud. This enables a seamless automated processing process from image masks to block segmentation.
[0058] (3) Block volume measurement
[0059] Then, the point cloud is input into an automated convex hull algorithm, which calculates the volume of the block based on the shape and geometric features of the block. The size of the block can be described by two dimensions: block volume ([L3]) and block size ([L]). The latter can be defined in various ways, such as the nominal diameter of an isovolume cube or the maximum length between two vertices of the block. In the present invention, the block volume is used to represent the size of the block because this is a commonly used metric in volume estimation. The present invention introduces an innovative volume estimation method based on the convex hull. This assumption provides a more accurate representation of the in-situ block surface and is more flexible because it is not restricted by the characteristics of the structural planes. Based on this, the proposed method directly constructs the convex hull of the block from the point cloud data and calculates its volume, ensuring a one-to-one correspondence with the original outcrop position. In addition, this method eliminates the need to identify structural plane information, thus simplifying the volume estimation process. Compared with the method of assuming a convex polyhedron, this method only estimates the exposed part of the block but avoids assumptions about the structural planes, such as the full penetration of the structural planes or the planar geometry of the structural planes, making it more applicable to practical scenarios.
[0060] This process is automated and can be applied efficiently in batches. The convex hull is the smallest convex shape that completely encloses these points and is a fundamental concept in mathematics and computational geometry.
[0061] Example 1:
[0062] To verify the in-situ block intelligent recognition method introduced in this paper, the present invention constructs an outcrop model and uses a cardboard box for model verification. In the present invention, a rectangular cardboard box is used to simulate three sets of orthogonally intersecting structural planes. In the model of the present invention, some structural planes are designed as opening features (i.e., structural planes with openings), and cardboard boxes of different sizes are used to represent different scales of blocks commonly found in natural rock masses. To distinguish cardboard boxes of different volumes, the present invention uses color coding.
[0063] In addition, the model shows rectangular blocks with different degrees of exposure to simulate two scenarios commonly found in real rock masses: the blocks are either completely exposed as boulders or partially buried in the rock mass. Since the actual volume of the cardboard box is known, the volume of the identified block can be calculated and compared with the known value to verify the method. Theoretically, the volume of the completely exposed block calculated by the method of the present invention should be very close to the true value, while the volume of the partially exposed block will be underestimated. (This difference occurs because the present invention identifies and estimates the block based on the exposed part of the rock mass outcrop.) In addition, the images used to construct the model are publicly provided as a benchmark for verifying algorithms related to structural plane extraction and other calculations in rock engineering.
[0064] This case used 188 UAV images taken by close-range photogrammetry to construct a digital outcrop model around the model, which was represented by 1,024,493 sparse point cloud points. In Figure 6 , the sparse point cloud of the outcrop model generated by 3D reconstruction software.
[0065] Figures 7a through 7d shows the results of the experiment, and a total of 20 blocks in the image were identified. Figure 7b shows the block detection results detected by the AI vision large model from the selected images. The colors represent three categories: black represents blocks with one visible face, green represents blocks with two visible faces, and blue represents blocks with three visible faces. Figure 7c shows the results of automated block segmentation, where the 3D block segments are represented in the form of point clouds. Then, the block mask is used as input to index through the structure from motion data file to obtain the corresponding point clouds of all blocks, thus completing the block recognition method based on the AI vision large model. To make the visualization clearer, the same view as the selected image is used, and the point cloud is superimposed on the image background.
[0066] Next, Figure 7d shows the convex hull generated from the 3D block segments, which is used to calculate the block volume. The viewing angle is slightly adjusted to better show the edges of the polyhedron. By combining the displayed sparse point cloud, it can be clearly seen that the convex hull fits tightly to the block point cloud. The results are roughly as expected: for fully exposed blocks, the calculated volume is close to the true value. The calculated volumes of the fully exposed paperboard blocks and their comparison with the true values are listed in Table 1.
[0067] Table 1 Comparison of measured values and the proposed method for the volumes of fully exposed blocks
[0068]
[0069] Example 2:
[0070] Taking a key slope at the inlet and outlet of the downstream reservoir of a pumped-storage power station in Qinghai Province, China as an example, the benefits of the present invention are further elaborated. The reservoir slope of this hydropower station needs to comprehensively evaluate the risk of falling rocks, and these risks mainly come from two aspects: in addition to partially exposed blocks, there are also a large number of blocks that may be displaced. However, existing methods have difficulties in extracting blocks without clear structural planes, which highlights a key advantage of the method used in the present invention. Figure 8 shows the point cloud of the study area, presenting the current situation of the reservoir slope from two perspectives, and marking the unfavorable geological bodies identified in the engineering geological survey in red.
[0071] This case focuses on the steep slope area of the downstream reservoir, where the fractured rock mass poses a considerable challenge to stability. Since the Middle Pleistocene, these slopes in volcanic clastic rocks, sandstones, slates, and granodiorites, with a slope height ranging from 200 meters to 220 meters and a slope angle between 35° and 40°, have undergone long-term weathering, unloading, and gravitational modification, resulting in frequent failures. On the slope behind the intake structure, dangerous rock masses such as loose rock blocks, isolated boulders, and groups of falling rocks are widely distributed. According to geological field surveys, the sizes of these dangerous rock masses vary, with volumes ranging from 80 cubic meters to 5800 cubic meters, and the total estimated volume is approximately 26000 cubic meters. The presence of these unstable rock masses, especially in areas affected by unloading and erosion, poses a significant risk of falling rocks, making it crucial to identify potential unstable blocks to ensure the safety of the hydropower station reservoir slope. The proposed method can be applied to similar adverse geological sites, especially by directly identifying dangerous rock masses.
[0072] The digital outcrop model data in this case is a sparse point cloud containing 1,007,088 points, covering a slope with an orthographic projection area of 11,222 square meters. These data were generated from 1,160 drone-taken images, although only 20 representative images were needed for the analysis to cover the study area. In Figure 9 , the red dashed line outlines the scope of the study area, and the red solid line divides it into 20 sub-areas, with each image corresponding to one sub-area. To avoid repeated identification of blocks, each image is marked with a non-overlapping area, and the image areas numbered 1 and 2 show the correspondence between the image and the area.
[0073] During the interactive block detection process, an automatic mask generator was used to implement an AI vision large model on the selected 20 images in "full mode", with the parameter settings of 32 points per side and 32 points per batch. Subsequently, results unrelated to the blocks, such as the sky background and grassland, were manually removed to obtain a large number of block masks for each image. After that, some irrelevant or inaccurate masks also needed to be manually screened out. In the subsequent fully automated block segmentation and block volume measurement process, the point cloud of each block was identified, and the block volume was calculated. A total of 7,303 blocks were identified, with a total volume of 3,462.7739 cubic meters, and 25 blocks exceeding 1 cubic meter contributed a volume of 3,324.5614 cubic meters. The results are as shown in Figure 10a and Figure 10b shown.
[0074] The results show that the largest block is approximately 1,327 cubic meters, which is smaller than the 5,800 cubic meters estimated by the geological field survey. This difference is mainly due to the limitations of the field survey, which is restricted by the inaccessibility of the slope and relies on images and scales to estimate the block volume. The box-fitting method is used to approximate the size of the block, which tends to overestimate the block volume because it assumes a too-simple geometric shape and does not consider the irregularity of the block geometry. In contrast, the improved block measurement method proposed in the present invention utilizes digital measurement techniques and provides a more realistic estimate of the block volume. In addition, the comparison of the total volume shows that the proposed block measurement result is approximately 3,463 cubic meters, which is significantly smaller than the 26,000 cubic meters reported in the field survey report. This is because digital measurement focuses on a segment of the slope for detailed analysis, while the field survey covers the entire geological outcrop, resulting in a larger estimated volume.
[0075] Statistically, the average volume of the blocks is 0.5048 cubic meters, while the median is 0.0012 cubic meters. In geotechnical engineering, the selection of representative or equivalent block sizes is not consistently recorded and applied in different studies. In this case, either the average block volume or the median block volume can be used as an estimate of the representative block size.
[0076] To evaluate the degree of jointing of the rock mass, relying solely on a single value (such as the representative block volume) is insufficient. Therefore, as Figure 10a and Figure 10b shown, the histogram and cumulative distribution diagram of the measured in-situ block volume can provide a more comprehensive evaluation. By applying the method proposed in the present invention, the variation of block sizes within a given location can be effectively represented by a block distribution diagram. This method can efficiently measure the block sizes one by one by directly observing individual blocks on a large area of rock mass, which was previously considered an impossible task. In addition, combined with the logarithmic-scale histogram in Figure 10a , it can also be observed that the block size distribution has a relatively narrow grading and is approximately symmetrically distributed around 1 dm 3 .
[0077] Analysis of several documented rockfall events indicates that the volume of rockfall fragments can be described by a power-law distribution. However, due to the difficulty of measuring in-situ block sizes by observation on the rock mass surface, discussions on the in-situ block size distribution are rare. This task is particularly challenging for extensive rock mass exposures that require individual block size measurements. In this case study, block sizes were measured by direct observation of a large area of the rock mass, and the data were fitted to a log-normal distribution and a Weibull distribution. The shape parameter of the Weibull distribution determines the tail behavior of the distribution. When the shape parameter equals 1, the Weibull distribution degenerates into an exponential distribution, showing a light-tailed characteristic. For the case where the shape parameter is less than 1, the tail becomes lighter; while for the case where the shape parameter is greater than 1, the tail becomes heavier, approaching the characteristics of a power-law distribution. In this case, the shape parameter is 0.3380, and the Weibull distribution shows a power-law-like tail behavior with a mean square error of 18.66. This analysis indicates that the block size distribution in the studied rock mass exposure area is approximately a power-law distribution. From the mean square error results, the block distribution in this area can be well fitted by a log-normal distribution (see Figure 10b ), with a mean square error of only 0.40. These in-situ block data extracted from the reservoir slope provide valuable original geological information for geological surveys.
[0078] Note: The actual scope of the present invention includes not only the specific embodiments disclosed above, but also all equivalent solutions that implement or execute the present invention under the claims.
Claims
1. An in-situ rock block recognition and volume estimation method based on an AI vision large model, characterized in that, It includes the following steps: (1) Interactive block detection An artificial intelligence vision large model is used to intelligently and interactively detect blocks from only one or more selected original images; the term "interactive" refers to its ability to optimize the block mask through mouse interaction prompts or filtering, thus helping to achieve accurate results; (2) Automatic block segmentation Using the intermediate file of the structure from motion technology, i.e., the structure from motion data file, the block fragments obtained in the previous step are used as input, and their corresponding block points are automatically indexed; It is applicable to digital outcrop model data based on photogrammetry and requires retaining the images taken by cameras or drones; the three-dimensional point cloud of the blocks is extracted and stored as a three-dimensional point cloud for directly identifying the blocks, making it applicable to complex rock masses with four or more structural plane groups, or conditions with random and dense structural planes; (3) Block volume measurement Then the point cloud is input into an automatic convex hull algorithm, which calculates its volume according to the shape and geometric features of the blocks; The size of the blocks is described by two dimensions: block volume ([L3]) and block size ([L]); the block volume is used to represent the size of the blocks.
2. The in-situ rock block recognition and volume estimation method based on the AI vision large model according to claim 1, wherein (1) In the interactive block detection, the model architecture of the artificial intelligence vision large model includes three core components, a mask decoder, an image encoder, and a prompt encoder; The image encoder generates image embeddings, enabling the real-time creation of block masks through various prompts; The mask decoder combines the information of the image embeddings and the prompt decoder to generate image masks, each with a confidence score; the prompt encoder implements interactive prompts, and each input type has a tailored encoding method to achieve the best representation; when applying the artificial intelligence vision large model to detect blocks, representative images are selected. When selecting the viewing angle, images that can clearly capture the appearance of the blocks and expose multiple surfaces are selected. For large outcrop areas, they are divided into small areas, and image regions are selected from each area for block recognition; The artificial intelligence vision large model can identify these possibilities and provide a confidence score for each detected block, which indicates the model's confidence in the accuracy of the predicted segmentation, reflecting the possibility that the blocks or objects identified in the image correspond to the real blocks or objects in the real world, thus helping to evaluate the quality of the segmentation and determine whether the result is acceptable or needs further refinement; To address the diversity of block configurations, two interaction modes are adopted. The first interaction mode is the prompt mode, which uses the model's prompt decoder to process various input prompts through mouse interaction; the second interaction mode is the all mode, which generates a large number of masks through automatic grid sampling and manual screening to eliminate misidentified blocks.
3. The in-situ rock block identification and volume estimation method based on the AI vision large model according to claim 1, characterized in that (2) In the automatic block segmentation, the image fragments are imported in batches, followed by an indexing algorithm to segment the blocks in the three-dimensional point cloud.
4. The in-situ rock block recognition and volume estimation method based on the AI vision large model according to claim 1, characterized in that (3) In the block volume measurement, the convex hull algorithm is to directly construct the convex hull of the blocks from the point cloud data and calculate its volume, ensuring a one-to-one correspondence with the original outcrop position.
Citation Information
Patent Citations
Three-dimensional point cloud semantic segmentation method and device, electronic equipment and storage medium
CN118898717A
Underwater in-situ sea cucumber volume measurement method based on binocular vision and electronic equipment
CN119295532A