A digital modeling method of river sediments using ISfM-PMVS technology based on smartphone-collected images
By combining the ISfM-PMVS method and convolutional neural networks with smartphones, the cost and operational complexity issues of constructing three-dimensional river sediment models in the field using traditional equipment were resolved, achieving low-cost, efficient three-dimensional modeling and accurate geological information interpretation.
Patent Information
- Application Number
- CN202411602548.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-11-11
AI Technical Summary
Existing technologies make it difficult to quickly and cost-effectively construct high-precision three-dimensional digital models of river sediments, especially in field environments. Traditional equipment is limited by cost and operational complexity and cannot fully cover the details of sediment characteristics.
Using a smartphone combined with the ISfM-PMVS method, high-density three-dimensional point clouds and digital orthophotos are generated through multi-view image acquisition, feature point matching, RANSAC and bundle adjustment optimization, and convolutional neural networks are used to interpret geological information.
It achieves low-cost and efficient three-dimensional modeling, generates digital models with centimeter-level accuracy, improves the flexibility and convenience of field data collection, reduces the limitations of traditional equipment, and improves the efficiency and accuracy of geological information interpretation.
Smart Images

Figure CN119600192B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of river sediment digitization, and in particular to an ISfM-PMVS river sediment digital modeling method based on images collected by a smartphone. Background Art
[0002] With the development of science and technology, the continuous progress of modern surveying and mapping technologies for earth observation and spatial information recording, such as three-dimensional laser scanners (Three Dimensional Laser Scanners) and unmanned aerial vehicles (UAVs), can record and track real-life three-dimensional scenes faster, easier and in more detail, effectively enriching the geological database.
[0003] Compared to two-dimensional images and textual information, three-dimensional digital models offer richer data diversity. Numerous researchers are working in geoscience fields, such as quantitative interpretation of geological outcrops, precise measurement of glacial geomorphological changes, and neotectonic activity and its hazards. These studies typically utilize highly overlapping images acquired by unmanned aerial vehicles (UAVs) and then constructed digital models using structure-from-motion (SfM) algorithms. With the rapid advancements in computer vision, researchers are increasingly experimenting with using high-definition smartphone-captured photos to create 3D models for easier storage, sharing, and interpretation of geological outcrops and geological relics. These models capture field information, enabling researchers to digitally interpret and analyze them. For example, Stefano et al. compared the accuracy of smartphone-generated models with 3D models generated by UAVs of a natural rock cliff in Pietrasecca, central Italy. The results demonstrated that smartphones can generate digital models that meet certain accuracy requirements. Chen et al. used image matching algorithms to construct 3D point cloud models of rock outcrops in Nanwang Mountain and Guilin to measure the orientation of discontinuities within the rock outcrops. An et al. used smartphones to digitally model muddy limestone or muddy siltstone in a lithologic landslide zone, measuring rock morphological characteristics through digitized images. Tavani et al. used smartphone images to digitally model the "Acquedotto Felice" section in Rome, Italy, and three geological outcrops in the Aman Mountains. Accuracy verification showed that high-definition images captured by smartphones could effectively interpret geological information. Tavani et al. used the inertial measurement unit (IMU) and magnetometer included in some smartphones to collect directional data, verifying that the error was less than 3.6°, which was essentially consistent with traditional compass measurements, demonstrating the significant potential of smartphones in assisting geological field work. To ensure the effectiveness of smartphone digital models of actual slag heaps in mining environments, Tungol et al. first used smartphones to model large, angular rocks and smaller, rounded rocks. They then quantitatively evaluated the error in three-dimensional reconstruction using SfM. The results showed that the accuracy of smartphone digital models was guaranteed, and that the more images in the same dataset, the smaller the scale error. Corradetti et al. conducted a digital modeling study on a trench approximately 20 meters long in Enemonzo, Italy. The results showed that although the digital model constructed by a smartphone had a 13° rotation error, it could meet the accuracy requirements of three-dimensional digital modeling to a certain extent.
[0004] In summary, many scholars at home and abroad have been conducting research on constructing digital models from smartphone photos using the SfM method, but many scholars only interpret digital models of a single dimension, such as point clouds or DOM. This paper selects three geological sites to quickly obtain photos of geological relics through smartphones, and uses intelligent algorithms to interpret the unique geological information of different geological sites based on digital models of different dimensions. Summary of the Invention
[0005] The present invention aims to provide an ISfM-PMVS river sediment digital modeling method based on smartphone-collected images. The ISfM method, combined with a smartphone, can acquire high-resolution, multi-view images in a short period of time, making it suitable for rapidly constructing three-dimensional digital models. Combining ISfM and PMVS technologies, the method can generate high-density three-dimensional point clouds and digital orthophotos (DOMs), ensuring that the data contain not only geometric information but also rich color and texture information. RANSAC and bundle adjustment (BA) are used to optimize the three-dimensional model parameters, ensuring the accuracy of the sparse point cloud.
[0006] In order to achieve the above technical objectives and the above technical effects, the present invention is implemented through the following technical solutions:
[0007] A method for digital modeling of river sediments using ISfM-PMVS (Infrastructure for Mobile Phase Visualization) based on smartphone-collected images includes the following steps:
[0008] S1: Use a smartphone to capture images from multiple perspectives at a predetermined location;
[0009] S2: Extract feature points from each image and establish the geometric relationship between images by matching these feature points.
[0010] S3: The ISfM method is used to estimate camera parameters and sparse point cloud coordinates, and the initial 3D model is constructed through RANSAC and bundle adjustment optimization.
[0011] S4: Apply the PMVS algorithm and use multi-view consistency verification to generate a high-density 3D point cloud to ensure the accuracy and integrity of the model; transfer the color and detail information of the photo to the 3D model through texture mapping to form a digital orthophoto (DOM).
[0012] S5: Using the generated 3D model and 2D DOM, the geometric morphology of river sediments is identified and analyzed through convolutional neural networks.
[0013] S6: Verify the model accuracy by comparing the digital model with actual measurement data, and use Deep Learning methods to detect the characteristics and distribution of sediments.
[0014] Furthermore, the step S1 specifically includes:
[0015] Image acquisition was performed using a smartphone, such as the HUAWEI IP 30 Pro. Equipped with a 6mm focal length lens, this phone is suitable for high-resolution close-up photography. This requires careful planning of the shooting strategy to ensure sufficient image overlap, which is crucial for subsequent 3D reconstruction. The researchers determined the viewing angle, distance, and direction of movement on site. Specifically, the smartphone was used to capture images along a semicircular path within the sediment area, at a distance of 3 meters. Starting from one side of the sediment area, the camera moved along a predetermined semicircular path, capturing images at regular intervals to ensure comprehensive coverage and capture information from different sides of the sediment. The smartphone was kept stable during this movement, ensuring 60% to 80% overlap between consecutive images for later feature matching and 3D reconstruction. Shooting was also performed under uniform lighting conditions to avoid strong shadows and light contrast that could affect image quality. A total of 81 images with a resolution of 2736×3648 were acquired at the geological site. The content and location of each image were carefully designed to ensure complete coverage and optimal image overlap, ensuring accuracy and integrity in subsequent processing.
[0016] Furthermore, the step S2 specifically includes: feature extraction and feature point matching of adjacent photos can determine the relative position and posture of the camera at different viewing angles, providing an accurate spatial coordinate basis for 3D reconstruction. At the same time, the feature points with the same name between different images can also be determined, thereby establishing geometric constraints between the images. Figure 2 Show, let m K-1 With m K is a pair of matching points with the same name in two adjacent images, and their coordinates are (X, Y, 1) and (X', Y', 1) respectively. The corresponding point m satisfies the geometric constraints:
[0017]
[0018] Where F is a 3×3 matrix with 7 degrees of freedom.
[0019] Assume that the camera parameters are C={C1,C2,C3·····C n}, the point cloud coordinates are X={X1,X2,X3·····X n}, project the point cloud coordinates into the image plane coordinate system through perspective projection transformation, and the projection error is assumed to be g(C,X).
[0020] Furthermore, step S3 specifically includes: by extracting and matching feature points in the image, the ISfM algorithm searches for points that meet the camera and scene parameter optimization goals from the image to solve the image matching sparse point cloud. ISfM mainly estimates the parameter information of the camera C and the three-dimensional coordinates of the point cloud X by nonlinearly minimizing g(C,X):
[0021]
[0022] Where: n is the number of images taken by the camera, m is the number of feature matching points. j On image i, then ω i,j =1, otherwise ω i,j =0,||q i,j -p(C i ,X j )|| 2 is the projection error of the point in the camera.
[0023] The Random Sample Consensus (RANSAC) algorithm and Bundle Adjustment (BA) are used to project the point clouds onto their respective images and find the corresponding points and pixel coordinates. After the BA iteration, a sparse point cloud of the whole scene is obtained.
[0024] Furthermore, the step S4 specifically includes:
[0025] PMVS uses camera pose information from sparse point clouds to generate a high-density 3D point cloud through multi-view consistency verification. Based on the feature points in the sparse point cloud, patches are generated. Multi-view verification methods such as photometric and parallax consistency are used. Photometric consistency verifies the color consistency of image patches across different viewpoints, while parallax consistency controls the consistency of geometric positions to ensure the matching accuracy of these image patches across different viewpoints. Normal vector optimization is also applied to update the normal information of each image patch, improving the accuracy of depth estimation, especially in areas with complex textures or sudden morphological changes. During the final dense point cloud construction process, PMVS integrates image patches that meet consistency requirements into 3D points. The coordinates f of each point are obtained by weighted averaging the multi-view projection results using Equation 6. Texture information is then mapped to the point cloud, completing the digital construction of the 3D model, image dense matching (DIM) point cloud, and digital orthophoto map (DOM).
[0026]
[0027] Where: N represents the number of viewpoints; Ri and ti represent the rotation matrix and translation vector of the i-th viewpoint respectively, pi is the image projection point in the viewpoint, and K is the camera intrinsic parameter matrix.
[0028] Furthermore, step S5 specifically includes: the constructed 3D model and the generated 2D DOM are used to interpret the geometric and morphological information of the river sediments. With the help of the Darknet-53 convolutional neural network, the model can accurately decompose and record the geometric features of geological points in the digital orthophoto.
[0029] Furthermore, step S6 specifically includes comparing and analyzing the digital model's measurement results with actual field measurement data to ensure the model's geometric accuracy is within acceptable limits. Sediment locations are identified using a Darknet-53 network detection box, and the sediment perimeter is revised based on the detection results. This process also requires statistical analysis of the distribution of sediments of different sizes to provide data support for river sediment analysis and geological interpretation.
[0030] Beneficial effects of the present invention:
[0031] The present invention realizes low-cost and high-efficiency three-dimensional modeling by using smartphones for image acquisition. Smartphones are common and easily accessible devices with high-definition cameras that can meet the needs of collecting high-resolution images. Combined with the incremental structure from motion (ISfM) algorithm, feature points can be gradually extracted and matched from the image to establish a sparse point cloud and estimate camera parameters, thereby generating an initial three-dimensional model. Compared with traditional professional equipment, such as three-dimensional laser scanners or drones, the equipment cost is high, the operating technology requirements are high, and the use and maintenance costs are high. In the field environment, the limitations of transportation and power support affect the breadth and sustainability of digitization, and in areas with complex or narrow terrain, the layout and operation of professional equipment are limited by environmental conditions, and the characteristic details of the sediments cannot be fully covered, resulting in partial loss of the digital model. This method does not require additional hardware investment, which reduces the cost of the project. At the same time, because smartphones are portable and easy to use, the flexibility and convenience of field data acquisition are greatly improved.
[0032] In terms of optimizing modeling accuracy, the present invention uses the random sampling consensus (RANSAC) algorithm and bundle adjustment (BA) for data optimization. RANSAC iteratively selects samples from the data set for model fitting and removes outliers to improve the accuracy of point cloud matching. Bundle adjustment further optimizes the camera parameters and the three-dimensional coordinates of the sparse point cloud, improves the accuracy of the model by minimizing the projection error, and forms a basic geometric structure. Next, a multi-view stereo (PMVS) algorithm is used to generate a high-density point cloud through multi-view consistency verification, ensuring the complete presentation of details and the geometric accuracy of the model reaches the centimeter level. This accuracy can provide reliable basic data support in the morphological analysis of river sediments.
[0033] The present invention applies convolutional neural networks (CNN), such as Darknet-53, to the automatic interpretation of digital models. By generating digital orthophotos (DOM), the image texture information is accurately mapped onto the three-dimensional model, providing detailed data input for CNN. The design of multi-layer convolution and residual connection in the Darknet-53 network structure enhances the depth and breadth of feature extraction, enabling it to exhibit good generalization ability when detecting river sediment characteristics, and to detect targets through feature maps of different scales, thereby realizing automatic recognition and classification of sediment geometry. Experiments show that its recognition accuracy rate reaches 0.82, which not only improves the efficiency of geological information interpretation, but also reduces the subjective errors of traditional manual interpretation.
[0034] This technical solution provides a method for quickly acquiring and analyzing geological information, which is of great significance for the recording of geological relics and the protection of natural heritage. The portability of smartphones allows for efficient data collection in complex terrain and harsh environments, and the generated data model can accurately reflect the actual conditions of geological points. The relative accuracy of the model is ensured through measurement and model comparison verification. Even in the absence of geographic coordinates, the geometric error of the model is controlled at the centimeter level. For heritage protection, it can not only assist management and monitoring, but also implement effective protection measures by dynamically tracking geological changes. This method can also be integrated with a variety of modern surveying and mapping technologies to form a multi-scale and multi-information level monitoring system, providing solid data support for the continuous research of geological and ecological environments.
[0035] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0037] Figure 1 It is a schematic diagram of image incremental calculation expression;
[0038] Figure 2 Schematic diagram for photo feature point matching;
[0039] Figure 3 Schematic diagram of the process of establishing a digital model for SfM;
[0040] Figure 4 This is a schematic diagram of the river sediment detection network structure;
[0041] Figure 5 Digital models and physical measurement diagrams of geological points;
[0042] Figure 6 Schematic diagram of river sediment profile;
[0043] Figure 7 Schematic diagram of training and validation loss;
[0044] Figure 8 Schematic diagram of river sediments; DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0046] Example 1
[0047] The ISfM-PMVS river sediment digital modeling method described in this embodiment based on smartphone-collected images includes the following steps:
[0048] S1: Use a smartphone to capture images from multiple perspectives at a predetermined location;
[0049] S2: Extract feature points from each image and establish the geometric relationship between images by matching these feature points.
[0050] S3: The SfM method is used to estimate camera parameters and sparse point cloud coordinates, and the initial 3D model is constructed through RANSAC and bundle adjustment optimization.
[0051] S4: Apply the PMVS algorithm and use multi-view consistency verification to generate a high-density 3D point cloud to ensure the accuracy and integrity of the model; transfer the color and detail information of the photo to the 3D model through texture mapping to form a digital orthophoto (DOM).
[0052] S5: Using the generated 3D model and 2D DOM, the geometric morphology of river sediments is identified and analyzed through convolutional neural networks.
[0053] S6: Verify the model accuracy by comparing the digital model with actual measurement data, and use Deep Learning methods to detect the characteristics and distribution of sediments.
[0054] Example 2
[0055] This example studies the Zhongxing Town-Tiger Leaping Gorge section of the Jinsha River in Yulong Naxi Autonomous County, Lijiang City, using smartphones to digitally model and interpret river sediments. The study area lies within the Zhongxing Town-Tiger Leaping Gorge section of the Jinsha River, with the left bank belonging to Shangri-La City, Diqing Tibetan Autonomous Prefecture, and the right bank to Yulong Naxi Autonomous County, Lijiang City. Since the Quaternary Period, this region has been subject to significant fluvial activity, influenced by factors such as neotectonic movements and climate change, resulting in the formation of V-shaped river valleys, gorges, diversions, side and center shoals, and river terraces. Based on the valley geomorphology and the origin of the sedimentary deposits within the Zhongxing Town-Tiger Leaping Gorge section, this section can be divided into the Zhongxing Town-Longpan Township section and the Tiger Leaping Gorge section. The Zhongxing Town-Longpan Township section has a wide valley with a gentle slope. The longitudinal gradient is approximately 0.67%, with well-developed side and center shoals and multiple terraces along both banks. The impact materials in the valley have a distinct binary structure. The well-rounded alluvial gravels, commonly known as pebbles, are formed by weathering and erosion of various rocks in the Jinsha River and its tributaries. They are carried long distances by flowing water, where they collide, rub, and become rounded, eventually depositing on the riverbed. The gravels are rich in color and complex in composition.
[0056] Smartphone model building
[0057] Smartphones are common electronic devices today, providing data to visual systems at a low cost, effectively recording and protecting cultural heritage. To preserve the scene without disrupting the site, while also enabling accurate measurement and analysis of different geological relics, this study determined the theoretical basis of close-range photogrammetry techniques and attempted to utilize mobile phone imaging. Based on the actual conditions observed on site, feasible and practical imaging perspectives, shooting distances, and mobile photography directions were designed. Geological relic data collection using mobile phone imaging with a certain degree of overlap was carried out and completed. Using a HUAWEI P30 Pro with a 6mm lens focal length, a total of 81 photos were captured from multiple perspectives at three geological sites, with a resolution of 2736*3648 pixels.
[0058] Incremental SfM iteratively fuses one or more images into the initial image pair model. In this process, the relevant algorithm is repeatedly used to calculate the three-dimensional information of the object and update the existing model until the image sequence is completed. Figure 1 Show.
[0059] To build a digital model based on ISfM and smartphone images, it is necessary to first extract the feature points of the photos and match them. This step is the basis for building a digital model and is used to determine the feature points with the same name between different images, thereby establishing geometric constraints between the images. Figure 2 Show, let m K-1 With m K is a pair of matching points with the same name in two adjacent images, and their coordinates are (X, Y, 1) and (X', Y', 1) respectively. The corresponding point m satisfies the geometric constraints:
[0060]
[0061] Where F is a 3×3 matrix with 7 degrees of freedom.
[0062] Assume that the camera parameters are C={C1,C2,C3·····C n}, the point cloud coordinates are X={X1,X2,X3·····X n}, project the point cloud coordinates into the image plane coordinate system through perspective projection transformation, and the projection error is assumed to be g(C,X).
[0063] The SfM algorithm
[25] is used to find the camera and scene parameter optimization targets from the image to solve the image matching sparse point cloud. SfM mainly estimates the parameter information of the camera C and the three-dimensional coordinates of the point cloud X by nonlinear minimization g(C,X):
[0064]
[0065] Where: n is the number of images taken by the camera, m is the number of feature matching points. j On image i, then ω i,j =1, otherwise ω i,j =0,||q i,j -p(C i ,X j )|| 2 is the projection error of the point in the camera.
[0066] At the same time, the Random Sample Consensus (RANSAC) algorithm and Bundle Adjustment (BA) are used to project the point clouds onto their respective images and find the corresponding points and pixel coordinates. After the BA iteration, a panoramic sparse point cloud is obtained.
[0067] Finally, the sparse point cloud is transformed into a high-density 3D point cloud through PMVS, using the camera's pose information and multi-view consistency verification. Based on the feature points in the sparse point cloud, patches are generated, and multi-view verification methods such as photometric and parallax consistency are used to ensure the matching accuracy of these image patches from different viewpoints. Verified image patches are further expanded, and the accuracy of depth estimation is improved through normal vector optimization. During the final dense point cloud construction process, PMVS integrates image patches that meet consistency requirements into 3D points. The coordinates f of each point can be obtained by weighted averaging the multi-view projection results using Equation 6. Texture information is then mapped to the point cloud, completing the digital construction of the 3D model, image dense matching (DIM) point cloud, and digital orthophoto map (DOM).
[0068]
[0069] Where: N represents the number of viewpoints; Ri and ti represent the rotation matrix and translation vector of the i-th viewpoint respectively, pi is the image projection point in the viewpoint, and K is the camera intrinsic parameter matrix.
[0070] Interpretation of geological information using smartphone models
[0071] ISfM is a cost-effective 3D reconstruction method that not only reconstructs the 3D morphology of a scene but also maps the image's color and texture onto the model's surface. This method restores the scene's true texture, providing a more intuitive and richer visual experience. Furthermore, this process generates digital models in multiple dimensions, including 3D DIM point clouds and 2D DOMs. The ISfM-PMVS-based digital models of various dimensions, including 3D real-world models and 2D DOMs, can effectively interpret stratigraphic profiles and morphological maps. Based on the DOM, the intelligent Darknet-53 network architecture accurately decomposes and records the geometric morphology of river sediments at geological sites.
[0072] DOM maps the color and texture information of an image onto a 3D model through orthorectification, eliminating geometric errors caused by factors such as terrain undulation, camera tilt, and lens distortion. It can effectively analyze the accumulation form, particle size distribution, and morphological characteristics of river sediments. Darknet53 is a convolutional neural network model proposed by Joseph Redmon, consisting of 53 convolutional layers and 5 max-pooling layers. The network structure is as follows: Figure 5 The network draws on the concept of pyramid feature maps: small feature maps are used to detect large objects, while large feature maps are used to detect small objects. Furthermore, multiple convolutional and residual connection layers are stacked to improve feature extraction. Three feature maps are output: the first is downsampled by a factor of 32, the second by a factor of 16, and the third by a factor of 8. The concat operation concatenates feature maps according to the channel dimension. For example, concatenating two 8x8x16 feature maps generates an 8x8x32 feature map.
[0073] The Darknet53 network training for DOM identification of geological points in river sediments is calculated according to Formula 7, which is the proportion of positive samples detected by the model.
[0074]
[0075] Where TP means that the prediction result matches the actual object, and FP means that the prediction result does not match the actual object, but the algorithm identifies it as the target object.
[0076] Accuracy verification of geological point models built using smartphones
[0077] The high-definition photos taken by smartphones are used to construct the 3D model of geological points through ISfM-PMVS. Figure 6The model demonstrates a high degree of texture restoration, with clear surface textures and realistic colors, delicately depicting the texture and light and shadow variations of the geological relics. Because mobile phone photos lack geographic coordinates, this paper compares the relative distances of different entities on the digital model with those measured using a tape measure to ensure the model meets requirements in practical applications.
[0078] In the absence of geographic coordinates, the accuracy verification of the model mainly relies on the assessment of relative accuracy and the inspection of geometric consistency. The length and width of the chopsticks and the scale ruler were measured with a tape measure and the difference analysis was performed with the digital model measurements. The results are shown in Table 1. The maximum error is 0.11 cm and the minimum error is 0.01 cm. The centimeter-level geometric error can, to a certain extent, reflect the effectiveness of the constructed digital model.
[0079] Table 1 Comparative analysis of digital model accuracy of geological points
[0080]
[0081] Smartphones build models to interpret geological information
[0082] The river terraces on the right bank of the Jinsha River in Shaba Village are located near Shaba Village downstream of Shigu Town. Four levels of river terraces have developed on the right bank of the Jinsha River. River sediments are also called alluvial deposits. When the longitudinal slope of the riverbed decreases and the width widens, the riverbed becomes flat, the flow rate slows, and the river's carrying capacity decreases, sedimentation occurs. The resulting sediments have a distinct binary structure, with floodplain deposits on the upper part and riverbed deposits on the lower part. A cross-section of river sediments on the right bank of the Jinsha River on the north slope of Shaba Village is shown in the figure below. Figure 6 It can be found that the upper part is composed of silt and mud with a thickness of about 50 to 60 cm, which belongs to the floodplain phase sediments; the lower part is composed of riverbed gravel (golden sandstone) with a thickness of more than 2.5 m. The gravel composition is complex (all three types of rocks are present, which should be consistent with the bedrock lithology of its upstream basin). The gravels are of different sizes, highly rounded, well sorted, and have a certain orientation and are arranged in a shingle-like manner.
[0083] Based on the calculation of the processor is 12th Gen Intel Core i7-12700H, the graphics card is NVIDIA GeForce RTX3070Ti Laptop GPU, calling the Darknet-53 network, where the epoch, batch size and chip_size are 200, 32 and 224 respectively. The verification is carried out with 20% of the data, and the results are as follows Figure 7The blue curve represents the loss change of the training set, and the orange curve represents the loss change of the validation set. From the change of the loss function, it can be seen that the model shows good convergence on both the training and validation sets. The loss function continues to decrease and tends to be stable. The similarity between the validation loss and the training loss shows that there is no overfitting phenomenon in the model. The accuracy of the model in detecting sediments on the validation set is 0.82, indicating that the Darknet-53 network has good generalization ability after the loss function converges.
[0084] The Darknet-53 network detection box is a rectangle, and the perimeter needs to be revised at the location of the identified sediment. A total of 397 sediments were detected, and the results are as follows Figure 8 The statistical results are shown in Table 2. Small shapes with a Circumference < 50 are densely distributed on the left and center of the image, indicating that the shapes in these areas have relatively small circumferences. Shapes with a Circumference of 50 ≤ < 100 are primarily concentrated in the center of the image, with a relatively even and dense distribution. These are small sediment shapes, but they account for the largest proportion, reaching 50.60% of the area. Shapes with a Circumference of 100 ≤ < 150 are primarily distributed in the center-left and upper-right corners of the image, indicating that these areas have a higher concentration of medium-sized sediments. Shapes with a Circumference of 150 ≤ < 200 are relatively rare, accounting for 6.05%, and are more dispersed, mainly concentrated on the right and upper sides of the image. Shapes with a Circumference of 200 ≤ < 300 are larger sediment shapes, distributed in the center and right sides of the image, and account for the smallest proportion, at 3.53% of the entire area.
[0085] Table 2 Statistics of river sediment perimeter
[0086]
[0087]
[0088] There are missing holes in the point cloud in the 3D model built using a smartphone, indicating that in field data collection, the viewing angle, direction, and distance of the mobile phone shooting need to be adjusted accordingly according to the actual conditions of the scene targets, so as to improve the missing matching feature point cloud in the SFM-MVS modeling process. Of course, the modeling holes of the tiny gaps between the inlaid rocks are insurmountable and unavoidable due to the mobile phone shooting itself, which requires necessary supplementary measurements during field surveys.
[0089] The processing method based on mobile phone imaging modeling technology used in this embodiment shows the advantages of convenient and rapid recording of geological relics, but it has its limitations for the observation and modeling of larger scenes. In this regard, it is possible to consider the application of integrated air-ground technology and attempt to produce results supported by air-ground observation methods to demonstrate the multi-granularity characteristics of scene modeling, that is, the multi-scale characteristics of spatial resolution. This can be used to describe the level of detail of targets such as outcrops and rock masses from the overall and local composition of the site at the landscape scale. In addition, the results produced with the support of air-ground methods are not only a description of their spatial location and the overall status of their environment, but also, from a chronological and functional perspective, are more conducive to the dynamic tracking description and scientific management of the quality status (changes and protection) of rock fossil relics and ruins. Overall, the geographical and geological ecological environment in which geological relics are located is extremely fragile, and future protection is very challenging and needs to be strengthened.
[0090] In summary, this paper proposes an ISfM-PMVS method using smartphone-captured imagery to construct a digital model of river sediments on the river terraces on the right bank of the Jinsha River in Shaba Village. By measuring the length and width of chopsticks and a scale with a tape measure and performing a difference test against the digital model measurements, the geometric error of the objects was found to be only at the centimeter level, demonstrating the effectiveness of the constructed digital model. Using digital models of different dimensions to interpret different river sediment information, the results showed that, based on analysis of the three-dimensional real-world digital model, the river sediments in the test area consisted of silt and mud, approximately 50 to 60 cm thick, representing floodplain sediments. The lower portion consisted of riverbed gravel (golden sandstone) with a thickness greater than 2.5 m. Sediment detection using DOM combined with the Darknet-53 network achieved an accuracy of 0.82, detecting a total of 397 sedimentary items. Statistical analysis of sediment perimeters revealed that sediments with a Circumference of 50 ≤ < 100 are concentrated in the center of the map, showing a relatively even and dense distribution. These sediments represent the largest proportion, reaching 50.60%. Sediments with a Circumference of 200 ≤ < 300 are larger, accounting for the smallest proportion, at 3.53% of the total area. While smartphone-generated digital models may exhibit some distortion in marginal areas, this method not only allows for the rapid and comprehensive acquisition of digital models of geoheritage sites, but also enables the intelligent interpretation of geological information, which is of great significance for the preservation of geological sites and geological heritage.
[0091] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. An ISfM-PMVS river sediment digital modeling method based on smartphone-collected images, characterized in that: The following steps are involved: S1: Use a smartphone to capture images from multiple perspectives at a predetermined location; S2: Extract feature points from each image and establish the geometric relationship between images by matching these feature points; S3: The ISfM method is used to estimate camera parameters and sparse point cloud coordinates, and the initial 3D model is constructed through RANSAC and bundle adjustment optimization. S4: Apply the PMVS algorithm and use multi-view consistency verification to generate a high-density 3D point cloud to ensure the accuracy and integrity of the model. Use texture mapping to transfer the color and detail information of the photo to the 3D model to form a digital orthophoto DOM. S5: Using the generated 3D model and 2D DOM, the geometric morphology of river sediments is identified and analyzed through convolutional neural networks; S6: Verify the model accuracy by comparing the digital model with actual measurement data, and use Deep Learning methods to detect the characteristics and distribution of sediments; The step S3 specifically includes: extracting and matching feature points in the image, and the ISfM algorithm searches for the camera and scene parameter optimization target from the image to calculate the image matching sparse point cloud. ISfM minimizes the nonlinear , estimated camera Parameter information and point cloud The three-dimensional coordinates of: (2) Where: n is the number of images taken by the camera, m is the number of feature matching points; if On image i, then =1, otherwise =0, is the projection error of the point in the camera; The random sampling consensus algorithm and bundle adjustment are used to project the point clouds onto their respective images and find the corresponding points and pixel coordinates. After the BA iteration, the sparse point cloud of the whole view is obtained.
2. The ISfM-PMVS river sediment digital modeling method based on smartphone-collected images according to claim 1, characterized in that: The step S1 specifically includes: Image acquisition was performed using a smartphone; the viewing angle, distance, and direction of movement were determined, and a total of 81 photos with a resolution of 2736x3648 were acquired at three different geological locations to achieve complete coverage and optimal image overlap.
3. The ISfM-PMVS river sediment digital modeling method based on smartphone image acquisition according to claim 1, characterized in that: The step S2 specifically includes: extracting feature points from the photos and matching each feature point to determine the feature points with the same name between different images, thereby establishing geometric constraints between the images; setting and is a pair of matching points with the same name in two adjacent images, and their coordinates are and , the corresponding points m satisfy the geometric constraints: (1) Where F is a 3×3 matrix with 7 degrees of freedom; Set the camera parameters to , the point cloud coordinates are , through perspective projection change, the point cloud coordinates are projected into the image plane coordinate system, and the projection error is assumed to be .
4. The ISfM-PMVS river sediment digital modeling method based on smartphone image acquisition according to claim 1, characterized in that: The step S4 specifically includes: using the camera posture information through PMVS to generate a high-density three-dimensional point cloud through multi-view consistency verification; based on the feature points in the sparse point cloud, generating patches, and using the multi-view verification method of photometric and parallax consistency, the photometric consistency verifies the color consistency of the image block under different view angles, and the parallax consistency controls the consistency of the geometric position to ensure the matching accuracy of these image blocks under different view angles; in areas with complex textures or sudden morphological changes, applying normal vector optimization to update the normal information of each image block to improve the accuracy of depth estimation; in the final dense point cloud construction process, PMVS integrates the image blocks that meet the consistency requirements into three-dimensional points, and the coordinates f of each point are obtained by weighted average in the multi-view projection results using formula 6; then, the texture information is mapped to the point cloud to complete the digital construction of the three-dimensional model, image dense matching point cloud and digital orthophoto; (6) Where: N represents the number of viewpoints; Ri and ti represent the rotation matrix and translation vector of the i-th viewpoint respectively, pi is the image projection point in the viewpoint, and K is the camera intrinsic parameter matrix.
5. The ISfM-PMVS river sediment digital modeling method based on smartphone image acquisition according to claim 1, characterized in that: Step S5 specifically includes: the constructed three-dimensional model and the two-dimensional DOM generated by it are used to interpret the geometric and morphological information of river sediments; with the help of the Darknet-53 convolutional neural network, the model can accurately decompose and record the geometric features of geological points on the digital orthophoto.
6. The ISfM-PMVS river sediment digital modeling method based on smartphone image acquisition according to claim 1, characterized in that: Step S6 specifically includes: ensuring the geometric accuracy of the model is within an acceptable range by comparing and analyzing the results of digital model measurements with actual on-site measurement data; using the detection frame of the Darknet-53 network to identify the location of sediments and revising the perimeter of the sediments based on the detection results; this process also requires statistical analysis of the distribution of sediments of different sizes to provide data support for river sediment analysis and geological interpretation.
Citation Information
Patent Citations
Unmanned aerial vehicle side slope and foundation pit excavation amount measuring device
CN111207728A
Construction method of urban live-action three-dimensional model based on aerial-triangulation parallel computing algorithm
CN113192200A