Method and system for automatically generating geological description based on geological photos
By preprocessing geological photos and deep learning model analysis, and combining with the geological knowledge base to generate geological descriptions, the problem of geological description relying on manual analysis is solved, and efficient and accurate generation and sharing of geological information is achieved.
Patent Information
- Application Number
- CN202510563017.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, geological description generation relies on low efficiency of manual analysis, is susceptible to environmental and time constraints, cannot be checked in real time, has low data sharing and utilization efficiency, lacks unified standards, and information omissions and errors are frequently recorded.
By pre-processing the geological photos, a geological image data model was constructed, and target detection and semantic segmentation were used using YOLO, Faster R-CNN, U-Net and DeepLabV3+ models, and geological descriptions were generated in combination with the geological knowledge base.
Real-time generation of geological descriptions is realized, artificial errors are reduced, data accuracy and consistency are improved, data sharing and management are supported, and professional thresholds are lowered.
Smart Images

Figure CN120472465A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of geological exploration technology, and in particular to a method and system for automatically generating geological descriptions based on geological photographs. Background Art
[0002] With the continuous development of geological exploration technology, traditional geological methods often rely on geologists' on-site records or manual analysis, which are inefficient and easily affected. In recent years, the rapid development of deep learning and computer vision technology has provided many new ideas for geological exploration in the field of geology.
[0003] However, the current technology for automatic generation of geological descriptions for mobile devices is not yet mature. There is an urgent need for an efficient and customized method that can quickly generate high-quality geological descriptions using mobile phone camera data to assist geological work. Summary of the Invention
[0004] The embodiments of the present invention provide a method and system for automatically generating geological descriptions based on geological photographs, aiming to improve the real-time performance and data sharing capabilities of field recording work and reduce the problems of information omission, misrecording and low efficiency in traditional manual recording.
[0005] In a first aspect, an embodiment of the present invention provides a method for automatically generating a geological description based on a geological photograph, comprising:
[0006] S1, performing image preprocessing on the geological field image to obtain a preprocessed image;
[0007] S2, constructing a geological image data model and training the geological image data model; the input of the model is an image containing geological attributes, and the output of the model is a description text describing geological information;
[0008] S3, inputting the pre-processed image into the geological image data model, and using the output result of the model as the initial description text;
[0009] S4, matching the initial description text and the coordinate information of the geological site image with a geological knowledge base; supplementing the initial description text with relevant information in the geological knowledge base to generate a final text entity.
[0010] Furthermore, in S1, the image preprocessing process includes: converting the geological field image into a preprocessed image by performing normalization processing, noise removal, contrast enhancement and ROI extraction on the geological field image;
[0011] The normalization process includes: adjusting the pixel range, image size and number of image channels to be consistent with the input of the geological image data model;
[0012] The noise removal process includes: eliminating noise points generated by environmental factors through Gaussian filtering or bilateral filtering;
[0013] The contrast enhancement process includes: adjusting brightness distribution and visual effects through histogram equalization or adaptive histogram equalization;
[0014] The ROI extraction process includes: dividing the region of interest through edge detection, color segmentation or deep learning segmentation.
[0015] Furthermore, in S2, the constructing of the geological image data model and the training of the geological image data model include:
[0016] S201, obtaining a self-test dataset and / or an open source dataset as training data;
[0017] S202, performing data enhancement on the training data by random rotation, random scaling, noise injection and / or color adjustment;
[0018] S203, performing image classification on the data-enhanced training data to obtain multiple data sets of different categories; according to the size of the data set, different data set models are adopted to optimize the data set;
[0019] S204, perform target detection on the images in the dataset using the YOLO or Faster R-CNN model;
[0020] S205, performing semantic segmentation on the detected image using a U-Net or DeepLabV3+ model to obtain a classification prediction result for each pixel;
[0021] S206 , segmenting the stratigraphic region according to the classification prediction result; and generating a description text describing the geological information according to the geological attributes of the stratigraphic region.
[0022] Furthermore, in S4, the process of matching the initial description text and the coordinate information of the geological site image with the geological knowledge base includes:
[0023] S401, determining coordinate information of a geological site image;
[0024] S402, querying the geological knowledge base to obtain geological background information of the corresponding area;
[0025] S403: Inferring information required for geological description of the corresponding area based on the initial description text and the geological background information.
[0026] Furthermore, in S4, the matching method of the geological knowledge base includes: feature extraction comparison, spatial and context correction, and / or type confidence;
[0027] The feature extraction comparison includes: color feature comparison, texture feature comparison, and / or shape feature comparison;
[0028] The spatial and contextual correction includes: spatial consistency correction, and / or neighborhood consistency correction;
[0029] The type confidence includes: calculating and adjusting the confidence of the model output result.
[0030] Furthermore, before S4, it also includes:
[0031] Performing feature extraction on the preprocessed image;
[0032] The feature extraction includes: color feature extraction, texture feature extraction, and / or shape feature extraction;
[0033] The steps of color feature extraction include: color space conversion, color histogram and principal component analysis;
[0034] The steps of extracting texture features include: gray level co-occurrence matrix, local binary pattern and wavelet transform;
[0035] The steps of extracting shape features include edge detection, region analysis and inclination angle calculation.
[0036] Furthermore, after S4, it also includes:
[0037] S5, encapsulate the image preprocessing and geological image data model calls into corresponding APIs.
[0038] In a second aspect, an embodiment of the present invention provides a system for automatically generating geological descriptions based on geological photographs, comprising:
[0039] A preprocessing unit, used for performing image preprocessing on the geological field image to obtain a preprocessed image;
[0040] A model building unit is used to build a geological image data model and train the geological image data model; the input of the model is an image containing geological attributes, and the output of the model is a description text describing geological information;
[0041] A model application unit, configured to input the pre-processed image into the geological image data model and use the output of the model as an initial description text;
[0042] The text generation unit is used to match the initial description text and the coordinate information of the geological site image with the geological knowledge base; after successful matching, the initial description text is supplemented with relevant information in the geological knowledge base to generate a final text entity.
[0043] In a third aspect, an embodiment of the present invention provides an electronic device, comprising:
[0044] at least one processor; and
[0045] a memory communicatively connected to the at least one processor; wherein,
[0046] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for automatically generating geological descriptions based on geological photographs according to any embodiment of the present invention.
[0047] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for automatically generating geological descriptions based on geological photographs as described in any embodiment of the present invention when executed.
[0048] In an embodiment of the present invention, after capturing an image of the address at the scene, a text description can be quickly generated. By matching the AI model with the knowledge base, a geological description corresponding to the photo is automatically generated, eliminating the tedious process of manual analysis and recording. Users can obtain geological information feedback in real time without having to wait for laboratory analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0050] Figure 1 1 is a flow chart of a method for automatically generating geological descriptions based on geological photographs according to an embodiment of the present invention;
[0051] Figure 2 This is a code diagram of a combination of YOLO and U-Net according to an embodiment of the present invention;
[0052] Figure 3 is a schematic diagram of a code by color feature comparison provided by an embodiment of the present invention;
[0053] Figure 42 is a schematic diagram of a system for automatically generating geological descriptions based on geological photographs according to an embodiment of the present invention;
[0054] Figure 5 It is a structural diagram of an electronic device implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0055] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0056] In addition, before introducing the embodiments of the present invention, it should be noted that the conventional geological data collection methods currently have at least the following problems:
[0057] Strong Advantages: Geological descriptions rely on personal experience and the expertise of the data collector, which can lead to discrepancies between different descriptions of the same geological body. A lack of unified standards means that the vocabulary and format for lithologic classification and structural characterization are not standardized, hindering subsequent data analysis and sharing.
[0058] Inefficiency: Manual recording is cumbersome. Traditional data collection methods typically require manual recording of information, including text and sketches, which is labor-intensive and inefficient. Subsequent compilation is complex, as field records must be manually organized and digitized, which can easily lead to omissions or misrecording, further increasing the workload.
[0059] Environmental and time constraints: Complex terrain impacts data collection. The intensity of data collection in rugged plateaus or high-altitude climates affects the comprehensiveness and accuracy of recorded information. Limited data collection range: The collector's reach limits data coverage, potentially leading to information loss.
[0060] Inability to verify and adjust in real time: Descriptions cannot be verified immediately. After field descriptions are completed, the data must be verified in the laboratory. If errors or omissions are found, data collection must be repeated, increasing costs. Traditional methods lack on-site analysis capabilities and cannot integrate historical data or knowledge bases for real-time analysis, thus missing opportunities for optimizing data collection.
[0061] Inefficient data sharing and utilization: Data formats are fragmented, with different collectors or data generated often stored in different formats, hindering unified integration. Previous analytical methods and traditional data collection methods often generated static data that lacked compatibility with technologies like geographic information systems (GIS) and 3D modeling, limiting its value in subsequent applications.
[0062] Example 1
[0063] Figure 1 The present invention provides a flowchart of a method for automatically generating geological descriptions based on geological photographs. This embodiment is applicable to situations where high-quality geological descriptions are quickly generated using photographic data. The method can be executed by a system for automatically generating geological descriptions based on geological photographs in an embodiment of the present invention. The system can be implemented in software and / or hardware and integrated into electronic devices.
[0064] like Figure 1 As shown, the method specifically includes the following steps:
[0065] S1, performing image preprocessing on the geological field image to obtain a preprocessed image.
[0066] The image preprocessing process includes: converting the geological field image into a preprocessed image by normalizing, removing noise, enhancing contrast and extracting ROI.
[0067] In an embodiment of the present invention, geological photos are taken by mobile devices (such as mobile phones, etc.) to record geological field images. The relevant photos need to contain fault information (such as GPS coordinates, etc.) and support multi-angle acquisition to increase the comprehensiveness of features. The collected geological field images are subjected to image preprocessing to provide high-quality input data for subsequent analysis. During the image preprocessing process, the relevant functions can be implemented using mainstream image processing libraries (such as OpenCV, Pillow, Scikit-Image, etc.) and / or functions of deep learning frameworks (such as TensorFlow, PyTorch, etc.). After the relevant processing, the preprocessed image that meets the requirements is returned.
[0068] Furthermore, the normalization process includes adjusting the pixel range, image size and number of image channels to be consistent with the input of the geological image data model.
[0069] Specifically, normalize the input image so that its pixel value range meets the model requirements (usually 0-1 or -1 to 1). Resize the image (e.g., 640×640 or 512×512) to accommodate the input requirements of the target detection model. If the image has multiple channels (e.g., RGB) or a single channel (e.g., grayscale), ensure that the number of image channels is consistent with the input of the geological image data model.
[0070] Furthermore, the noise removal process includes: eliminating noise points generated by environmental factors through Gaussian filtering or bilateral filtering.
[0071] Specifically, field photos may be affected by environmental factors such as uneven lighting, dust, and reflections, resulting in noise. Both Gaussian filtering and bilateral filtering can eliminate noise. The principles and effects of noise elimination differ, but there is a certain connection. Gaussian filtering: It smoothes noise by weighting the values of surrounding pixels while retaining the main image structure. It is suitable for photos with uniform noise in the scene. Bilateral filtering: It considers the similarity and color similarity of spatially adjacent pixels at the same time, smoothing noise while retaining the edge features of the image. It is particularly suitable for bedding lines or crack images in geological photos.
[0072] Furthermore, the contrast enhancement process includes adjusting brightness distribution and visual effects through histogram equalization or adaptive histogram equalization.
[0073] Specifically, photos taken may appear blurry or flat due to insufficient light, and contrast enhancement can improve the visual effect of the photo. Among them, Histogram Equalization (HE) and Adaptive Histogram Equalization (CLAHE) are closely related to restoration enhancement, aiming to improve the brightness distribution and visual effect of the image. Histogram Equalization: By redistributing the brightness of pixel values, it increases the details of high-contrast areas. It is especially effective for scenes with strong light and dark differences (such as shadows or reflections) in the photo; Adaptive Histogram Equalization (CLAHE): Equalizes the local area of the image to avoid the problem of over-brightening that may be caused by histogram equalization. It works better in complex geological backgrounds (such as the coexistence of multiple rock textures and cracks).
[0074] Furthermore, the ROI extraction process includes dividing the region of interest through edge detection, color segmentation or deep learning segmentation.
[0075] Specifically, the entire photo may contain a large amount of irrelevant background (such as vegetation, sky, or human factors). By extracting the area of interest (such as part of the rock formation, etc.), you can focus on the analysis target. Edge detection: Use the Canny edge detection method to identify the boundaries of rock formations and extract geometric features such as cracks and joints. The accuracy of boundary recognition can be enhanced by adjusting the detection threshold. Color segmentation: Based on the color characteristics of the rock area, the K-means clustering method is used to divide the photo into different color groups, and the main rock area is retained. Deep learning segmentation: Use the pre-trained Mask R-CNN model to segment the photo, automatically mark the rock area and exclude the background.
[0076] In an embodiment of the present invention, the geological field image is first converted into the input format of the geological image data model through normalization processing; then, noise points generated by environmental factors in the image are eliminated through noise removal; then, the visual effect of the image is improved through contrast enhancement; and finally, the region of interest in the image is divided through ROI extraction.
[0077] S2, constructing a geological image data model and training the geological image data model; the input of the model is an image containing geological attributes, and the output of the model is a description text describing the geological information.
[0078] The construction of the geological image data model and the training of the geological image data model include:
[0079] S201: Obtain a self-test dataset and / or an open source dataset as training data.
[0080] Among them, self-check datasets: geological photos collected in the field, including different types of lithology and geological features (such as rock type, bedding, cracks, folds, etc.); open source datasets: public datasets specifically used for rock classification, and public datasets containing mineral classification information.
[0081] S202 , performing data enhancement on the training data by random rotation, random scaling, noise injection, and / or color adjustment.
[0082] Random rotation simulates shooting from different angles; random scaling improves the model's adaptability to scenes at varying distances; noise injection enhances the model's robustness under complex lighting or environmental conditions; and color adjustment addresses issues such as inconsistent ambient lighting (e.g., white balance). These random transformations can be implemented through function calls in several popular Python libraries, further enhancing image data. Popular libraries include Albumentations, imgaug, OpenCV, and TensorFlow / Keras.
[0083] Albumentations: Provides fast data manipulation enhancements. Albumentations is an efficient image enhancement library designed specifically for deep learning tasks, providing a variety of simple and fast image processing functions. The library seamlessly integrates with deep learning frameworks such as PyTorch and TensorFlow, supports common image enhancement operations such as rotation, flipping, scaling, cropping, and adding noise, and provides a variety of transformations including geometric transformations, color adjustments, and blurring, making it easy to create complex enhancement pipelines while keeping the code concise.
[0084] imgaug: supports complex enhancement strategies (such as affine transformation and restoration enhancement). imgaug is also a powerful image enhancement library, suitable for tasks that require complex enhancement strategies. It provides a series of enhancement operations, supports more custom transformation operations, has excellent scalability, and provides various image enhancement operations such as rotation, scaling, shearing, translation, color adjustment, etc. At the same time, it supports advanced operations such as geometric transformation (affine transformation, perspective transformation) and image blending. It can combine multiple transformations to build a complex enhancement process, and supports batch operations, which can enhance multiple images at once.
[0085] OpenCV: An open-source computer vision library, it includes a rich set of image processing functions widely used in computer vision tasks. It can be used not only for image enhancement but also for complex operations such as image preprocessing, feature extraction, and object detection. It provides a variety of basic image processing functions, including rotation, flipping, translation, scaling, color conversion, and noise addition. It also supports a variety of custom image enhancement operations, allowing you to write more complex enhancement strategies as needed. OpenCV provides strong support for efficient image processing and real-time processing.
[0086] TensorFlow / Keras: ImageDataGenerator is a built-in data augmentation tool provided by TensorFlow / Keras, specifically designed for deep learning tasks. It performs real-time augmentation by generating batches of augmented datasets, minimizing memory consumption. It also provides a variety of common augmentation operations, such as rotation, translation, scaling, and flipping. Furthermore, it automatically augments images in conjunction with the fit function, making it suitable for real-time data augmentation during model training. It also supports real-time batch generation and augmentation, significantly reducing memory consumption.
[0087] The present invention increases sample diversity by randomly transforming the training data, improving the model's generalization ability, reducing overfitting, and enhancing the model's robustness in real-world scenarios. At the same time, part of the training data will be used as test data to complete the model testing.
[0088] S203, performing image classification on the data-enhanced training data to obtain a plurality of data sets of different categories; and optimizing the data sets using different data set models according to the size of the data sets.
[0089] The training data is then classified into specific rock types (e.g., sandstone, shale, limestone, etc.). For medium-sized datasets (hundreds to thousands of samples), the ResNet model is used. This effectively addresses the vanishing gradient problem in deep networks, making it suitable for complex classification tasks and suitable for cloud-based inference or high-performance devices. For large datasets, the EfficientNet model is used, optimizing the model structure for greater accuracy and efficiency.
[0090] S204, performing target detection on the images in the dataset using the YOLO or Faster R-CNN model.
[0091] Specifically, the YOLO or Faster R-CNN model is selected to process the image input, mark the location and category of geological objects (such as cracks, joints, etc.) in the image, extract the detected geological structure information (such as cracks, joints, faults, etc.), and record its location, size, direction and category.
[0092] YOLO: A single-stage object detection model with fast inference, suitable for field applications requiring efficient labeling. It also performs full-image detection, predicting bounding boxes and class labels through full-image regression. This fast model is suitable for rapid field detection of geological features, such as real-time identification of cracks, joints, and faults. The YOLO model directly outputs bounding boxes and class labels for an image.
[0093] Faster R-CNN: This model, based on a two-stage detection process (region proposal generation and refined classification), is suitable for detecting complex geological structures. It offers flexibility and generates candidate boxes using a region proposal network (RPN), which are then further classified and repositioned. This model is suitable for complex features that require detailed classification and labeling, such as small cracks or irregular structures. The Faster R-CNN model uses the RPN and classifier to generate candidate boxes and classify images, respectively.
[0094] S205, performing semantic segmentation on the detected image through the U-Net or DeepLabV3+ model to obtain the classification prediction result of each pixel.
[0095] Specifically, the task of semantic segmentation is to classify each pixel in the image, perform pixel-level classification, and mark different areas (such as layers and rock formation boundaries).
[0096] U-Net: This model uses an encoder-decoder architecture and is well-suited for small-sample tasks. It supports high-precision pixel-level predictions, especially for regions with clear boundaries. It is also suitable for small-sample segmentation tasks, such as labeling rock formation boundaries or stratigraphic sequences. The U-Net model is used for lithologic stratification and sedimentary structure analysis.
[0097] DeepLabV3+: Featuring high resolution, DeepLabV3+ uses atrous convolution and an encoder-decoder architecture to generate high-resolution segmentation results in complex environments. It also incorporates multi-scale information, making it suitable for segmenting irregular regions and complex geological environments. For example, it can identify tilted or broken rock formations, allowing for stratigraphic region delineation in field photos and accurate segmentation of complex structures.
[0098] S206, segmenting the stratigraphic region according to the classification prediction result; and generating a description text describing the geological information according to the geological attributes of the stratigraphic region.
[0099] Specifically, after selecting the U-Net or DeepLabV3+ model, you can use a deep learning framework (such as TensorFlow or PyTorch) to train the U-Net or DeepLabV3+ model. After the model training is complete, the test data is input into the model to generate a category prediction for each pixel. The prediction results are superimposed on the original image for visualization and output to form a segmentation map. By analyzing the segmented stratum area, the geological attribute information of each layer, such as lithology, thickness, sequence, and inclination, is finally extracted, and a description text describing the geological information is generated.
[0100] In this embodiment of the present invention, the collected data can be deeply analyzed to generate geological image data models and make predictions, helping geologists make more informed decisions. Furthermore, through big data analysis, potential geological risks (such as landslides and earthquakes) can be identified in advance, providing risk warnings and management strategies.
[0101] In the embodiment of the present invention, any model in steps S201 - S206 may be adjusted to construct different geological image data models, thereby enhancing the diversity and selectivity of the description text.
[0102] In an embodiment of the present invention, image recognition technology is used, and deep learning algorithms such as target detection and semantic segmentation are applied to geological images to accurately extract geological features and realize the transformation from visual information to structured description.
[0103] S3, inputting the preprocessed image into the geological image data model, and using the output of the model as the initial description text.
[0104] Specifically, the preprocessed image is first subjected to target detection using the YOLO or Faster R-CNN model to quickly identify the location of geological features. The detected image is then subjected to semantic segmentation using the U-Net or DeepLabV3+ model to achieve fine segmentation of the detection area. Finally, the accuracy of the overall analysis can be improved by using the results of target detection and semantic segmentation as complementary information.
[0105] Figure 2 This is a code diagram of the combination of YOLO and U-Net provided by the embodiment of the present invention. Figure 2 The code implementation includes two parts: YOLO extraction and detection of geological structure information, and U-Net analysis and segmentation of stratum areas.
[0106] After YOLO is used to extract and detect geological structure information (such as cracks, joints, faults, etc.), its location, size, direction and category are recorded.
[0107] For example, the geological structure information is described as follows:
[0108] Gap: width 5cm, direction N45°E; Fault: type is normal fault, dip angle 35°.
[0109] The segmented stratigraphic area is analyzed by U-Net to extract the lithology, thickness, sequence, dip and other attribute information of each layer.
[0110] For example, the attribute information of each layer is described as follows:
[0111] The first layer: sandstone, thickness 3.2m, inclination 15°; the second layer: shale, thickness 2.8m, inclination 10°.
[0112] In the embodiment of the present invention, the geological structure information and the attribute information of each layer are both used as part of the output of the model to describe the geological information and serve as the initial description text.
[0113] S4, matching the initial description text and the coordinate information of the geological site image with the geological knowledge base; supplementing the initial description text with relevant information in the geological knowledge base to generate the final text entity.
[0114] The process of matching the initial description text, the coordinate information of the geological site image, and the geological knowledge base includes:
[0115] S401, determining the coordinate information of the geological site image.
[0116] The geographic coordinates of the capture point can be determined using the device's GPS data or the photo's EXIF metadata. For example, the latitude and longitude are: N30.12345°, E120.67890°.
[0117] S402: Query the geological knowledge base to obtain geological background information of the corresponding area.
[0118] The geological knowledge base is queried to obtain the geological background information of the region (such as regional lithology distribution, historical tectonic movement, etc.). For example, the region is mainly composed of middle-age sandstone layers with obvious fault structures.
[0119] S403: Inferring the information required for the geological description of the corresponding area based on the initial description text and the geological background information.
[0120] For example, it can be described as follows: Basic description: The collection point is located in the Mesozoic sedimentary rock area, with the main rock types being sandstone and shale; Structural characteristics: There is an active crack 5 m wide, oriented N45°E, which may be related to a regional fault; Stratigraphic information: The strata are sandstone (3.2 m) and shale (2.8 m) in sequence, with an overall dip of 15°.
[0121] It should be noted that different geological image data models will produce different output texts. Similarly, different geological databases store geological information in different template formats, resulting in differences in the information required for geological description.
[0122] For example, it can also be described as follows: Lithology description: the classified lithology name (such as "sandstone, likely composed of quartz and feldspar"), and the confidence score (such as "sandstone 92%"); Geological characteristics: detailed information on detected fractures, joints, and other features. For example, "the fracture direction is NW-SE, the length is approximately 5 meters, and the dip angle is 30°."
[0123] Furthermore, by matching relevant information in the geological knowledge base, the geological environment and historical data can be inferred. The knowledge base content includes: rock fracture types, formation causes, bedding morphology, color, texture, shape characteristics, geological composition, chemical indicators; geological stratigraphic data, the distribution area of various rocks, and rock characteristics statistics in different geological environments; characteristics of known samples, typical characteristics of each rock type, historical collection data, and related annotation information.
[0124] Furthermore, the matching method of the geological knowledge base includes: feature extraction comparison, spatial and context correction, and / or type confidence.
[0125] Furthermore, the feature extraction comparison includes: color feature comparison, texture feature comparison, and / or shape feature comparison.
[0126] Specifically, the features in the initial description text generated by target detection and semantic segmentation are matched with feature templates of corresponding types in the geological knowledge base.
[0127] Use color features: Compare the color distribution of the target area with the color features of different rock types in the knowledge base. For example, calculate the similarity of HSV histograms (such as cosine similarity or histogram intersection).
[0128] Figure 3 is a schematic diagram of a code by color feature comparison provided by an embodiment of the present invention, combined with Figure 3 ,If the similarity of the relevant feature similarity matching is higher than the set threshold (such as 0.8), the match is considered successful.
[0129] Use texture features: Based on GLCM, LBP or depth features, compare the texture information of the segmented area with the statistical features of the knowledge base.
[0130] Use shape features: Use methods such as Hu moments and edge features to match the shape of the segmented area with the shape model in the knowledge base.
[0131] Furthermore, before S4, it also includes:
[0132] Perform feature extraction on the preprocessed image; feature extraction includes: color feature extraction, texture feature extraction, and / or shape feature extraction.
[0133] Specifically, the preprocessed images can be converted into quantitative data for describing and distinguishing geological features. The extracted features are mainly divided into color features, texture features, and shape features, which can be extracted using OpenCV, Scikit-Image, and Numpy related functions.
[0134] The steps of color feature extraction include: color space conversion, color histogram and principal component analysis.
[0135] Specifically, different rock types usually have specific color distributions, for example: basalt is usually dark gray or black, and limestone is usually light gray or white. The steps of color feature extraction include: color space conversion, color histogram, principal component analysis (PCA) and other related steps. Color space conversion: convert the image from RGB space to HSV (hue, saturation, brightness) or LAB (perceptual brightness, green-red, blue-yellow) space to facilitate color analysis. Color histogram: calculate the distribution of each color component in the image, for example, count the proportion of pixels with the main color within a certain range. Principal component analysis (PCA): reduce the dimension of color features and extract the main color components for rock type classification.
[0136] The steps of texture feature extraction include: gray-level co-occurrence matrix, local binary pattern and wavelet transform.
[0137] Specifically, the texture features in geological photos reflect the mineral composition, sedimentary environment, etc. of the rock. For example, sandstone may show a rough granular texture, while shale shows a fine layered texture. The steps of texture feature extraction include: gray level co-occurrence matrix (GLCM), local binary pattern (LBP), wavelet transform and other steps. Gray level co-occurrence matrix: extracts the texture features of the image by calculating the grayscale relationship between pixels, including: contrast: reflects the difference in depth of the texture; homogeneity: indicates the smoothness of the texture; energy: indicates the regularity of the texture. Local binary pattern (LBP): used to describe the local features of the texture, and extracts the microscopic texture structure by encoding the grayscale pattern of the pixel points. Wavelet transform: used to analyze the multi-scale characteristics of the texture, decompose the image into low-frequency and high-frequency components, and capture the global and local characteristics of the texture.
[0138] The steps of shape feature extraction include edge detection, region analysis and inclination angle calculation.
[0139] Specifically, the geometric shape of the rock layer (such as bedding, cracks, joints, etc.) can reflect the geological structure and stress environment. The steps of shape feature extraction include: edge detection, regional analysis, and inclination calculation related steps. Edge detection: Use the Canny or Sobel algorithm to extract obvious geometric shapes in the image, and use the Hough transform to identify straight lines (bedding direction) or circles (mineral composition distribution). Regional analysis: Calculate the area, perimeter, and shape factor of the connected area to describe the morphology of the rock. Dip calculation: Based on the detected bedding line direction, the scale is combined to infer the inclination of the rock layer.
[0140] The embodiment of the present invention extracts features from the preprocessed image and adopts a matching method of feature extraction and comparison when matching with the geological knowledge base, thereby being able to quickly establish a connection between the image and the geological knowledge base, thereby improving the geological description of the geological field image.
[0141] Furthermore, the spatial and contextual correction includes: spatial consistency correction, and / or neighborhood consistency correction.
[0142] Specifically, based on the segmented or detected geographic location information in combination with GPS coordinates, the results are cross-validated with the geological stratigraphic data in the knowledge base to determine whether the geological type of the location is consistent with the knowledge base record.
[0143] Spatial consistency correction: Whether the detected rock type is consistent with the known geological distribution of the area. For example, if the knowledge base shows that the area is mainly sandstone, but the result is identified as shale, a probability adjustment is performed.
[0144] Neighborhood consistency correction: Use the detection results of adjacent areas to assist in judgment. For example, sandstone is usually not directly adjacent to basalt. If such a combination occurs, annotation correction can be performed.
[0145] Furthermore, type confidence adjustment involves calculating and adjusting the confidence of the model output. The confidence of the model output is recalculated based on statistical data from the knowledge base. If the confidence of the matching result is low but the geological knowledge base supports the classification, the confidence is increased. If the confidence of the matching result is high but inconsistent with the knowledge base, the confidence is lowered or a manual review is prompted.
[0146] For example, it is displayed in json format as follows:
[0147] {
[0148] "location":{"latitude":34.05,"longitude":-118.25},
[0149] "detected_type":"Sandstone",
[0150] "adjusted_confidence":0.92,
[0151] "segment_area":1500,
[0152] "matched_with_knowledge":true
[0153] }
[0154] In this embodiment of the present invention, a unified data structure (such as JSON) can be used to store recognition results, facilitating subsequent analysis, storage, and sharing. Furthermore, the model and data are cyclically optimized, and the collected field data can be fed back into the training dataset, continuously optimizing the model's detection and classification capabilities. The integration of the knowledge graph and the geological knowledge base enables intelligent, context-sensitive geological descriptions by building a geological knowledge base and combining it with AI model output.
[0155] After the final text entity is generated in step S4, the following steps are also included:
[0156] S5, encapsulate the image preprocessing and geological image data model calls into corresponding APIs.
[0157] The image preprocessing and geological image data model calls are encapsulated into corresponding APIs, and the API calls are made after the mobile terminal takes a photo.
[0158] By formatting the text entities output by the model, the geological description information of the current location can be parsed. At this time, it is determined whether the mobile terminal collection form exists. If it does not exist, the collection form is created first; if it exists, the relevant information is directly filled into the relevant collection form.
[0159] Field data collectors check the data on relevant collection forms and then upload and save them.
[0160] In this embodiment of the present invention, the collected photos and generated descriptions can be stored in digital form, promoting the transformation of geological information into a comprehensive information-based system. Furthermore, standardized technical interfaces can be used to facilitate integration with other field collection systems or geological databases, enhancing collaborative capabilities.
[0161] In summary, the technical value of automatically generating geological descriptions by collecting photos through mobile devices in the geological collection APP is mainly reflected in improving data processing capabilities, automation levels, wide application and innovation. It can promote the development and progress of related disciplines, and at the same time improve the efficiency and accuracy of geological surveys and research.
[0162] The technical solution in the embodiment of the present invention integrates the function of automatically generating geological descriptions by collecting photos through mobile devices into the geological field collection app. Compared with the existing technology, it has the following beneficial effects:
[0163] Improved Work Efficiency: This embodiment of the present invention can quickly generate text descriptions. By matching AI models with knowledge bases, it automatically generates geological descriptions corresponding to photos, eliminating the tedious process of manual analysis and record-keeping. Real-time geological information feedback is available to users during the acquisition process, eliminating the need to wait for laboratory analysis.
[0164] Enhanced data accuracy and consistency: The embodiments of the present invention can reduce human errors. The automated generation process relies on training data and algorithms to avoid omissions or subjective errors in manual records. At the same time, the geological description is standardized and output in a unified description format through knowledge base correction, which facilitates subsequent data aggregation and analysis.
[0165] Facilitates geological data management and sharing: The embodiment of the present invention adopts digital storage, and the generated geological description is directly stored in the form of structured data (such as JSON or table), which is convenient for archiving and retrieval, and cross-team sharing. Photos and generated descriptions can be uploaded to the cloud in real time for other team members to view and use.
[0166] Lowering the professional threshold: The embodiments of the present invention can assist non-professionals in field surveys. Even users without geological background knowledge can obtain accurate geological descriptions through the app, lowering the threshold for field surveys. Training aids, as teaching tools, help novices learn geological types and feature descriptions.
[0167] In summary, geological description-related data can be automatically generated based on photos collected by mobile devices, and combined with large-scale model analysis, the efficiency and accuracy of data collection and processing can be significantly improved.
[0168] Example 2
[0169] Figure 4 Schematic diagram of a system for automatically generating geological descriptions based on geological photographs according to an embodiment of the present invention. Figure 4 As shown, the system includes:
[0170] The preprocessing unit 10 is used to perform image preprocessing on the geological field image to obtain a preprocessed image;
[0171] The model building unit 20 is used to build a geological image data model and train the geological image data model; the input of the model is an image containing geological attributes, and the output of the model is a description text describing the geological information;
[0172] The model application unit 30 is used to input the pre-processed image into the geological image data model and use the output of the model as the initial description text;
[0173] The text generation unit 40 is used to match the initial description text and the coordinate information of the geological site image with the geological knowledge base; after the matching is successful, the initial description text is supplemented with relevant information in the geological knowledge base to generate a final text entity.
[0174] The technical solution in the embodiments of the present invention, which automatically generates geological descriptions from photos collected by mobile devices within a geological collection app, has the potential to improve data processing capabilities, automation, widespread application, and innovation. It has promoted the development and progress of related disciplines while also improving the efficiency and accuracy of geological surveys and research.
[0175] Example 3
[0176] Figure 5 The present invention is a block diagram of an electronic device for implementing a method for automatically generating geological descriptions based on geological photographs according to an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided for example only and are not intended to limit the implementation of the present inventions described and / or claimed herein.
[0177] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0178] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0179] Processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as the method for automatically generating geological descriptions based on geological photographs.
[0180] In some embodiments, the method for automatically generating geological descriptions based on geological photographs can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for automatically generating geological descriptions based on geological photographs described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the method for automatically generating geological descriptions based on geological photographs in any other suitable manner (e.g., via firmware).
[0181] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0182] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0183] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0184] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0185] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0186] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0187] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0188] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for automatically generating geological descriptions based on geological photographs, characterized in that: include: S1, performing image preprocessing on the geological field image to obtain a preprocessed image; S2, constructing a geological image data model and training the geological image data model; the input of the model is an image containing geological attributes, and the output of the model is a description text describing geological information; S3, inputting the pre-processed image into the geological image data model, and using the output result of the model as the initial description text; S4, matching the initial description text and the coordinate information of the geological site image with a geological knowledge base; supplementing the initial description text with relevant information in the geological knowledge base to generate a final text entity.
2. The method according to claim 1, characterized in that In S1, the image preprocessing process includes: converting the geological field image into a preprocessed image by performing normalization processing, noise removal, contrast enhancement and ROI extraction on the geological field image; The normalization process includes: adjusting the pixel range, image size and number of image channels to be consistent with the input of the geological image data model; The noise removal process includes: eliminating noise points generated by environmental factors through Gaussian filtering or bilateral filtering; The contrast enhancement process includes: adjusting brightness distribution and visual effects through histogram equalization or adaptive histogram equalization; The ROI extraction process includes: dividing the region of interest through edge detection, color segmentation or deep learning segmentation.
3. The method according to claim 1, characterized in that In S2, the step of constructing a geological image data model and training the geological image data model includes: S201, obtaining a self-test dataset and / or an open source dataset as training data; S202, performing data enhancement on the training data by random rotation, random scaling, noise injection and / or color adjustment; S203, performing image classification on the data-enhanced training data to obtain multiple data sets of different categories; according to the size of the data set, different data set models are adopted to optimize the data set; S204, perform target detection on the images in the dataset using the YOLO or Faster R-CNN model; S205, performing semantic segmentation on the detected image using a U-Net or DeepLabV3+ model to obtain a classification prediction result for each pixel; S206 , segmenting the stratigraphic region according to the classification prediction result; and generating a description text describing the geological information according to the geological attributes of the stratigraphic region.
4. The method according to claim 1, wherein In S4, the process of matching the initial description text, the coordinate information of the geological site image, and the geological knowledge base includes: S401, determining coordinate information of a geological site image; S402, querying the geological knowledge base to obtain geological background information of the corresponding area; S403: Inferring information required for geological description of the corresponding area based on the initial description text and the geological background information.
5. The method according to claim 1, wherein In S4, the matching method of the geological knowledge base includes: feature extraction comparison, spatial and context correction, and / or type confidence; The feature extraction comparison includes: color feature comparison, texture feature comparison, and / or shape feature comparison; The spatial and contextual correction includes: spatial consistency correction, and / or neighborhood consistency correction; The type confidence includes: calculating and adjusting the confidence of the model output result.
6. The method according to claim 1, characterized in that Prior to S4, it also included: Performing feature extraction on the preprocessed image; The feature extraction includes: color feature extraction, texture feature extraction, and / or shape feature extraction; The steps of color feature extraction include: color space conversion, color histogram and principal component analysis; The steps of extracting texture features include: gray level co-occurrence matrix, local binary pattern and wavelet transform; The steps of extracting shape features include edge detection, region analysis and inclination angle calculation.
7. The method according to claim 1, characterized in that After S4, it also includes: S5, encapsulate the image preprocessing and geological image data model calls into corresponding APIs.
8. A system for automatically generating geological descriptions based on geological photographs, for implementing the steps of the method for automatically generating geological descriptions based on geological photographs according to any one of claims 1 to 7, characterized in that: include: A preprocessing unit, used for performing image preprocessing on the geological field image to obtain a preprocessed image; A model building unit is used to build a geological image data model and train the geological image data model; the input of the model is an image containing geological attributes, and the output of the model is a description text describing geological information; A model application unit, configured to input the pre-processed image into the geological image data model and use the output of the model as an initial description text; The text generation unit is used to match the initial description text and the coordinate information of the geological site image with the geological knowledge base; after successful matching, the initial description text is supplemented with relevant information in the geological knowledge base to generate a final text entity.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the steps of the method for automatically generating geological descriptions based on geological photographs according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the steps of the method for automatically generating geological descriptions based on geological photographs according to any one of claims 1 to 7 when executed.