Drilling rock core image and test parameter automatic identification and recording method based on multi-modal large model
Through the multimodal large model combining image, numerical and text data for automatic core identification and cataloging, the problems of high labor intensity, artificial accuracy dependence, and insufficient multimodal fusion in traditional methods are solved, and efficient and accurate core cataloging and model adaptability enhancement is achieved.
Patent Information
- Application Number
- CN202510688364.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art has problems in drilling core catalogs with high labor intensity, personal experience dependence on personal experience, insufficient multimodal fusion, low degree of automation, and poor adaptability to model updates.
A multimodal large model-based method is adopted, combining drilling core images, numerical data and text data for feature extraction and fusion, using deep learning technology for automatic identification and cataloging, and supporting visual presentation and online/offline fine-tuning of the model.
It improves the accuracy and efficiency of core cataloging, enhances the adaptability of the model, supports on-site low-computing power deployment and cloud batch processing, and provides standardized interfaces to facilitate docking with geological software.
Smart Images

Figure CN120493075A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image recognition, data processing, deep learning and other technologies, and specifically to a method for automatically identifying and cataloging drill core images and test parameters based on a large modal model. Background Art
[0002] In numerous fields, including mining exploration, water conservancy and hydropower, oil and gas development, and civil engineering, coring is a key tool for gaining insight into underground geological structures. Documenting the mineralogy, fracture characteristics, and lithologic distribution of core samples provides a scientific basis for subsequent project planning and risk assessment. However, traditional documentation relies primarily on visual observation and manual measurement by geologists, who then record the results in tables or written reports. This method is labor-intensive, and its accuracy depends on individual experience, which can lead to subjective errors.
[0003] With technological advancements, some computer vision-based rock core image analysis methods have emerged. These methods use traditional image processing (such as threshold segmentation and edge detection) or simple convolutional neural network (CNN) models to identify features such as core fractures and rock particles, which to some extent reduces the burden of manual cataloging. However, they still have many limitations: (1) Most of these solutions only process a single image modality and cannot be combined with logging data or text description information. The recognition accuracy is insufficient when dealing with complex geological phenomena, and the recognition results lack a complete geological interpretation.
[0004] (2) There is a lack of real-time visual annotation and statistics. Although some software can perform image segmentation or simple recognition, it is unable to display the results in conjunction with geomechanical parameters, and it is also difficult to automatically output statistical analysis such as structural area ratio and crack distribution curve.
[0005] (3) Insufficient model updating and adaptability. Traditional methods lack an incremental learning mechanism after model deployment and cannot be quickly adjusted according to new core types or special structures that appear on site. Subsequent maintenance costs are high and the results are unstable.
[0006] In summary, existing technologies have obvious deficiencies in multimodal fusion, automatic labeling and visualization, and model update adaptability, making it difficult to meet the needs of mining exploration and other fields for efficient, accurate, and automated core cataloging. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to propose a method for automatically identifying and cataloging drill core images and test parameters based on a large modal model, so as to solve the problem that the existing technology is difficult to meet the requirements of efficient, accurate and automated core cataloging.
[0008] The technical solution adopted by the present invention to solve the above technical problems is: A method for automatically identifying and cataloging drill core images and test parameters based on a large modal model comprises the following steps: S1. Collect image data, numerical data and text data of drill core; S2. Preprocess the collected data; S3. Using the preprocessed image data and associated numerical and textual data as input, the trained multimodal large model is used for analysis and processing to obtain core structure detection results, classification results, and explanatory text; S4. Visualize the analysis results and allow geologists to manually revise the analysis results.
[0009] Furthermore, in step S1, the method of collecting image data, numerical data and text data of the drill core includes: Drill core image data acquisition: Capture the core box and local details of the core, and record the shooting time, borehole ID, core box number, burial depth range and a brief description of the lithology; Numerical data acquisition: collecting core-related logging data and rock mechanics parameters; Text data collection: input geological descriptions and historical geological reports compiled by geologists.
[0010] Furthermore, in step S2, the collected data is preprocessed, including: The image data is normalized and format converted, the numerical data is cleaned and standardized, and the text data is cleaned and segmented. Using the borehole ID and burial depth range as the primary key, the corresponding relationship between the image data, numerical data, and text data is established and stored.
[0011] Furthermore, in step S3, the multimodal large model includes: Input layer, used to receive preprocessed image, numerical and text data; The encoding layer includes an image encoder, a value encoder, and a text encoder, which are used to extract feature vectors of images, values, and texts respectively; The fusion layer is used to interactively fuse the feature vectors of the three modalities: image, numerical value, and text, through a multi-head attention mechanism; The output layer includes a detection head, a classification head, and a text generation module, which are used to output the detection results, classification results, and explanatory text of the core structure respectively.
[0012] Furthermore, the training method of the multimodal large model includes: Construct a training dataset, where the data samples include core images, numerical data, and text data; Preprocess the data samples; Input the pre-processed core images and associated numerical and textual data into the multimodal macromodel; Obtain core structure detection results, classification results and explanatory text based on the multimodal large model; Calculating a loss function, including the loss function including a weighted sum of target detection loss, classification loss, and text generation loss; Update the model parameters according to the loss function.
[0013] Furthermore, the target detection loss adopts Focal Loss and Bounding Box regression loss; the classification loss adopts Cross-Entropy loss; and the text generation loss adopts sequence-to-sequence Cross-Entropy loss.
[0014] Furthermore, the deployment and update methods of the multimodal large model include: The trained multimodal large model is converted into ONNX format, quantified, and then deployed on-site terminals or in the cloud. By collecting revision data from geologists as incremental training data, it can regularly guide the online or offline fine-tuning of the model.
[0015] Furthermore, the deployment and update method of the multimodal large model also includes: Assign a new model version number to each update and record the source of the dataset, revision amount, and training parameters; if an anomaly is found in the new version model, quickly roll back to the previous version.
[0016] Furthermore, in step S4, the analysis results are visualized, including: Overlay the annotation results on the image to display the detected core structure location, category and attribute information; Allow geologists to manually modify the identification results and record the revision information; Export recognition results to JSON, XML, or GeoTIFF file formats, and generate reports in TXT, DOCX, PDF, or HTML formats.
[0017] Furthermore, in step S4, the visual display of the analysis results further includes: When the marked numerical parameters conflict with prior knowledge, a reminder or an abnormality mark will be issued to prompt geologists to conduct a secondary check.
[0018] The beneficial effects of the present invention are: (1) Improve recognition accuracy: The present invention combines image, numerical and text data, and integrates multiple information sources for analysis, avoiding the limitations of single modality data and improving the accuracy and reliability of recognition.
[0019] (2) Improve cataloging efficiency: The present invention automatically identifies and annotates core structures through a large multimodal model, reducing the workload of manual cataloging and significantly improving cataloging efficiency. The system can display the identification results in real time, including the annotation and statistical information of structures such as fractures, fragment segments, and columnar segments, making it easier for geologists to quickly understand the core characteristics.
[0020] (3) Improve the adaptability of the model: Through online / offline fine-tuning and incremental learning mechanisms, the model can be continuously optimized based on new data and revision results of geologists, adapting to different geological scenarios and core types, thereby improving the adaptability of the model.
[0021] (4) Support on-site low-computing power deployment and cloud batch processing: This paper uses quantization and ONNX Runtime technology to deploy the model on-site terminals, realizing real-time inference of the model on ordinary CPUs. The model can also be deployed on the cloud, so as to centrally process large amounts of core data and conduct large-scale training or batch analysis in the cloud, taking into account both the needs of engineering sites and batch application of scientific research.
[0022] (5) Provide standardized and modular interfaces: The present invention can provide multiple result export formats such as JSON, XML, GeoTIFF, etc., which is easy to connect with other geological software or engineering platforms; it can be integrated with geological databases or knowledge bases to build a complete automated cataloging and decision support system. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is the technical roadmap of the solution of the present invention.
[0024] Figure 2 This is a flow chart of the method for automatically identifying and cataloging drill core images and test parameters in the present invention. DETAILED DESCRIPTION
[0025] This invention aims to provide a method for automatically identifying and cataloging drill core images and test parameters based on a modal macromodel, addressing the challenges of existing technologies in meeting the requirements for efficient, accurate, and automated core cataloging. Its core concept is to utilize multimodal macromodels from deep learning technology to comprehensively process drill core images, logging data (such as density and porosity), rock mechanical parameters (such as compressive strength and elastic modulus), and textual descriptions such as geological reports and field notes acquired during geological exploration. Feature extraction and interactive fusion of these diverse modal data enable automatic identification and annotation of core structures (such as fractures, fragmented segments, and columnar segments), and generate reports containing test results, classification results, and explanatory text. The system also supports visualization, manual revision, and online / offline model fine-tuning and continuous learning to adapt to diverse geological scenarios and improve model accuracy and adaptability. This solution can provide a scientific basis and technical support for geological research and engineering decision-making in fields such as mining exploration, water conservancy and hydropower, oil and gas development, and civil engineering.
[0026] In terms of specific implementation, the technical roadmap of the present invention can be found in Figure 1 , which will be explained in detail below.
[0027] 1. Multi-source data collection In this process, the multi-source data collected mainly include core image data, numerical data and text data.
[0028] 1. Core image data acquisition: (1) Shooting environment: When shooting indoors or outdoors, the lighting should be uniform and the contrast should be moderate to reduce the noise caused by direct strong light or large shadows.
[0029] (2) Camera or video device: Choose a digital camera, industrial camera, or high-resolution mobile phone with a high resolution (e.g., 20 megapixels or above). If conditions permit, use a fixed bracket to ensure consistent lens angle and distance.
[0030] (3) Shooting method: Shoot each box of cores in depth order, ensuring that one photo completely covers the entire box of cores. A reference ruler (such as a scale) should be placed in the image to facilitate subsequent image measurement and size calibration.
[0031] (4) Image recording information: metadata such as shooting time, borehole number, core box number, burial depth range, and brief description of lithology are collected and registered using a handwriting tablet or spreadsheet, and then associated in subsequent file naming or databases.
[0032] (5) Photographing local details such as cracks and layers: Zoom in on key areas such as cracks, pores, and fine-grained mineral distribution, using a macro lens or high-power digital microscope to capture detailed textures. After photographing, record the corresponding depth and core number. Maintaining a consistent photographic style within a single project facilitates subsequent large-scale data analysis and model training.
[0033] 2. Multi-source numerical and text data collection: (1) Numerical experimental parameters: Well logging data: density, porosity, water content, sonic time difference, etc.
[0034] Rock mechanics parameters: compressive strength, tensile strength, elastic modulus, Poisson's ratio, etc.
[0035] Mineral composition detection parameters: such as mineral content in X-ray diffraction (XRD) results.
[0036] Numerical parameters are recorded in Excel, CSV or professional logging data formats, and information such as the borehole number, depth section, sampling location, etc. are clearly defined to achieve one-to-one correspondence with the core image.
[0037] (2) Text description and geological labels: Field geologists create detailed text descriptions (including rock formation attribution, rock names, and descriptions of special structures) during the cataloging process, combined with historical geological reports. These descriptions are typically stored in Word documents or databases to ensure that the different descriptions correspond to the captured core images or depth segments.
[0038] During the multi-source data acquisition process, after photographing the core box and collecting numerical / text parameters, these data need to be associated through a database or file structure: for example, a storage structure of borehole ID → depth segment → corresponding core image file name + corresponding logging / experimental parameters + text description is used. This allows all information to be in the same coordinate or identification system, laying the foundation for subsequent multimodal data processing.
[0039] 2. Multimodal Preprocessing During this process, the collected image data, numerical data, and text data are preprocessed separately to reduce noise interference and improve the accuracy and stability of subsequent model training and reasoning.
[0040] 1. Image preprocessing: (1) Size normalization: Due to differences in shooting distance, resolution, and core length, images can be scaled according to preset sizes (512×512, 1024×1024, etc.).
[0041] (2) Color correction: If the core color tone is important for analysis, simple white balancing or color correction can be performed.
[0042] 2. Numerical and text data preprocessing: (1) Numerical cleaning: Eliminate outliers from logging or laboratory test data, convert units, and unify dimensions based on project needs. If continuous logging data or segmented data exists, segment or aggregate it to ensure it aligns with the depth segments corresponding to the core images.
[0043] (2) Text segmentation: Normalize or synonymize geological terms. For Chinese descriptions, use a word segmentation tool (such as Jieba) and add a custom dictionary of geological terms to eliminate content unrelated to the core structure to prevent interference in subsequent model analysis.
[0044] Based on this preprocessed data, a database or unified index is established, using (drillhole ID, depth start and end, box number) as the primary key to store image preprocessing results, numerical parameters, and text descriptions. During candidate model training and inference, this primary key can be used to simultaneously retrieve images, numerical values, and text, meeting the requirements of multimodal fusion input.
[0045] 3. Feature Encoding In this process, the encoding layer in the multimodal large model is used to extract the feature vectors of preprocessed image data, numerical parameters and text descriptions respectively.
[0046] The multimodal large model includes: input layer, encoding layer, fusion layer and output layer. The specific implementation of each layer is as follows: 1. Input layer: This layer is used to receive preprocessed image, numerical, and text data as multimodal input.
[0047] (1) Image modality input: preprocessed core photos.
[0048] (2) Numerical modal input: pre-processed continuous or discrete data such as density, porosity, compressive strength, acoustic time difference, mineral composition ratio, etc.
[0049] (3) Text modal input: pre-processed geological descriptions, layer names, historical records, and other word segmentation results.
[0050] 2. Coding layer: This layer includes an image encoder, a numerical encoder, and a text encoder to extract feature vectors of different modalities respectively.
[0051] (1) Image encoder: The Vision Transformer (ViT) architecture is used to extract high-dimensional features of core images.
[0052] (2) Numerical encoder: converts logging / mechanical parameters into vectors through embedding.
[0053] (3) Text encoder: Use BERT or a similar language model to semantically encode the word segmentation results to obtain a context vector representation.
[0054] 3. Fusion layer: This layer uses a multi-head attention structure to interactively fuse image, numerical, and text features, learning their relationships for drill core identification and attribute inference. When encountering specialized geological terms or numerical ranges, prior knowledge can be embedded to enhance the model's understanding of this specialized information.
[0055] 4. Output layer: This layer consists of three parts: detection head, classification head and text generation module.
[0056] (1) Detection head: Detects targets of different types of core segments in the core image, thereby marking the locations of fragment segments, cracks, columnar segments, etc.
[0057] (2) Classification head: Outputs the corresponding results according to the geological classification task (lithology category, fracture grade, etc.) or numerical prediction (porosity, strength, etc.) requirements.
[0058] (3) Text generation module: Generates text descriptions of core features and outputs structured or natural language descriptions, such as “This section of core is mainly sandstone, with moderate crack development and moderate weathering.
[0059] 4. Fusion Training In this process, multi-task learning is used to continuously optimize model parameters in combination with loss calculation to obtain a trained multimodal large model.
[0060] 1. Multimodal input assembly: For each sample, an image tensor, a numerical vector, and a text sequence are simultaneously prepared, and any necessary padding or segmentation is performed. The organized multimodal data is fed into the image encoder, numerical encoder, and text encoder, which each output a sequence of feature vectors.
[0061] 2. Feature fusion and loss calculation: After the feature vector sequences of the three modalities interact in the fusion layer, the fusion results are passed to the downstream task head (object detection, classification, text generation) to calculate the corresponding loss function. The loss function includes the following types of detection functions: (1) Target detection: Use Focal Loss + Bounding Box regression loss; Focal Loss: ; in, Represents the predicted probability of the true category (if the true category is 1, then ; If the true category is 0, then ); is the balance coefficient, which is used to adjust the imbalance between positive and negative samples; It is a focusing parameter used to reduce the loss weight of easy-to-classify samples and pay more attention to difficult-to-classify or small object samples.
[0062] Bounding Box Regression Loss: ; in, is the coordinate difference between the predicted bounding box and the true box.
[0063] (2) Classification: Use Cross-Entropy loss;
[0064] in, is the number of categories, used for multi-classification; is the one-hot encoding representation corresponding to the true label (or true word); is the predicted probability output by the model.
[0065] (3) Text generation: Use sequence-to-sequence Cross-Entropy loss;
[0066] Among them, "text" refers to the contextual information output from the Encoder (such as multimodal features); is the overall trainable parameter; The target text is Position words, Represents a word that has been generated before.
[0067] For multi-task scenarios that require simultaneous output of image target detection and text description, the weighted multi-task loss function mentioned above is used to comprehensively optimize the losses of each task.
[0068] 3. Optimization and hyperparameter setting: Use optimization algorithms such as AdamW and SGD, set the learning rate, batch size, number of training rounds, etc.; in large model scenarios, use gradient accumulation and mixed precision training (FP16 / BF16) to balance video memory overhead and training speed; regularly save model checkpoints, perform validation set evaluations, observe indicators such as recognition accuracy or text generation effect, and tune hyperparameters or model structure based on the results.
[0069] 5. Engineering Application In this process, the trained multimodal large model is deployed and applied in actual scenarios.
[0070] 1. Model deployment: Convert the trained multimodal model to the ONNX (Open Neural Network Exchange) format. During the conversion process, ensure that all components of the model architecture (convolution, attention, activation functions, etc.) are compatible with the ONNX specification and adapt to possible dynamic dimensions or unsupported operators.
[0071] To enable real-time inference on CPUs or low-computing devices, ONNX models are quantized to INT8, FP16, and other precisions, significantly reducing model size, memory usage, and improving inference speed. The model can be deployed on-site as needed, enabling routine intelligent analysis and processing of core images, including real-time cataloging and 3D visualization, on a standard CPU workstation or laptop. Alternatively, the model can be deployed in the cloud, centrally processing large amounts of core data and performing large-scale training or batch analysis.
[0072] 2. Model iterative update: Due to geological complexity, the core structure and feature distribution are also diverse. In order to improve the model's recognition accuracy and adaptability to different geological scenarios, the model needs to be continuously iterated and updated.
[0073] This solution incorporates a manual correction feature for model output (e.g., correcting crack locations, adding missing fragments, etc.). The system records these corrections in real time and stores them in a revision database or incremental annotation file. Each revision entry includes the original image ID, the model prediction result, the final result after user correction, the time of correction, the user ID, and a description of the reason.
[0074] By regularly merging the collected revised data with the original training data, a new incremental data set is formed; this incremental data can reflect the latest field sample distribution or newly emerging structural types (such as new features discovered in certain special strata), guiding model updates.
[0075] Model updates can be performed online or offline. First, online fine-tuning involves receiving revised data in the backend or cloud, initiating small-batch training to incorporate newly annotated images and numerical / textual information into the model for continuous learning. Second, offline fine-tuning involves performing full training or periodic retraining within a fixed period to enhance the model's ability to adapt to new scenarios. If performance indicators are evaluated to be superior to the previous version, the on-site model will be replaced.
[0076] 3. Model version management and rollback: A new model version number is assigned to each update, and the revision amount and training parameters are recorded. Once an anomaly is found in the new version, it can be quickly rolled back to the previous version to ensure stable operation of the system on site.
[0077] The above is an explanation of the technical route of the present invention. In actual scene applications, the method for automatically identifying and cataloging drill core images and test parameters based on the modal large model provided by the present invention is implemented as follows: Figure 2 As shown, the following steps are included: S1. Collect image data, numerical data and text data of drill core; In this step, the image data, numerical data and text data of the drill core are obtained according to the "multi-source data acquisition" method in the above technical route.
[0078] S2. Preprocess the collected data; In this step, the image data, numerical data and text data of the core are preprocessed according to the "multimodal preprocessing" method in the above technical route.
[0079] S3. Use the trained multimodal large model for analysis and processing; In this step, the preprocessed image data and associated numerical data and text data are used as input. The multimodal large model extracts feature vectors of different modalities and performs feature fusion, and analyzes the contextual information provided by the fused features to obtain the prediction results.
[0080] S4. Visualize the output results: In this step, the output results include image recognition results and explanatory output of numerical pre-text information.
[0081] 1. Image recognition output: Based on the detection output, the coordinate range or outline of each target (such as a "fragment" or "column") is obtained. A common visualization method for this outline is a rectangular box, which encircles the region of interest on the image and annotates the classification label nearby. Furthermore, the detected target set can be counted, sized, and evaluated for distribution, serving as input for subsequent three-rate or structural indicator analysis. For example, the area or length distribution of all "fragmented" targets can be summarized to calculate the proportion of fragments in the entire core section.
[0082] 2. Explanatory output of numerical and textual information: (1) Mechanical parameter mapping: When a structural area is detected, it can be linked with numerical modal information (such as compressive strength and porosity) to annotate the structure with strength level or porosity based on thresholds or intervals. For example, if the compressive strength of a columnar core is high, the model can automatically output the text "The strength level of this section of core is relatively high" and distinguish the area with color (for example, red represents high strength).
[0083] (2) Text description generation: Integrating text generation capabilities within the multimodal large model automatically outputs "core structure descriptions" or "geological interpretations," providing reference for compiling reports or making decisions. For example, "This section of core rock is primarily sandstone, with two cracks, each approximately 2 mm wide, and mechanical tests showing a compressive strength of 80 MPa."
[0084] During the visualization process, when numerical parameters conflict with prior knowledge (such as abnormally high porosity or obvious deviation of mechanical data), the system automatically issues a reminder or marks an anomaly to assist geologists in conducting secondary verification.
[0085] In addition, the system also provides a means correction function: if misidentification or omission occurs, geologists can make manual modifications on the interface; the system records the correction results in the revision database or incremental annotation file to provide a basis for subsequent model fine-tuning.
[0086] Finally, it should be noted that the above embodiments are merely preferred embodiments and are not intended to limit the present invention. It should be noted that those skilled in the art would be able to make modifications, equivalent substitutions, and improvements without departing from the spirit and scope of the present invention and the claims, all of which should be included within the scope of protection of the present invention.
Claims
1. A method for automatically identifying and cataloging drill core images and test parameters based on a large modal model, characterized in that: The following steps are involved: S1. Collect image data, numerical data and text data of drill core; S2. Preprocess the collected data; S3. Using the preprocessed image data and associated numerical and textual data as input, the trained multimodal large model is used for analysis and processing to obtain core structure detection results, classification results, and explanatory text; S4. Visualize the analysis results and allow geologists to manually revise the analysis results.
2. The method for automatically identifying and cataloging drill core images and test parameters based on a large modal model according to claim 1, characterized in that: In step S1, the method for collecting image data, numerical data and text data of the drill core includes: Drill core image data acquisition: Capture the core box and local details of the core, and record the shooting time, borehole ID, core box number, burial depth range and a brief description of the lithology; Numerical data acquisition: collecting core-related logging data and rock mechanics parameters; Text data collection: input geological descriptions and historical geological reports compiled by geologists.
3. The method for automatically identifying and cataloging drill core images and test parameters based on a modal large model according to claim 1, characterized in that: In step S2, the collected data is preprocessed, including: The image data is normalized and format converted, the numerical data is cleaned and standardized, and the text data is cleaned and segmented. Using the borehole ID and burial depth range as the primary key, the corresponding relationship between the image data, numerical data, and text data is established and stored.
4. The method for automatically identifying and cataloging drill core images and test parameters based on a large modal model according to claim 1, wherein: In step S3, the multimodal large model includes: Input layer, used to receive preprocessed image, numerical and text data; The encoding layer includes an image encoder, a value encoder, and a text encoder, which are used to extract feature vectors of images, values, and texts respectively; The fusion layer is used to interactively fuse the feature vectors of the three modalities: image, numerical value, and text, through a multi-head attention mechanism; The output layer includes a detection head, a classification head, and a text generation module, which are used to output the detection results, classification results, and explanatory text of the core structure respectively.
5. The method for automatically identifying and cataloging drill core images and test parameters based on a large modal model according to claim 4, characterized in that: The training method of the multimodal large model includes: Construct a training dataset, where the data samples include core images, numerical data, and text data; Preprocess the data samples; Input the pre-processed core images and associated numerical and textual data into the multimodal macromodel; Obtain core structure detection results, classification results and explanatory text based on the multimodal large model; Calculating a loss function, including the loss function including a weighted sum of target detection loss, classification loss, and text generation loss; Update the model parameters according to the loss function.
6. The method for automatically identifying and cataloging drill core images and test parameters based on a large modal model according to claim 5, characterized in that: The target detection loss adopts Focal Loss and Bounding Box regression loss; the classification loss adopts Cross-Entropy loss; the text generation loss adopts sequence-to-sequence Cross-Entropy loss.
7. The method for automatically identifying and cataloging drill core images and test parameters based on a modal large model according to claim 4, characterized in that: The deployment and update methods of the multimodal large model include: The trained multimodal large model is converted into ONNX format, quantified, and then deployed on-site terminals or in the cloud. By collecting revision data from geologists as incremental training data, it can regularly guide the online or offline fine-tuning of the model.
8. The method for automatically identifying and cataloging drill core images and test parameters based on a modal large model according to claim 7, characterized in that: The deployment and update method of the multimodal large model also includes: Assign a new model version number to each update and record the source of the dataset, revision amount, and training parameters; if an anomaly is found in the new version model, quickly roll back to the previous version.
9. The method for automatically identifying and cataloging drill core images and test parameters based on a modal large model according to claim 1, wherein: In step S4, the analysis results are visualized, including: Overlay the annotation results on the image to display the detected core structure location, category and attribute information; Allow geologists to manually modify the identification results and record the revision information; Export recognition results to JSON, XML, or GeoTIFF file formats, and generate reports in TXT, DOCX, PDF, or HTML formats.
10. The method for automatically identifying and cataloging drill core images and test parameters based on a large modal model according to claim 9, characterized in that: In step S4, the visual display of the analysis results further includes: When the marked numerical parameters conflict with prior knowledge, a reminder or an abnormality mark will be issued to prompt geologists to conduct a secondary check.
Citation Information
Patent Citations
Rock core data classification and layering method
CN118747201A
Rock-soil core sample information recording system based on spectrum and image recognition
CN119540595A
Multimodal Learning from Structured and Unstructured Data
US20240386321A1