Multi-modal large model-based borehole histogram information identification and extraction method
By constructing a borehole histogram information recognition method based on a multimodal large model, intelligent recognition and automatic extraction of borehole histogram information are achieved, solving the problems of low efficiency and insufficient accuracy in traditional methods, and improving the efficiency and data quality of geological database construction.
Patent Information
- Application Number
- CN202510761543.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional manual identification and manual entry of borehole histogram information are inefficient, time-consuming and easily affected by subjective factors. Existing technologies are insufficient in information recognition accuracy and universality, and cannot meet the needs of geological database construction.
A drilling histogram information recognition method based on a multimodal large model is constructed, including image preprocessing, multimodal large model parameter tuning, prompt fine-tuning, data cleaning and user interactive verification, to achieve intelligent recognition and automatic extraction of drilling histogram information.
It improves the flexibility and adaptability of drilling histogram information processing, reduces labor costs, improves work efficiency, ensures the accuracy and completeness of data, and meets the needs of geological database construction.
Smart Images

Figure CN120635934A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of geological information processing, and more particularly to a method for identifying and extracting borehole histogram information based on a multimodal large model. Background Art
[0002] With the rapid development of infrastructure construction, survey and design institutes have accumulated vast amounts of geological data through long-term geological surveys. This data contains rich engineering geological information and is a crucial resource for supporting engineering design and decision-making. As the core carrier of geological survey results, borehole histograms intuitively present key information such as stratum lithology, thickness, and geological structure, and serve as fundamental data for building geological databases. However, traditional borehole histograms often exist in the form of paper drawings or unstructured electronic documents. The geological information they contain is a mixture of text, symbols, and graphics, with complex spatial distribution and a dense concentration of specialized terminology, making information extraction difficult.
[0003] Currently, the industry still primarily processes drillhole histogram information through manual identification and data entry. This method is not only inefficient and time-consuming, but also subject to significant subjective influences, prone to data errors, and unable to meet the requirements for data integrity, accuracy, and standardization in geological database construction. Although some studies have attempted to use technologies such as image recognition and natural language processing to automatically extract information, due to the diverse graphic symbols in drillhole histograms, non-standardized annotation layout, and regional differences in geological terminology, existing technologies still lack accuracy and universality in information recognition, making them unable to effectively adapt to the mapping standards and data characteristics of different survey units. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for identifying and extracting drilling histogram information based on a multimodal large model, which realizes the intelligent identification and automatic extraction of drilling histogram information, can adapt to drilling histograms of different formats and complexities, improves the flexibility and adaptability of the system, greatly reduces labor costs, and improves work efficiency.
[0005] The technical solution adopted by the present invention to solve the technical problem is to construct a method for identifying and extracting drilling histogram information based on a multimodal large model, comprising the following steps:
[0006] S1, performing image enhancement preprocessing on the input drilling histogram image;
[0007] S2. Construct the prompt fine-tuning corpus of the multimodal large model and convert the information in the drill bar chart into a structured prompt template;
[0008] S3: Parameter tuning of the multimodal model and prompt fine-tuning: using the optimized prompt words to fine-tune the multimodal model;
[0009] S4. Using the fine-tuned multimodal large model to identify and extract drilling histogram information, performing data cleaning on the information output by the multimodal large model, and outputting data in a specified format;
[0010] S5. Display the cleaned data information, and the user checks and modifies it, and stores the finalized bar chart information in the database.
[0011] According to the above solution, in step S1, the method for preprocessing the input drilling histogram image for image enhancement includes the following steps:
[0012] S101, reading an input drilling histogram image file to obtain image data;
[0013] S102 , performing noise reduction processing on the image to remove random noise in the image, enhance the recognizability of the drilling histogram, and output the drilling histogram after noise reduction.
[0014] According to the above scheme, in step S2, the method for constructing the prompt fine-tuning corpus of the multimodal large model includes the following steps:
[0015] S201, analyzing the type of information to be identified in the drill histogram, and determining the location area and meaning of the information according to the drill histogram compilation rules;
[0016] S202, converting the information into a standardized prompt word format; designing a corresponding prompt word template according to the determined information type and location area, extracting text and symbol information from the image, and converting it into a standardized prompt word format.
[0017] According to the above scheme, the information types include layer bottom depth, layer thickness, layer bottom elevation, rock and soil name and characteristic information.
[0018] According to the above scheme, in step S3, the method of fine-tuning the multimodal large model using the optimized prompt words includes the following steps:
[0019] S301. Set the temperature (generated randomness) of the multimodal large model to 0.3, the TopP (cumulative probability) to 0.8, and the frequency penalty (repeated sentence penalty) to 1.1.
[0020] S302: Use the prompt word in step S2 as the system prompt word of the large model.
[0021] According to the above scheme, the multimodal large model includes Qwen2.5-VL-72B-Instruct and Doubao-1.5-thinking-pro visual multimodal large model.
[0022] According to the above solution, in step S4, the method for performing data cleaning on the information output by the multimodal large model includes the following steps:
[0023] S401: Calling the multimodal large model API interface, and using the drill histogram and the system prompt word system prompt after the data enhancement processing in step S1 as input information of the multimodal large model;
[0024] S402: Perform data cleaning on the markdown format information output by the multimodal large model.
[0025] According to the above solution, in step S402, data cleaning includes the following:
[0026] Input preprocessing: Use replace( / \\n / g,'\n') to process escaped newline characters, and split('\n') to split the input content into lines;
[0027] Table row identification: Use the regular expression / ^\s*\|(.+)\|\s*$ / to accurately match table rows, filter non-table rows, and retain valid data;
[0028] Header parsing: extract the first row as the header, use split('|') to split the columns and remove spaces;
[0029] Separator line detection: Use the isSeparatorLine function to identify separator lines of the type |---|:---|:---:|---:| and dynamically determine the starting row of the data (skip the separator line).
[0030] Data conversion: Process cell line breaks, build key-value pair objects, automatically handle missing columns, and retain original line break information.
[0031] According to the above scheme, in step S5, the method of displaying the cleaned data information, allowing the user to check and modify the information, and storing the finalized histogram information in the database includes the following steps:
[0032] S501. Generate a visualization table based on the structured data after data cleaning, and display the layer bottom depth, layer thickness, layer bottom elevation, rock and soil name and characteristic information of the drill hole histogram through an interactive interface;
[0033] S502: Provide a user interaction function to support the user to manually correct or supplement the recognition results;
[0034] S503: Store the drilling histogram information confirmed by the user into a drilling histogram information database.
[0035] According to the above scheme, in step S502, an editable text box is set for each field in the interactive interface, and the user can directly modify the value in the text box, so that the user can adjust the recognition result according to the actual drilling histogram information to ensure the accuracy of the data.
[0036] The implementation of the present invention's method for identifying and extracting drilling histogram information based on a multimodal large model has the following beneficial effects:
[0037] By introducing a multimodal large model, the present invention realizes the intelligent recognition and automatic extraction of borehole histogram information, can adapt to borehole histograms of different formats and complexities, improves the flexibility and adaptability of the system, greatly reduces labor costs, improves work efficiency, and solves the problem of time-consuming and labor-intensive traditional manual extraction of borehole histogram information, which can meet the urgent needs of geological database construction and intelligent application of geological data. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0039] Figure 1 It is a flow chart of the method for identifying and extracting drilling histogram information based on a multimodal large model of the present invention;
[0040] Figure 2 It is a schematic diagram of the multimodal large model of the present invention identifying key information in the drilling histogram after data cleaning. DETAILED DESCRIPTION
[0041] In order to have a clearer understanding of the technical features, purposes and effects of the present invention, specific embodiments of the present invention are now described in detail with reference to the accompanying drawings.
[0042] like Figure 1-2 As shown, the method for identifying and extracting drilling histogram information based on a multimodal large model of the present invention includes the following steps:
[0043] S1. Image preprocessing: image enhancement of the input drilling histogram image to facilitate subsequent reading of the multimodal large model. Specifically, the following steps are included:
[0044] S101 , reading an input drilling histogram image file to obtain image data.
[0045] S102 , performing noise reduction processing on the image to remove random noise in the image, enhance the recognizability of the drilling histogram, and output the drilling histogram after noise reduction.
[0046] S2. Construct a multimodal large model prompt fine-tuning corpus and convert the information in the drill bar chart into a structured prompt template. This specifically includes the following steps:
[0047] S201. Analyze the types of information to be identified in the drill hole histogram, including layer bottom depth, layer thickness, layer bottom elevation, rock and soil name, and characteristic information. Determine the location and meaning of each type of information based on the drill hole histogram compilation rules.
[0048] S202: Convert the information into a standardized prompt word format. According to the information type and location area, a corresponding prompt word template is designed to extract the text and symbol information in the image and convert it into a standardized prompt word format.
[0049] The prompt word is:
[0050] #Geological data extraction instructions
[0051] Input Handling
[0052] - If there is a double table structure, it is processed strictly in the order of [left first, then right]
[0053] - Automatically identify table boundaries and cell contents in images
[0054] **Data Extraction Rules**
[0055] - Column order is forced to be: Layer bottom depth (m) | Layer thickness (m) | Layer bottom elevation (m) | Geotechnical name | Characteristics
[0056] -Numerical processing:
[0057] - Layer bottom depth: Keep negative format (such as -1.5)
[0058] -Layer thickness / bottom elevation: Keep positive format
[0059] -Text separation:
[0060] - When "Rock and Soil Name" and "Characteristics" are in the same cell, separate them by the last space (e.g. "Clay Hard" → Rock and Soil Name: Clay | Characteristics: Hard)
[0061] - Leave blank if no feature
[0062] **Output Specification**
[0063] -Strictly generate standard Markdown tables
[0064] - Preserve original data precision (no rounding)
[0065] - No explanatory text allowed
[0066] - Ignore table contents not related to the target column
[0067] Sample output:
[0068] |Layer bottom depth (m)|Layer thickness (m)|Layer bottom elevation (m)|Rock and soil name|Characteristics|
[0069] |---|---|---|---|---|
[0070] |-1.5|1.5|5.2|Silty Clay|Plastic|
[0071] |-3.0|1.5|3.7|Muddy soil|Plastic|
[0072] ```.
[0073] S3, multimodal model parameter tuning and prompt fine-tuning, using the optimized prompt words to fine-tune the multimodal model. Specifically, the following steps are included:
[0074] S301. Set the temperature (randomness of generation) of the multimodal large model to 0.3, the Top P (cumulative probability) to 0.8, and the frequency penalty (penalty for repeated sentences) to 1.1. Keep all other settings as default.
[0075] S302: Use the prompt word in step S2 as the system prompt word of the large model.
[0076] Multimodal large models include Qwen2.5-VL-72B-Instruct, Doubao-1.5-thinking-pro, and a series of other visual multimodal large models.
[0077] S4. Using the fine-tuned multimodal large model to identify and extract drilling histogram information, clean the information output by the multimodal large model, and output data in a specified format. This specifically includes the following steps:
[0078] S401: Call the multimodal large model API interface and use the drill histogram after data enhancement processing in step S1 and the system prompt word system prompt as input information of the multimodal large model.
[0079] S402: Perform data cleaning on the markdown format information output by the multimodal large model based on the JavaScript language.
[0080] Data cleaning includes the following:
[0081] Input preprocessing: Use replace( / \\n / g,'\n') to process escaped newline characters, and split('\n') to split the input content into lines;
[0082] Table row identification: Use the regular expression / ^\s*\|(.+)\|\s*$ / to accurately match table rows, filter non-table rows, and retain valid data;
[0083] Header parsing: extract the first row as the header, use split('|') to split the columns and remove spaces;
[0084] Separator line detection: Use the isSeparatorLine function to identify separator lines of the type |---|:---|:---:|---:| and dynamically determine the starting row of the data (skip the separator line).
[0085] Data conversion: Process cell line breaks, build key-value pair objects, automatically handle missing columns, and retain original line break information.
[0086] S5: Display the cleaned data, and the user will check and modify it, and then store the finalized bar chart information in the database. This includes the following steps:
[0087] S501. Generate a visual table based on the structured data after data cleaning, and display information such as the layer bottom depth, layer thickness, layer bottom elevation, rock and soil name, and characteristic information of the drill hole histogram through an interactive interface.
[0088] S502. Provide user interaction functions to support users to manually correct or supplement the recognition results. In the interactive interface, set an editable text box for each field, and users can directly modify the value in the text box, so that users can adjust the recognition results according to the actual drilling histogram information to ensure data accuracy.
[0089] S503. The drill hole histogram information confirmed by the user is stored in the drill hole histogram information database. After the user completes the verification and modification of the recognition results, the system stores the data including the layer bottom depth, layer thickness, layer bottom elevation, rock and soil name, and characteristic information in the drill hole histogram information database according to the database table structure.
[0090] Example
[0091] The method for identifying and extracting drilling histogram information based on a multimodal large model of the present invention comprises the following steps:
[0092] S1. Image preprocessing: input drilling histogram image, such as Figure 2 As shown, image enhancement is performed to facilitate subsequent reading of the multimodal large model. The specific steps include:
[0093] S101. Read the JPG / PNG format image of the drilling histogram through the OpenCV library and convert it into RGB three-channel matrix data.
[0094] S102 , performing noise reduction processing on the image, using a non-local means noise reduction algorithm to remove random noise and moiré patterns in the scanned image, enhancing the contrast of the text area in the histogram, and outputting an optimized image.
[0095] S2. Construct a multimodal large model prompt fine-tuning corpus and convert the information in the drill bar chart into a structured prompt template. This specifically includes the following steps:
[0096] S201. Analyze the types of information to be identified in the drill hole histogram, including layer bottom depth, layer thickness, layer bottom elevation, rock and soil name, and characteristic information. Determine the location and meaning of each type of information based on the drill hole histogram compilation rules.
[0097] S202: Convert the information into a standardized prompt word format. According to the information type and location area, a corresponding prompt word template is designed to extract the text and symbol information in the image and convert it into a standardized prompt word format.
[0098] S3, multimodal model parameter tuning and prompt fine-tuning, using the optimized prompt words to fine-tune the multimodal model. Specifically, the following steps are included:
[0099] S301. Configure the Doubao-1.5-thinking-pro model parameters: temperature = 0.6, Top_p = 0.7, frequency penalty = 0.0, max_tokens = 4096.
[0100] S302: Use the prompt word in step S2 as the system prompt word of the large model to enhance the multimodal large model's ability to understand key information in the drilling histogram.
[0101] S4. Using the fine-tuned multimodal large model to identify and extract drilling histogram information, clean the information output by the multimodal large model, and output data in a specified format. This specifically includes the following steps:
[0102] S401. Call the multimodal large model API interface, use the drilling histogram after data enhancement processing in step S1 and the system prompt word system prompt as input information of the interface, and finally return the histogram information recognized by the multimodal large model.
[0103] S402: Data cleaning is performed on the markdown format information output by the multimodal large model, which is performed in the following steps:
[0104] Input preprocessing: Use replace( / \\n / g,'\n') to process escaped newline characters, and split('\n') to split the input content into lines;
[0105] Table row identification: Use the regular expression / ^\s*\|(.+)\|\s*$ / to accurately match table rows, filter non-table rows, and retain valid data;
[0106] Header parsing: extract the first row as the header, use split('|') to split the columns and remove spaces;
[0107] Separator line detection: Use the isSeparatorLine function to identify separator lines of the type |---|:---|:---:|---:| and dynamically determine the starting row of the data (skip the separator line).
[0108] Data conversion: process cell line breaks, build key-value pair objects, automatically handle column missing, and retain original line break information. The output example after data cleaning is as follows Figure 2 shown.
[0109] S5: Display the cleaned data, and the user will check and modify it, and then store the finalized bar chart information in the database. This includes the following steps:
[0110] S501. Based on the structured data after data cleaning, use Vue to generate a visual table to display the layer bottom depth, layer thickness, layer bottom elevation, rock and soil name, and characteristic information of the drill hole histogram.
[0111] S502. Provide user interaction functions to support users to manually correct or supplement the recognition results. In the interactive interface, set an editable text box for each field, and users can directly modify the value in the text box, so that users can adjust the recognition results according to the actual drilling histogram information to ensure data accuracy.
[0112] S503: Select MySQL database as the storage medium, and store the drilling histogram information confirmed by the user in the database.
[0113] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A method for identifying and extracting drilling histogram information based on a multimodal large model, characterized in that: The following steps are involved: S1, performing image enhancement preprocessing on the input drilling histogram image; S2. Construct the prompt fine-tuning corpus of the multimodal large model and convert the information in the drill bar chart into a structured prompt template; S3: Parameter tuning of the multimodal model and prompt fine-tuning: using the optimized prompt words to fine-tune the multimodal model; S4. Using the fine-tuned multimodal large model to identify and extract drilling histogram information, performing data cleaning on the information output by the multimodal large model, and outputting data in a specified format; S5. Display the cleaned data information, and the user checks and modifies it, and stores the finalized bar chart information in the database.
2. The method for identifying and extracting drilling histogram information based on a multimodal large model according to claim 1, characterized in that: In step S1, the method for preprocessing the input drilling histogram image for image enhancement includes the following steps: S101, reading an input drilling histogram image file to obtain image data; S102 , performing noise reduction processing on the image to remove random noise in the image, enhance the recognizability of the drilling histogram, and output the drilling histogram after noise reduction.
3. The method for identifying and extracting drilling histogram information based on a multimodal large model according to claim 2, characterized in that: In step S2, the method for constructing the prompt fine-tuning corpus of the multimodal large model includes the following steps: S201, analyzing the type of information to be identified in the drill histogram, and determining the location area and meaning of the information according to the drill histogram compilation rules; S202, converting the information into a standardized prompt word format; designing a corresponding prompt word template according to the determined information type and location area, extracting text and symbol information from the image, and converting it into a standardized prompt word format.
4. The method for identifying and extracting drilling histogram information based on a multimodal large model according to claim 3 is characterized in that: The information types include layer bottom depth, layer thickness, layer bottom elevation, rock and soil name and characteristic information.
5. The method for identifying and extracting drilling histogram information based on a multimodal large model according to claim 4 is characterized in that: In step S3, the method of fine-tuning the multimodal large model using the optimized prompt words includes the following steps: S301, set the randomness of the multimodal large model generation to 0.3, the TopP cumulative probability to 0.8, and the frequency penalty to 1.1; S302: Use the prompt word in step S2 as the system prompt word of the large model.
6. The method for identifying and extracting drilling histogram information based on a multimodal large model according to claim 5, characterized in that: The multimodal large model includes Qwen2.5-VL-72B-Instruct and Doubao-1.5-thinking-pro visual multimodal large model.
7. The method for identifying and extracting drilling histogram information based on a multimodal large model according to claim 6, characterized in that: In step S4, the method for performing data cleaning on the information output by the multimodal large model includes the following steps: S401: Calling the multimodal large model API interface, and using the drill histogram and the system prompt word system prompt after the data enhancement processing in step S1 as input information of the multimodal large model; S402: Perform data cleaning on the markdown format information output by the multimodal large model.
8. The method for identifying and extracting drilling histogram information based on a multimodal large model according to claim 7, characterized in that: In step S402, data cleaning includes the following: Input preprocessing: Use replace( / \\n / g,'\n') to process escaped newline characters, and split('\n') to split the input content into lines; Table row identification: Use the regular expression / ^\s*\|(.+)\|\s*$ / to accurately match table rows, filter non-table rows, and retain valid data; Header parsing: extract the first row as the header, use split('|') to split the columns and remove spaces; Separator line detection: Use the isSeparatorLine function to identify separator lines of the type |---|:---|:---:|---:| and dynamically determine the starting row of the data; Data conversion: Process cell line breaks, build key-value pair objects, automatically handle missing columns, and retain original line break information.
9. The method for identifying and extracting drilling histogram information based on a multimodal large model according to claim 8, characterized in that: In step S5, the method for displaying the cleaned data, checking and modifying it by the user, and storing the finalized histogram information in the database includes the following steps: S501. Generate a visualization table based on the structured data after data cleaning, and display the layer bottom depth, layer thickness, layer bottom elevation, rock and soil name and characteristic information of the drill hole histogram through an interactive interface; S502: Provide a user interaction function to support the user to manually correct or supplement the recognition results; S503: Store the drilling histogram information confirmed by the user into a drilling histogram information database.
10. The method for identifying and extracting drilling histogram information based on a multimodal large model according to claim 7, characterized in that: In step S502, an editable text box is set for each field in the interactive interface, and the user can directly modify the value in the text box, so that the user can adjust the recognition result according to the actual drilling histogram information to ensure the accuracy of the data.
Citation Information
Patent Citations
Efficient drilling histogram information automatic extraction method
CN110826393A
Drilling histogram data processing method and device, electronic equipment and storage medium
CN116166729A
Drilling histogram digitization method and device based on structural pattern recognition
CN118658174A
Multi-modal model generation method, multi-modal processing method, and device
WO2025031090A1