Method of converting a page comprising elements of different types to a defined format
A processor-based method for geo-data digitization identifies keywords in pages to automate the extraction and correction of text elements, addressing the inefficiencies of manual methods and improving accuracy and consistency in geo-data processing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-04-02
AI Technical Summary
The manual digitization of geo-data, such as borehole logs, is time-consuming, prone to human error, and lacks repeatability and consistency, with existing OCR technologies struggling to accurately recognize complex geological symbols and unique formats.
A method utilizing a processor to extract and correct text elements from geo-data pages by identifying a specific keyword, applying machine learning for error correction, and outputting to a defined format, enhancing repeatability and accuracy.
Automates the digitization process, reducing manual effort, ensuring high accuracy and consistency across datasets, and facilitating faster decision-making.
Smart Images

Figure EP2025076571_02042026_PF_FP_ABST
Abstract
Description
METHOD OF CONVERTING A PAGE COMPRISING ELEMENTS OF DIFFERENT TYPES TO A DEFINED FORMATFIELD OF THE INVENTION
[0001] The present disclosure generally relates to a method of geodata processing, more specifically, to a method of converting a page comprising elements of different types to a defined format. Unlocking insights from Geo-Data, the present invention further relates to improvements in sustainability and environmental developments: together we create a safe and liveable world.BACKGROUND OF THE INVENTION
[0002] Geo-data encompasses a wide range of data types related to the Earth's subsurface, such as borehole logs, which are critical for various applications in geology, mining, oil and gas exploration, and environmental science.
[0003] There are millions of Geo-data points that date back decades ago and are only available in paper or PDF format. Although this Geo-data is very useful for ground modelling or general engineering purposes, the data is not digital and as such, a tedious, manual effort is required to digitalize these Geo-data.
[0004] Traditionally, the digitization of geo-data, such as borehole logs, is a manual task, involving significant human labour to transcribe or convert physical logs into digital formats. This process is not only time-consuming but also prone to human error, leading to inconsistencies in the data. The manual nature of the task introduces significant delays in data processing, which can be detrimental in scenarios where timely access to digitized logs is critical for decision-making processes.
[0005] Due to its reliance on human input, the manual process suffers from a lack of repeatability and consistency. Different individuals may interpret or transcribe the same log differently, resulting in variations in the output quality. This inconsistency poses a significant problem for applications requiring high levels of accuracy and uniformity across datasets.
[0006] While some users may turn to Optical Character Recognition (OCR) technologies embedded within PDF or document reader applications to alleviate manual burdens, these tools present their own set of challenges.
[0007] Generic OCR technologies are not specialized for borehole logs, which often contain complex geological symbols, handwritten notes, and unique formats that can bedifficult for standard OCR algorithms to accurately recognize and interpret. Consequently, even with OCR, a substantial amount of manual intervention is required for quality assurance and quality control (QA / QC) to correct OCR errors and ensure the fidelity of the digitized data.
[0008] In consideration of the above, there is a need for a solution that significantly reduces the need for manual input, enhances repeatability and consistency, and improves the overall efficiency and reliability of the digitization process.BRIEF SUMMARY OF THE INVENTION
[0009] In one aspect of the invention, there is provided a method of converting a page comprising elements of different types to a defined format, the method performed by a processor and comprising the steps of:
[0010] - extracting text elements of the page, the text elements being recognized using a recognition module configured for converting the text elements to a machine-readable format;
[0011] - recording the extracted text elements of the page and their respective locations;
[0012] - identifying a specific keyword in the recorded text elements of the page;
[0013] - extracting text elements from contours located within a threshold distance to the identified specific keyword;
[0014] - identifying a potential error in the extracted text elements;
[0015] - correcting the identified error by applying a correction;
[0016] - outputting the recognized elements with the identified error corrected to the defined format.
[0017] The present disclosure is based on the inventors’ insight that a page comprising elements of different types, such as a borehole log or a laboratory data, can be converted to a defined format in an automated and efficient way by extracting information from the page based on identification of a specific keyword. The specific keyword indicates how recorded data is organized in the page. Text elements which are located within a threshold distance to the identified specific keyword are the most relevant part of the page, which can provide detailed information about geological formations as recorded in the page, are extracted.
[0018] The method of the present disclosure therefore involves extracting the interested information from the page based on the identified specific keyword. Thereafter, potential error in the recognized elements is identified and corrected with a correction method. The recognized elements with the identified error corrected are then output to a defined format.
[0019] By further incorporating error correction operations into the method, the present disclosure can provide reliability digitized data with repeatability and consistency.
[0020] The method of the present disclosure employs a standardized method to recognize and interpret complex geological symbols, handwritten notes, and varied formats unique to borehole logs. This tailored approach ensures higher accuracy and reliability in digitization.
[0021] By automating the digitization process and minimizing manual tasks, the method significantly accelerates data processing. This efficiency is a direct improvement over the manual, labour-intensive methods of the past, enabling faster decision-making and reducing bottlenecks in data workflow.
[0022] The application of uniform algorithms ensures that each borehole log is processed in a consistent manner, eliminating the variability introduced by human interpretation in the prior art. This standardized approach guarantees that the output quality is repeatable across different datasets.
[0023] In an example of the present disclosure, the recognition module comprises an optical character recognition, OCR, module.
[0024] An OCR module is a readily available means for recognizing text elements in the page, which can be used to recognize the text elements in the page.
[0025] In the present disclosure, the OCR module can be enhanced with a machine learning model which is specifically trained based on training pages having relevant elements of the page, such as borehole logs. The method therefore can identify not only regular text but also specialized elements, such as geological symbols, handwritten notes, and unique layouts often found in these logs.
[0026] In another example of the present disclosure, the recognition model comprises a machine learning-based vision module.
[0027] The machine-learning based vision module can be a machine-learning based vision module for recognizing text from PDFs or images, it can for example leverage advanced deep learning models to detect, classify, and extract text from complex document layouts.
[0028] Such a module may utilize convolutional neural networks (CNNs) and transformerbased architectures to accurately identify and interpret text a scanned image or a PDF page. Machine learning enhances the accuracy of text extraction by adapting to different fonts, languages, and handwritten annotations, while also learning to correct recognition errors and handle complex visual formats. This approach allows for efficient and scalable text extraction with minimal manual intervention, ensuring high accuracy even in documents with challenging layouts.
[0029] In an example of the present disclosure, the step of extracting text elements from contours located within a threshold distance to the identified specific keyword comprises:
[0030] - obtaining a list of contours located within the threshold distance to the identified specific keyword; and
[0031] - iteratively extracting text elements in the obtained list of contours.
[0032] Due to the fact that documents such as borehole logs are usually recorded in a specific format, when the specific keyword indicating recorded data is identified, text elements which are located within a threshold distance to the identified specific keyword and can provide detailed information about geological formations as recorded in the page, are extracted.
[0033] In practice, for example, if a borehole log is being processed, when the specific keyword “description” is identified, it can be known, based on the format of the borehole log, that texts below the “description” are about strata of the borehole. Such texts are located horizontally within a defined threshold to the keyword and therefore should be extracted.
[0034] In an example of the present disclosure, the method further comprises:
[0035] - extracting, for each contour, all capital text in recognized text elements in that contour.
[0036] In borehole logs, "soil type" is a crucial element that provides information about the different types of soil or rock encountered at various depths. Identifying this keyword helps in extracting and analyzing data related to the composition and characteristics of the subsurface material. Soil type information is often categorized in borehole logs to classify the different strata encountered during drilling. Extracting this data allows for better organization and interpretation of geological information.
[0037] In a borehole log, the column stratigraphic description usually gives a description of the soil based on examination of the samples and / or laboratory test results. The soil type is usually recorded in a font format different than other texts. As an example, texts related to the soil type can be recorded in capital letter or in upper case letters, while other text providing more detail information can be recorded using lower case letters. Therefore, by extracting texts recorded in a different format, such as extracting all capital texts, the soil type is obtained.
[0038] In an example of the present disclosure, the step of identifying a potential error in the recognized elements comprises:
[0039] comparing a recognized element with a known pattern error to identify the potential error.
[0040] By comparing recognized elements with known pattern errors, the system can pinpoint discrepancies that may have been missed initially, leading to more accurate finaloutput. Identifying and correcting potential errors proactively helps in minimizing the number of mistakes in the recognized data, improving the overall quality of the output.
[0041] In an example of the present disclosure, the steps of identifying a potential error in the recognized elements and correcting the identified error are performed by an error correction module based on machine learning.
[0042] Machine learning models can learn from past errors and improve their error detection and correction capabilities over time. This adaptability helps in handling new types of errors that may not have been anticipated initially.
[0043] Moreover, machine learning can identify complex patterns and anomalies that traditional rule-based systems might miss. This capability enhances the accuracy of error detection and correction by recognizing subtle deviations from expected patterns.
[0044] Furthermore, ML-based error correction can handle large volumes of data efficiently. As the amount of data increases, the ML model can scale accordingly and maintain high performance in error detection and correction.
[0045] In an example of the present disclosure, the defined format is specified by a user.
[0046] As a result, users can tailor the output format to meet their specific needs and preferences, ensuring that the data is presented in a way that best fits their particular use case or application. By specifying the format, users ensure that the output is directly relevant to their requirements, which can enhance usability and efficiency in how the data is utilized or analyzed.
[0047] In an example of the present disclosure, the defined format comprises industrystandard formats including, but not limited to AGS, Excel, CSV.
[0048] These are exemplary formats that can be used. Other formats can be specified by the user where necessary.
[0049] In an example of the present disclosure, the page comprises at least one of borehole logs and laboratory data.
[0050] Borehole logs and laboratory reports typically have complex layouts with structured data (e.g., depth measurements, geological descriptions) and unstructured data (e.g., observations, comments). The method’s ability to recognize and correct errors ensures that such complex formats are accurately processed.
[0051] In an example of the present disclosure, the elements of different types comprise at least two of text data, visual graphs and symbols.
[0052] Borehole logs and laboratory reports often include a mix of text descriptions, visual graphs, and specialized symbols, including text, numerical data, symbols, and diagrams. Eachtype of data provides different insights, and the ability to process all of these elements ensures a complete and accurate interpretation of the information.
[0053] In an example of the present disclosure, the specific keyword comprises the word “description”.
[0054] Using "description" as a keyword in the context of processing borehole logs and laboratory data is particularly relevant. In borehole logs and laboratory reports, the term "description" often precedes or is associated with detailed explanations of geological formations, soil types, or experimental results. Identifying this keyword helps in pinpointing sections where critical descriptive information is located.
[0055] Therefore, the keyword "description" can help in systematically extracting relevant sections from the document. For example, it might be used to locate detailed descriptions of geological layers in borehole logs or analytical results in laboratory reports, which are essential for understanding the data.
[0056] In an example of the present disclosure, the extracted text elements comprise Client Name, Borehole Date, Borehole Latitude & Longitude, Coordinates, Water Depth.
[0057] Such text elements represent key pieces of information typically found in borehole logs and related documents. As an example, Client Name identifies the individual or organization requesting or commissioning the borehole drilling. It is crucial for record-keeping and linking the data to specific projects or stakeholders. As another example, The date is important for tracking the timing of data collection, which can be relevant for monitoring changes over time or correlating with other events or data. Extracting such information therefore ensures that all relevant data points are accurately identified and organized, facilitating better analysis, reporting, and decision-making related to the borehole and associated projects.
[0058] In an example of the present disclosure, a large language model, LLM, is used to extract text elements of the page. This is detailed in the description.
[0059] A second aspect of the present disclosure provides a device for converting a page comprising elements of different types to a defined format, the device comprising a processor for performing the method according to the first aspect of the present disclosure.
[0060] A third aspect of the present disclosure provides a computer program product, comprising a computer readable storage medium storing instructions which, when executed on at least one processor, cause the at least one processor to carry out the method according to the first aspect of the present disclosure.
[0061] The above mentioned and other features and advantages of the disclosure will be best understood from the following description referring to the attached drawings. In the drawings, like reference numerals denote identical parts or parts performing an identical or comparable function or operation.BRIEF DESCRIPTION OF THE DRAWINGSIn order to describe the manner in which the above-recited and other advantages and features of the disclosure can be obtained, a more particular description of the principles briefly described above will be rendered by reference to specific embodiments thereof which are illustrated in the appended drawings. Understanding that these drawings depict only exemplary embodiments of the disclosure and are therefore not to be considered to be limiting of its scope, the principles herein are described and explained with additional specificity and detail through the use of the accompanying drawings in which:
[0062] FIG. 1 schematically illustrates an exemplary borehole log page.
[0063] FIG. 2 schematically illustrate, in a flow chart, a method of converting a page comprising elements of different types to a defined format, in accordance with the present disclosure.
[0064] FIG. 3 schematically illustrates a schematic diagram of a device for converting a page comprising elements of different types to a defined format.DESCRIPTION OF ILLUSTRATIVE EMBODIMENTS
[0065] Embodiments contemplated by the present disclosure will now be described in more detail with reference to the accompanying drawings. The disclosed subject matter should not be construed as limited to only the embodiments set forth herein. Rather, the illustrated embodiments are provided by way of example to convey the scope of the subject matter to those skilled in the art.
[0066] The present disclosure provides a method for converting a page comprising elements of different types into a defined format using a processor. This method operates to enhance the recognition, error correction, and formatting of complex documents, such as borehole logs and laboratory reports.
[0067] The method of the present disclosure will be detailed below using borehole log data pages as examples. Those skilled in the art will understand that the method as described hereinis applicable to other pages or documents comprising elements of different types, such as for example laboratory data, structural data, metocean data, time-series based data.
[0068] The method of the present disclosure addresses the tedious, manual effort needed to digitize borehole logs. The method addresses the issues described in the background section by digitalizing the entire log, capturing as much information as possible (depending on the quality of the log) and exporting this in a standard output template. A manual, human effort is reduced from two hours to several minutes using the method of the present disclosure.
[0069] The method of the present disclosure employs standardized algorithms specifically trained to recognize and interpret the complex geological symbols, handwritten notes, and varied formats unique to borehole logs. Machine learning models maybe incorporated into some modules of the method.
[0070] The application of uniform algorithms ensures that each borehole log is processed in a consistent manner, eliminating the variability introduced by human interpretation in the prior art. This standardized approach guarantees that the output quality is repeatable across different datasets.
[0071] By leveraging consistent processes for each digitization task, the present disclosure ensures that the output is uniform across different logs, regardless of their original format or condition. This repeatability is a significant improvement over the prior art, where output quality could vary widely due to manual processes.
[0072] The method is developed to process and digitise documents including borehole logs and laboratory data, and to extract the relevant information in a structured manner. The extracted information comprises:
[0073] Project metadata: coordinates, project name, report reference, date of issue, water depth, depth of borehole, Client name, contractor’s name, etc.
[0074] Soil stratigraphy: depth intervals between the different layers, soil description (main soil type, secondary soil type, soil consistency, soil colour, etc.)
[0075] Figure 1 schematically illustrates an exemplary borehole log page. As shown in Figure 1, the exemplary borehole log page comprises project-related information including client name, project name, project number, borehole number, contractor name, date, drill method, location, elevation.
[0076] The exemplary borehole log page further comprises information related to soil description of different layers of the borehole, in which soil type is recorded in capital letters. It can be understood by those skilled in the art that soil type can be recorded in a format differentthan other content of the soil description, such as in a different colour (when the borehole log is a colour copy), in bold font, in italic font or in underlined letter.
[0077] Figure 2 schematically illustrate, in a flow chart, a method of converting a page comprising elements of different types to a defined format, in accordance with the present disclosure.
[0078] The Geo-data, presented as a borehole log or laboratory summary table, is ingested in the forms of images or PDF files into the algorithm of the present disclosure. The algorithm can both be hosted locally on user machines or called on through an API in the cloud.
[0079] The ingested documents, such as a PDF file, is split into separate pages and saved as images.
[0080] As a preparation step 21, each image is preprocessed to enhance its quality. Advanced image processing techniques can be applied to the extracted images to enhance their quality. Methods such as de-noising, contrast enhancement, and edge detection improve image clarity, preparing the images for more accurate OCR.
[0081] At step 22, a recognition module is employed to capture information from the enhanced images. In other words, this step involves extracting text elements of the page using a recognition module configured for converting the text elements to a machine-readable format.
[0082] The recognition module may be for example a customized optical character recognition, OCR, engine tailored to recognize and extract detailed geotechnical data, focusing on accuracy and completeness. The OCR engine may be enhanced with a machine learning, ML, model to capture the maximum amount of information from the enhanced images.
[0083] As an example, the OCR engine detects and extracts text from the preprocessed image. This includes identifying and converting characters, words, and phrases into a digital format. Alongside text, other elements such as symbols, diagrams, and visual graphs are extracted if present.
[0084] Each extracted text element is associated with a bounding box, which is a rectangular area in the image that encloses the text. This bounding box is defined by coordinates (typically the top-left, a contour width, and a contour height) that indicate the precise location of the text within the image.
[0085] For more complex elements such as symbols or graphical data, contours or polygonal regions can be used to define the spatial area where these elements are located.
[0086] According to the present disclosure, the ML model is trained based on a dataset of training pages having one or more of the elements of the page. In practice, for geo-data applications, a comprehensive dataset of geotechnical documents, including borehole logs,laboratory summaries, and related materials, are collected. This dataset encompasses a wide range of text formats, symbols, and layouts specific to the geotechnical field.
[0087] The method of the present disclosure extracts features from the training dataset that are relevant for recognizing and interpreting geotechnical data. Features may include text patterns, symbol shapes, and layout structures.
[0088] ML models, such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs), are then trained using the prepared dataset. The models learn to identify and classify text elements, symbols, and other data types based on the features extracted. The ML model is further fine-tuned by incorporating feedback and new data. This iterative process improves the model’s ability to handle variations and nuances in the data.
[0089] Moreover, the ML model provides contextual understanding by interpreting the recognized text and symbols in relation to the document's content. For instance, it can distinguish between different types of geological symbols based on their context within the document.
[0090] An example of the ML model is a large language model, LLM, fine-tuned on a custom dataset of borehole logs. The ML model can be used to interpret the text extracted by the OCR. This model leverages its multi-modal capabilities to understand and accurately structure the data.
[0091] The LLM may also take a whole image as an input and understands the structure of the data and extracts the requested information, in this case, the LLM is integrated with a vision model. In other words, the machine learning model comprises a vision module and a language module. The vision module may comprise for example a CNN or any vision algorithm, and the language module is typically a large language module.
[0092] As an example, the fine-tuning process of the LLM is as follows:
[0093] Dataset Preparation: Curating a comprehensive dataset of borehole logs, including various formats and terminologies used in the field.
[0094] Model Training: Fine-tuning a pre-trained multi-modal LLM on the curated dataset to adapt it to the specific context of borehole logs.
[0095] Validation and Testing: Conducting rigorous validation and testing to ensure the model's performance meets the desired accuracy and reliability standards.
[0096] Moreover, the text from the OCR can be processed using a customised prompt template designed through extensive prompt engineering. This template captures the geotechnical aspects of the document, ensuring that all relevant information is extracted and correctly interpreted. The prompt template includes:
[0097] Contextual Prompts: Phrasing prompts to provide context about the type of information being extracted, such as "Describe the soil type at depth X."
[0098] Geotechnical Keywords: Incorporating domain-specific keywords and phrases to guide the LLM in accurately interpreting geotechnical data.
[0099] An exemplary prompt template may include the following:“Your task is to take the unstructured text provided and convert it into a well-organized xx format.Given below is the extracted graphical data from a pdf file of Borehole logs. The task is to interpret and analyze the data extracted from complex, varied formats of borehole logs and convert it into a structured format capturing all the information.Strata description is very important information in the borehole logs.I have extracted the data using an OCR tool with page segmentation 4 which means OCR assumes a single column of text of variable sizes. Here is the extracted data: {extracted text} Analyse the data and capture the important information utilizing all the information.Be sure to include information for all the different layers.Use xx format{{\"source_report\": {{\"report_title\": \"\",\"date_of_report\": \"\",\"project_number\": \"\",\" country _code\": \"\",\"contractor_id\": \"\",\"contractor_report_number" : \"\",\"revision\": \"\",\"client_id\": \"\",\"client_report_number\" : \"\",\"company_id\": \"\",\"company_report_number\" : \"\",\"total_page\" : \"\",\"report_type\": \"\",\"object_point\": [{{\"source_id\": V'type of file extensionV,\"point_name\":\"remarks\": \" Additional information and insightsV,\"lithology\": [{{\" secondary _description\": \"\"»]»]»»Summarize the insight of the borehole logs to remarks, keep it a small paragraph of 5 lines max.If any information is not present then substitute it with N / A The final output must strictly only consist of a xx format.Ensure that the data is accurately represented and properly formatted within the xx format structure and no other text either before or after.The resulting xx file should provide a clear, structured overview of the information presented in the original text.”
[0100] The LLM's response is structured into a standardized format and stored in a database. This uniform structuring enables consistent data formatting across different borehole logs, facilitating efficient data aggregation and analysis.
[0101] The data structuring process involves:
[0102] Data Normalization: Converting extracted data into a common schema, including standardizing units of measurement and terminology.
[0103] Database Design: Designing a relational or NoSQL database schema optimized for storing and querying geotechnical data.
[0104] Data Ingestion: Implementing data ingestion pipelines to automate the process of storing structured data into the database.
[0105] At step 23, the extracted text elements of the page and their respective locations are recorded or stored.
[0106] Storing the extracted text elements and their respective locations from a page facilitates subsequent data processing, especially when dealing with complex documents such as borehole logs or laboratory reports. This step involves not just capturing the content of the page but also maintaining the spatial relationships of each element to preserve the context and structure of the original document.
[0107] Each extracted text element is paired with metadata that includes its bounding box coordinates, the page number (if the document consists of multiple pages), and any relevant context or formatting information.
[0108] Text elements and their locations are often organized hierarchically to reflect the structure of the original document. For example, headings, subheadings, and body text might be stored in a nested format to maintain their relative positions.
[0109] As an example, for each extracted text element, the method can record its bounding box coordinates. For example, the text “Water Depth” might be located within a bounding box defined by the top-left coordinates (xl, yl), a contour height h, and a contour width w..
[0110] Metadata such as the page number, font size, and text style may also be stored along with the text content.[OHl] At step 24, a specific keyword in all stored text elements of the page is identified.
[0112] The primary goal is to locate specific sections of the document page that contain relevant information, this is needed in extracting and organizing relevant information from documents such as borehole logs. In the context of a borehole log, the keyword "description" or "stratigraphic description" or “soil description”, either in English or in a different language, plays a significant role in pinpointing detailed geological information, therefore, “description”is the specific keyword searched for borehole log documents. It will be understood by those skilled in the art that other keywords may be used for other types of documents.
[0113] Keywords help in distinguishing between different types of information within a document. In a borehole log, various elements such as measurements, annotations, and descriptions are present. Keywords like "description" guide the extraction process towards the sections containing detailed stratigraphic information. By identifying a keyword such as "description" or "stratigraphic description," the system can focus on extracting and analyzing sections of the log that describe the geological strata encountered during drilling.
[0114] Various information is extracted from the OCR results following the identification of the specific keyword. Extraction of information might include parallel processes involving extraction of general information and extraction of strata data.
[0115] For the extraction of general information, including for example, client name, borehole date, coordinates, water depth, this can be done by searching and locating the relevant text in the stored text elements of the page.
[0116] As for extracting strata data, it is organized in a borehole log in a way for providing a detailed and systematic representation of the geological formations encountered during drilling. This organization ensures that the data is easily interpretable and useful for various applications, including geological analysis, resource exploration, and construction planning.
[0117] Each strata or geological layer is described in detail. This includes the type of material (e.g., sand, clay, gravel), color, texture, and any notable features or characteristics.
[0118] Strata data organization can include the following information:
[0119] A. Stratigraphic Description
[0120] Stratigraphic Unit: Sometimes, the data is organized into stratigraphic units or formations, which group together layers with similar properties or origins.
[0121] B. Lithology
[0122] Lithological Information: Details the composition and physical characteristics of the strata, such as grain size, mineral content, and rock type.
[0123] Symbols and Codes: Standard symbols or codes may be used to represent different lithological types for ease of interpretation.
[0124] C. Geotechnical Properties
[0125] Physical Properties: Includes information on the strength, density, porosity, and permeability of the strata.
[0126] Laboratory Test Results: Data from tests conducted on samples, such as grain size analysis or compaction tests, may be included.
[0127] To extract such information in a way such that proper interpretation of the strata data is possible based on the output in the defined format using the recognized elements of the page, a special way of extracting this information is used in the present disclosure.
[0128] Specifically, at step 25, a list of contours located within a threshold distance to the identified specific keyword is obtained. This is based on the fact that strata data is organized at known locations relevant to the found keyword, in the case of borehole log, the keyword is “description”.
[0129] In other words, the keyword "description" is used as a reference point to locate relevant sections of the log. This keyword typically precedes or is associated with detailed descriptions of the geological strata.
[0130] The keyword "description" indicates that the following text contains detailed information about the geological layers, such as their composition, color, and other characteristics.
[0131] The method identifies contours that are within a certain threshold distance from the identified keyword. This distance is set to ensure that the contours being considered are closely related to the keyword and likely to contain relevant information.
[0132] For a keyword like "description," the method looks for contours around this keyword to locate the associated strata data. Since detailed descriptions of geological strata typically follow or are close to the keyword, this method helps in isolating the relevant sections. In an example, the threshold distance is a horizontal distance with reference to the identified keyword “description”.
[0133] Once the contours within the threshold distance are identified, the method creates a list of these contours. Each contour represents a region in the document that may contain information related to the keyword.
[0134] At step 26, text elements in the obtained list of contours are iteratively extracted. That is, the method then extracts text elements from these identified contours. This ensures that the information captured is directly associated with the keyword and relevant to the strata data.
[0135] Furthermore, at step 27, a second keyword is identified from the extracted text elements. In the case of borehole log, the second keyword is “soil type”. Identifying a keyword like "soil type" allows the system to focus on precise data points relevant to the geological analysis, facilitating a more detailed understanding of the soil or rock layers encountered.
[0136] The “Soil Type” is inferred from the strata description since it is typically provided as an all-caps word(s), so by for the all-caps word(s) in the extracted strata description the soil type is determined.
[0137] It can be contemplated by those skilled in the art that soil type may be recorded in another special font, such as in bold, italic or with underline. The method identify the soil type by relying on such special font.
[0138] At step 28, potential errors in the recognized elements are identified. This is done using for example by comparing a recognized element with a known pattern error to identify the potential error.
[0139] At step 29, the identified error by is corrected by applying a correction.
[0140] In the present disclosure, the steps of identifying a potential error in the recognized elements and correcting the identified error are performed by an error correction module. It can be contemplated by those skilled in the art that the extracted text may be cleaned from any leading or trailing non-alphanumeric characters.
[0141] The method compares recognized elements (such as text or symbols) with these known error patterns. This comparison helps in identifying inconsistencies or deviations that may indicate errors in the recognition process.
[0142] As an example, when the OCR system recognizes the text "Coarse sand with occaional gravel" where "occaional" is a likely misspelling of "occasional." The error detection module compares this recognized text with known patterns of spelling errors and identifies "occaional" as a potential error.
[0143] The method may also use contextual information to aid in error detection. For example, if the recognized text should follow a standard format or known geological terminology, deviations from this standard can be flagged as potential errors.
[0144] Machine learning models are trained on large datasets that include various examples of common errors and correct data. This training helps the model recognize and predict errors based on patterns observed in the data.
[0145] Over time, machine learning models can adapt to new types of errors that were not present in the initial training data. This adaptability enhances the accuracy of error detection and ensures that the system can handle a wide range of error types.
[0146] Once an error is identified, the method applies a correction based on predefined rules or learned patterns. This can involve correcting spelling mistakes, adjusting misrecognized symbols, or reformatting data to match the expected output.
[0147] As an example, if the method identifies a spelling error like "occaional," it can correct it to "occasional" based on a dictionary or learned spelling patterns in a specified language, including but not limited to English, Spanish, French, Portuguese, German, Chinese, Korean, Japanese or Modern Standard Arabic. As another example, if a recognized geologicalterm is incorrect or inconsistent with standard terminology, the system can replace it with the correct term based on context.
[0148] Machine learning models are capable of recognizing complex patterns and anomalies that might be missed by traditional rule-based systems. Error correction can efficiently handle large volumes of data, maintaining high performance even as the amount of data increases.
[0149] At step 30, the recognized elements with the identified error corrected is output to the defined format.
[0150] As can be contemplated by those skilled in the art, the defined format can be specified by a user and can comprise industry-standard formats including AGS, Excel, CSV.
[0151] In an example of the present disclosure, the output is an excel file, which contains information extracted using the above method. Specifically, for the stratigraphic description related text, it can be organized into columns including SOIL TYPE, MAIN DESCRIPTION, SECONDARY DESCRIPTION, which allows the information to be used in a more effective way.
[0152] The above describes the method of the present disclosure, which can be summarized briefly with the following exemplary process:
[0153] 1. The Geo-data, presented as a borehole log or laboratory summary table, is ingested in the forms of images or PDF files into the algorithm of the present disclosure. If PDF is uploaded, each page of the borehole log PDF is extracted for individual processing. The algorithm can both be hosted locally on user machines or called on through an API in the cloud.
[0154] 2 Advanced image processing techniques are applied to the extracted images to enhance their quality. Methods such as de-noising, contrast enhancement, and edge detection improve image clarity, preparing the images for more accurate OCR.
[0155] 3. A custom OCR setting is employed to capture the maximum amount of information from the enhanced images. The OCR system is tailored to recognize and extract detailed geotechnical data, focusing on accuracy and completeness.
[0156] 4. A state-of-the-art LLM, fine-tuned on a custom dataset of borehole logs, interprets the text extracted by the OCR. This model leverages its multi-modal capabilities to understand and accurately structure the data.
[0157] 5. The text from the OCR is processed using a custom prompt template designed through extensive prompt engineering. This template captures the geotechnical aspects of the document, ensuring that all relevant information is extracted and correctly interpreted.
[0158] 6. The LLM's response is structured into a standardized format and stored in a database. This uniform structuring enables consistent data formatting across different borehole logs, facilitating efficient data aggregation and analysis.
[0159] 7 Users specify the desired output format(s) before processing. Upon completion of OCR and error correction, the software converts and organizes the data into the selected formats. The software automatically formats the digitized data into various industry-standard formats (AGS, Excel, CSV), according to user preferences.
[0160] A device for implanting the method of the present disclosure is described in the following with reference to Figure 3.
[0161] Figure 3 schematically illustrates a schematic diagram of the device 300 for converting a page comprising elements of different types to a defined format.
[0162] The device 300 comprises a loading module 301, which is configured to load documents 300 including for example borehole logs into the device. The loaded document may go through preprocessing 311 where needed to enhance the image quality. The loaded and optionally preprocessed document in images is then fed to a recognition module 302 of the device. The recognition module 302 is configured to capture various elements including text elements of the document. The captured texts are then fed into a LLM and custom prompt module 303 of the device. The LLM and custom prompt module 303 is configured to enhance the quality of the captured texts using a fine-tuned LLM on custom data. The data processed by the LLM and custom prompt module 303 is then fed to an error correction module 304, which is configured to identify and correct potential errors present in the processed text. The corrected texts are then formatted by a data formation module 305, which is then output to a storage 313.
[0163] The invention has been described by reference to certain embodiments discussed above. It will be recognized that these embodiments are susceptible to various modifications and alternative forms well known to those of skill in the art.
[0164] Further modifications in addition to those described above may be made to the structures and techniques described herein without departing from the spirit and scope of the invention. Accordingly, although specific embodiments have been described, these are examples only and are not limiting upon the scope of the invention.
Claims
CLAIMS1. A method of converting a page comprising elements of different types to a defined format, the method performed by a processor and comprising the steps of: extracting text elements of the page, the text elements being recognized using a recognition module configured for converting the text elements to a machine-readable format; recording the extracted text elements of the page and their respective locations; identifying a specific keyword in the recorded text elements of the page; extracting text elements from contours located within a threshold distance to the identified specific keyword; identifying a potential error in the extracted text elements; correcting the identified error by applying a correction; outputting the recognized elements with the identified error corrected to the defined format.
2. The method according to claim 1, wherein the recognition module comprises an optical character recognition, OCR, module.
3. The method according to claim 1, wherein the recognition model comprises a machine learning-based vision module.
4. The method according to any of the previous claims, wherein the step of extracting text elements from contours located within a threshold distance to the identified specific keyword comprises: obtaining a list of contours located within the threshold distance to the identified specific keyword; and iteratively extracting text elements in the obtained list of contours.
5. The method according to claim 4, further comprising: extracting, for each contour, all capital text in recognized text elements in that contour.
6. The method according to any of the previous claims, wherein the step of identifying a potential error in the recognized elements comprises:comparing a recognized element with a known pattern error to identify the potential error.
7. The method according to any of the previous claims, wherein the defined format is specified by a user.
8. The method according to claim 7, wherein the defined format comprises industrystandard formats including AGS, Excel, CSV.
9. The method according to any of the previous claims, wherein the page comprises at least one of borehole logs, laboratory data.
10. The method according to any of the previous claims, wherein the elements of different types comprise at least two of text data, visual graphs and symbols.
11. The method according to claim 8 or 9, wherein the specific keyword comprises the word “description”.
12. The method according to any of the previous claims 8 to 11, wherein the extracted text elements comprise Client Name, Borehole Date, Borehole Latitude & Longitude, Coordinates, Water depth.
13. The method according to any of the previous claims, wherein a large language model, LLM, is used to extract text elements of the page.
14. A device for converting a page comprising elements of different types to a defined format, the device comprising a processor for performing the method according to any of the previous claims 1 to 13.
15. A computer program product, comprising a computer readable storage medium storing instructions which, when executed on at least one processor, cause the at least one processor to carry out the method according to any of the claims 1 to 13.
Citation Information
Patent Citations
A method of digitalising engineering documents
EP3104302B1
Method and system for configuring devices of a control system based on engineering graphic objects
US20170228589A1
System and method for detection and auto-validation of key data in any non-handwritten document
US20230205800A1