Drilling and completion report review method, apparatus, device, and medium
Patent Information
- Application Number
- CN202610964518.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]然而,人工审核处理效率较低,面对大批量报告时易受人员经验与工作负荷影响;自动化工具主要适用于字段、格式等显性内容校验,对跨章节数据关联、工序时序关系及非结构化文本的逻辑审查能力有限,且在报告模板或审核规范变化时适配成本较高
[0071]本申请实施例提供的钻完井报告审核方法、装置、设备及介质,通过获取待审核钻完井报告文档,并将其解析为包含章节层级信息的结构化文本,能够为后续审核提供统一、清晰的数据组织基础;通过分别对结构化文本执行业务规则校验和语义审核,能够同时覆盖显性规则检查与文本语义层面的审核需求;进而通过融合规则校验结果与语义审核结果生成钻完井报告审核结果,达到能够提高钻完井报告审核的全面性、准确性和处理效率,并增强对不同报告内容及审核要求变化的适应能力的效果。
Smart Images

Figure CN122821582A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent document review and information processing in oil and gas extraction, and in particular to a method, apparatus, equipment and medium for reviewing drilling and completion reports. Background Technology
[0002] In the process of oil and gas development, the review of drilling and completion reports usually adopts a combination of manual verification and automated verification based on preset rules to check the completeness of the report content, the compliance of the format, and the consistency of the data.
[0003] However, manual review is inefficient and easily affected by staff experience and workload when dealing with a large number of reports. Automated tools are mainly suitable for validating explicit content such as fields and formats, but have limited ability to logically review cross-chapter data associations, process sequence relationships, and unstructured text. Furthermore, they have high adaptation costs when report templates or review standards change.
[0004] Therefore, how to improve the efficiency, accuracy, and adaptability of drilling and completion report review while meeting data security requirements has become a technical problem that needs to be solved. Summary of the Invention
[0005] This application provides a drilling and completion report review method, apparatus, equipment, and medium to solve the aforementioned technical problems. This solution is designed for drilling and completion report documents to be reviewed, focusing on the structured understanding of report content, comprehensive judgment of review information, and unified output of review results. This enables the review of not only explicit content such as format and fields, but also the review of hierarchical, related, and semantic content within the report text, thereby improving the efficiency, accuracy, and adaptability of drilling and completion report review.
[0006] In a first aspect, embodiments of this application provide a method for reviewing drilling and completion reports, including:
[0007] Obtain the drilling and completion report document pending review;
[0008] The drilling and completion report documents to be reviewed are parsed in a structured manner to obtain structured text containing chapter and hierarchical information;
[0009] Perform business rule validation on the structured text and obtain the rule validation results; and perform semantic audit on the structured text and obtain the semantic audit results.
[0010] The results of rule verification and semantic review are combined to generate the drilling and completion report review results.
[0011] In one possible embodiment, the drilling and completion report document to be reviewed is subjected to structured parsing to obtain structured text containing chapter-level information, including:
[0012] Iterate through the entire content of the drilling and completion report document awaiting review and identify the heading level information of each paragraph;
[0013] Based on the heading level information of each paragraph, the heading level score of each paragraph is calculated separately;
[0014] Chapter boundaries are determined based on the title level score;
[0015] Based on chapter boundaries, the drilling and completion report documents awaiting review are grouped by chapter and the main text content is extracted to generate structured text.
[0016] In one possible embodiment, based on the heading level information of each paragraph, a heading level score for each paragraph is calculated, including:
[0017] Assign first-level weight to paragraphs whose heading level information is level one heading;
[0018] Assign second-level weights to paragraphs with heading level information of level two;
[0019] Assign third-level weights to paragraphs with heading level information of level three;
[0020] Among them, the weight of the first level is greater than the weight of the second level, and the weight of the second level is greater than the weight of the third level.
[0021] In one possible embodiment, business rule validation is performed on the structured text to obtain the rule validation result, including:
[0022] Perform NOT NULL checks on key fields in structured text to obtain field integrity determination results;
[0023] Perform format pattern matching on key data in structured text to obtain the format conformity judgment result;
[0024] Perform keyword retrieval on standard engineering terms in structured text to obtain content completeness assessment results;
[0025] Perform a consistency comparison on cross-chapter related data in structured text to obtain data consistency determination results;
[0026] The actual chapter titles in the structured text are traversed and matched based on a preset standard chapter directory list to obtain the structural integrity judgment result;
[0027] The rule verification results are obtained by summarizing the results of field integrity judgment, format standardization judgment, content integrity judgment, data consistency judgment, and structural integrity judgment.
[0028] In one possible embodiment, semantic auditing is performed on the structured text to obtain the semantic auditing result, including:
[0029] The target chapters suitable for semantic review in the structured text are concatenated into the standardized input text, and the local lightweight language model is called to perform semantic review chapter by chapter to obtain the semantic quality judgment results of each target chapter.
[0030] Calculate the ratio of the effective text length of each target chapter to the total length of the chapter, and determine the content confidence of each target chapter based on the ratio and the content quality coefficient to obtain the semantic credibility judgment result;
[0031] The local optical character recognition model is called to extract the text information of the images in the drilling and completion report document to be reviewed. The text information is then compared with the text content of the corresponding chapter to obtain the image consistency judgment result.
[0032] The semantic quality assessment results, semantic credibility assessment results, and image consistency assessment results are combined to obtain the semantic audit results.
[0033] In one possible embodiment, the method further includes:
[0034] Receive natural language rules from user input;
[0035] Parse natural language rules into executable validation logic;
[0036] Add executable validation logic to the preset audit rule library; the preset audit rule library is used to perform business rule validation on structured text.
[0037] In one possible embodiment, the rule verification result and the semantic review result are merged to generate the drilling and completion report review result, including:
[0038] The rule verification results and semantic review results are categorized into layers based on structural integrity, data consistency, format standardization, and text semantic quality.
[0039] Abnormal issues are automatically numbered and listed, compliant content is marked, and the drilling and completion report review results are output.
[0040] Secondly, embodiments of this application provide a drilling and completion report review device, comprising:
[0041] The acquisition module is used to retrieve drilling and completion report documents pending review.
[0042] The parsing module is used to perform structured parsing on the drilling and completion report documents to be reviewed, and obtain structured text containing chapter and hierarchical information;
[0043] The rule validation module is used to perform business rule validation on structured text and obtain the rule validation results;
[0044] The semantic auditing module is used to perform semantic auditing on structured text and obtain the semantic auditing results;
[0045] The fusion module is used to merge rule verification results and semantic review results to generate drilling and completion report review results.
[0046] In one possible implementation, the target chapter shall include at least one or more of the following: a summary of construction techniques, a summary of special techniques, and explanatory notes.
[0047] In one possible implementation, the parsing module is specifically used for:
[0048] Iterate through the entire content of the drilling and completion report document awaiting review and identify the heading level information of each paragraph;
[0049] Based on the heading level information of each paragraph, the heading level score of each paragraph is calculated. Paragraphs with heading level information of level 1 are assigned first-level weight, paragraphs with heading level information of level 2 are assigned second-level weight, and paragraphs with heading level information of level 3 are assigned third-level weight. The first-level weight is greater than the second-level weight, and the second-level weight is greater than the third-level weight.
[0050] Chapter boundaries are determined based on the title level score;
[0051] Based on chapter boundaries, the drilling and completion report documents awaiting review are grouped by chapter and the main text content is extracted to generate structured text.
[0052] In one possible implementation, the rule validation module is specifically used for:
[0053] Perform NOT NULL checks on key fields in structured text to obtain field integrity determination results;
[0054] Perform format pattern matching on key data in structured text to obtain the format conformity judgment result;
[0055] Perform keyword retrieval on standard engineering terms in structured text to obtain content completeness assessment results;
[0056] Perform a consistency comparison on cross-chapter related data in structured text to obtain data consistency determination results;
[0057] The actual chapter titles in the structured text are traversed and matched based on a preset standard chapter directory list to obtain the structural integrity judgment result;
[0058] The rule verification results are obtained by summarizing the results of field integrity judgment, format standardization judgment, content integrity judgment, data consistency judgment, and structural integrity judgment.
[0059] In one possible implementation, the semantic auditing module is specifically used for:
[0060] The target chapters suitable for semantic review in the structured text are concatenated into the standardized input text, and the local lightweight language model is called to perform semantic review chapter by chapter to obtain the semantic quality judgment results of each target chapter.
[0061] Calculate the ratio of the effective text length of each target chapter to the total length of the chapter, and determine the content confidence of each target chapter based on the ratio and the content quality coefficient to obtain the semantic credibility judgment result;
[0062] The local optical character recognition model is called to extract the text information of the images in the drilling and completion report document to be reviewed. The text information is then compared with the text content of the corresponding chapter to obtain the image consistency judgment result.
[0063] The semantic quality assessment results, semantic credibility assessment results, and image consistency assessment results are combined to obtain the semantic audit results.
[0064] In one possible implementation, the device further includes:
[0065] The rule extension module is used to receive natural language rules input by the user, call the local lightweight language model to parse the natural language rules into executable verification logic, and append the executable verification logic to the preset audit rule library; the preset audit rule library is used to perform business rule verification on structured text.
[0066] In one possible implementation, the fusion module is specifically used for:
[0067] The rule verification results and semantic review results are categorized into layers based on structural integrity, data consistency, format standardization, and text semantic quality.
[0068] Abnormal issues are automatically numbered and listed, compliant content is marked, and the drilling and completion report review results are output.
[0069] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the methods provided above.
[0070] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed, implements the methods provided above.
[0071] The drilling and completion report review method, apparatus, equipment, and medium provided in this application, by acquiring the drilling and completion report document to be reviewed and parsing it into structured text containing chapter-level information, can provide a unified and clear data organization foundation for subsequent review; by performing business rule verification and semantic review on the structured text respectively, it can simultaneously cover the review requirements at the explicit rule check and text semantic level; and by fusing the rule verification results and semantic review results to generate the drilling and completion report review results, it can improve the comprehensiveness, accuracy, and processing efficiency of drilling and completion report review, and enhance the adaptability to changes in different report content and review requirements. Attached Figure Description
[0072] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0073] Figure 1 A schematic diagram of the overall system architecture of a drilling and completion report review method provided in this application;
[0074] Figure 2 Flowchart of the drilling and completion report review method provided for this application Figure 1 ;
[0075] Figure 3 Flowchart of the drilling and completion report review method provided for this application Figure 2 ;
[0076] Figure 4 Flowchart of the drilling and completion report review method provided for this application Figure 3 ;
[0077] Figure 5 Flowchart of the drilling and completion report review method provided for this application Figure 4 ;
[0078] Figure 6 A schematic diagram of the drilling and completion report review device provided in this application;
[0079] Figure 7 A schematic diagram of the structure of the electronic device provided in this application.
[0080] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0081] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0082] First, let me explain the terms used in this application:
[0083] Drilling and Completion Report: In oil and gas extraction projects, a standardized document that records the technical data, construction techniques, equipment operating status, and results summary of the entire drilling and completion process. It is usually compiled in Word format and includes a cover, table of contents, multiple chapters of text, data tables, and figures.
[0084] Business rule verification: Based on preset rigid audit rules, this is the process of quantitatively judging and logically checking the data format, field integrity, cross-chapter consistency, and structural standardization of drilling and completion reports.
[0085] Rule Engine: A driver component that abstracts business review rules into executable logic code. In this embodiment, it is implemented through Python functions and supports the automatic execution of rule types such as non-empty judgment, regular expression matching, keyword retrieval, and multi-field association validation.
[0086] Semantic review: Based on natural language processing technology, the process of reviewing the unstructured text content (such as technical summaries, process descriptions, etc.) in drilling and completion reports for standardization of expression, logical rationality, completeness of content, and correctness of text.
[0087] Lightweight language model: A large language model that has undergone parameter compression or architecture optimization, which can run offline on a local ordinary computing device. In this embodiment, the Qwen3.5:9b model is deployed through the Ollam framework.
[0088] Large Language Model (LLM): A large-scale parameter neural network model trained based on deep learning technology, which has the ability to understand, generate and reason about text. In this application, it is used to perform semantic review and natural language rule parsing.
[0089] Heading Level Weighted Recognition Algorithm: A document structure parsing algorithm that identifies the heading level attributes (first-level heading, second-level heading, third-level heading) of paragraphs in a Word document, assigns different weight coefficients to calculate the heading level score, and then determines the chapter boundaries.
[0090] Structured text: Parses and converts elements such as paragraphs, tables, and headings in the original Word document into standardized text data with clear hierarchical relationships and format specifications. In this embodiment, Markdown format is used as an intermediate representation.
[0091] Chapter boundaries: The starting and ending positions of document chapters are identified based on the heading level score. This is used to divide continuous document content into independent chapter units for group review and data extraction.
[0092] Structured parsing: The process of converting unstructured or semi-structured Word document content into standardized data that is machine-readable, hierarchically clear, and uniformly formatted, including operations such as heading recognition, paragraph extraction, table splitting, and chapter reorganization.
[0093] Dynamic extension of natural language rules: A mechanism in which users describe new review rules in natural language, which are then parsed into executable verification logic by a local lightweight language model and added to the rule base, enabling zero-code dynamic maintenance of business rules.
[0094] The three-layer review architecture proposed in this application is a collaborative mechanism of "basic rules as a safety net, natural language dynamic expansion, and AI review as a backup". The rule engine is responsible for rigid verification, while the AI model is responsible for semantic review and serves as the dynamic expansion layer and final review correction layer of the rule engine.
[0095] Content confidence: A quantitative indicator used to measure the reliability of semantic review results. It is calculated by combining the ratio of effective text length to total chapter length with the content quality coefficient. It is used to filter vague or low-quality text and suppress model illusions.
[0096] Model illusion: Large language models tend to generate false information that is inconsistent with the input facts, logically contradictory, or fabricated when generating content. This application suppresses this by using low temperature parameters, negative constraint instructions, and block review strategies.
[0097] Temperature parameter: A parameter for controlling the randomness of text generation by a large language model. The lower the value, the more deterministic and conservative the model output. In this embodiment, temperature=0.2 is set to reduce the risk of hallucination.
[0098] Regular expression: a symbolic syntax used to describe string matching patterns. In this application, it is used to perform format specification validation on fields such as hash symbols, numbers, units, and dates.
[0099] Cross-chapter consistency: When the same key parameter appears in different chapters or tables in the drilling and completion report, its value, unit or description should remain the same. This application achieves cross-chapter data consistency comparison through multi-field association verification.
[0100] Classified data: Information involving core technologies, geological parameters, and production data related to oil and gas exploration and development that have confidentiality requirements. This application ensures that classified data does not leave the domain through a fully local offline process.
[0101] Zero data loss: Data remains on a local physical device or private network throughout its entire lifecycle of generation, transmission, processing, and storage, and is not transmitted to external servers or public cloud platforms via the Internet.
[0102] Ollama: A tool for local deployment of private large language models, supporting the download, running, and management of lightweight language models on local devices, ensuring that data does not leave the domain.
[0103] Qwen: A local semantic moderation model.
[0104] RapidOCR: A local optical character recognition engine used to extract text information from images in drilling and completion report documents, which can run without an internet connection.
[0105] python-docx: A third-party Python library for reading, parsing, and modifying Microsoft Word documents (.docx format). In this application, it is used to traverse document paragraphs, identify heading styles, and extract table content.
[0106] Markdown: A lightweight markup language that uses concise syntax to represent text format. In this application, it is used as an intermediate data format after structured parsing, which facilitates unified processing by the rule engine and the lightweight model.
[0107] Well number: A unique identifier for an oil and gas well. It is a key index field that runs throughout the drilling and completion report and must remain consistent throughout the entire report.
[0108] Well depth: The vertical depth or measured depth of the wellbore trajectory during drilling is a core engineering parameter in the drilling and completion report, and its numerical continuity and cross-chapter consistency must be verified.
[0109] Start-up point: In directional drilling, the starting depth point at which the wellbore trajectory deviates from the vertical direction from the vertical section is the critical engineering parameter.
[0110] Target point: The target location that the wellbore trajectory needs to reach in directional drilling. It is usually described in the form of coordinates and depth and must be consistent with the wellbore trajectory data.
[0111] Construction Log: A log document that records the daily work content, process time, equipment operation and abnormal situations at the drilling site, and must be consistent with the key process time in the actual drilling data.
[0112] Actual drilling data: Measurement while drilling (MWD) data or logging data collected during the actual drilling process, reflecting the real downhole conditions, and need to be compared with the time records in the construction log for consistency.
[0113] NOT NULL check: A basic method for validating the integrity of required fields. The condition is that Field ≠ empty set. If the field value is empty, it is determined that key information is missing.
[0114] Keyword location search: A method to search for the presence of standard engineering terms or key fields in text. The judgment formula is Keyword ⊂ Text. If no match is found, the content is considered missing.
[0115] Multi-field association verification: A method for verifying the numerical and logical consistency of the same type of parameters distributed in different chapters or tables in a document. The judgment formula is |Data1 − Data2| ≤ ε.
[0116] Chapter matching degree: The matching ratio between the actual parsed document chapters and the standard template chapters, used to quantitatively evaluate the integrity of the document structure.
[0117] Cover: The first page of the well completion report, which usually contains key information such as well number, author, reviewer, and date. It needs to be identified and skipped during parsing to avoid interfering with the determination of chapter boundaries.
[0118] Table of Contents: An index page in the drilling and completion report that lists the titles and page numbers of each chapter. It needs to be identified and skipped during parsing.
[0119] Data structure: The process of extracting table data from a Word document and converting it into a structured array or dictionary format, which facilitates numerical calculations and rule validation.
[0120] Drilling and completion report review technology falls under the field of engineering document processing and quality control in the oil and gas development process. It is applied in business scenarios where oilfield units, drilling construction units, or technical management departments conduct pre-submission review, pre-archiving examination, and batch quality checks of drilling and completion reports. The relevant documents typically include report documents containing chapter titles, table data, and explanatory text. The review process must focus on the completeness of content, formatting standards, and data accuracy.
[0121] In existing technologies, the common practice is for auditors to check drilling and completion reports using preset verification tools. The manual method mainly relies on experience to check the well number, well depth, construction records and chapter content item by item, while the automated method mostly uses field non-empty judgment, format matching or fixed rule comparison to find explicit errors. Its basic principle is to perform formal verification of the report text within the scope of established templates and rules.
[0122] However, in practical applications, manual review has a long processing cycle when dealing with a large number of reports, and is prone to omissions due to fatigue or differences in experience. While automated tools with fixed rules can complete some basic checks, they struggle to handle issues such as cross-chapter data correlation, logical progression of construction procedures, and the reasonableness of explanatory text. When report templates, review criteria, or document structures change, existing tools often need to be re-adapted, resulting in high maintenance costs and making it difficult to guarantee the comprehensiveness and stability of review results.
[0123] The aforementioned shortcomings directly impact the efficiency and quality of drilling and completion report review, especially in scenarios requiring a balance between accuracy, batch processing capabilities, and data security. Existing methods struggle to simultaneously meet the practical needs of engineering document review. Therefore, how to collaboratively review the structured and semantic content of reports while ensuring review quality and improving processing efficiency has become a pressing technical challenge.
[0124] In view of this, a drilling and completion report review method is provided. After obtaining the drilling and completion report document to be reviewed, the document is first parsed in a structured manner to obtain structured text containing chapter and hierarchical information. Then, business rule validation and semantic review are performed on the structured text, yielding rule validation results and semantic review results respectively. Finally, the rule validation results and semantic review results are merged to generate the drilling and completion report review result. This technical approach, with the collaborative processing of structured parsing, rule validation, and semantic review as its core, can improve the efficiency, accuracy, and adaptability of report review.
[0125] Figure 1 This is a schematic diagram of the overall system architecture for a drilling and completion report review method provided in this application. Figure 1 As shown, the drilling and completion report review method provided in this application is deployed on a local computing device and runs in a completely offline private environment. It is used to provide fully automated business rule verification and semantic review collaborative management for drilling and completion reports.
[0126] The local computing devices are ordinary office computers or field workstations, requiring no dedicated GPU servers or cloud computing power, ensuring zero out-of-domain leakage of classified data. The drilling and completion report review system is developed based on the Python open-source environment. Through the collaborative operation of modules such as document parsing, rule validation, semantic review, and result fusion, it achieves automated conversion from raw Word documents to structured review results.
[0127] In this embodiment of the application, the system architecture is divided into five functional layers from left to right: input layer, parsing layer, auditing layer, support layer and output layer.
[0128] Specifically, the input layer contains drilling and completion report documents to be reviewed, usually Word files in .docx format, which are imported into the system by users through a file interface and serve as the original data source for the review process.
[0129] The parsing layer includes a document parsing module responsible for structured parsing of drilling and completion reports awaiting review. This module uses the python-docx library to traverse the entire document, identifying the cover, table of contents, and paragraph heading levels. It employs a weighted heading level recognition algorithm to calculate heading level scores and determine chapter boundaries. It then extracts main text paragraphs and table data by chapter, merging scattered sections to ultimately generate standardized Markdown structured text containing chapter level information. This provides a consistent, clearly defined, and length-controlled data foundation for subsequent review.
[0130] The review layer comprises a rule verification module and a semantic review module, forming a two-way collaborative mechanism. The rule verification module is responsible for executing business rule verification. Relying on a pre-set review rule library, it uses methods such as non-empty checks, regular expression matching, keyword retrieval, and multi-field association verification to rigidly verify the data accuracy, format compliance, structural integrity, and cross-chapter consistency of structured text. The semantic review module is responsible for performing semantic review. By calling a local lightweight model, it performs flexible review of the expression standardization, logical rationality, and content completeness of unstructured text chapters such as construction technology summaries and featured technology summaries. Simultaneously, it uses an OCR recognition module to extract text information from document images and compares it with the corresponding chapter text for consistency. The rule verification module and the semantic review module collaborate through a two-way interactive channel: the rule verification module sends difficult-to-match abnormal rules or boundary cases to the semantic review module for secondary review, while the semantic review module feeds back potential rule conflicts or new rule requirements to the rule verification module, forming a three-layer review architecture of "basic rules as a safety net, dynamic expansion using natural language, and AI-based secondary review as a backup."
[0131] The support layer comprises locally deployed support components, including a local lightweight model (LLM), an OCR recognition module, and an extensible audit rule base. The local lightweight model deploys lightweight large language models such as qwen3.5:9b through the Ollam framework, providing offline inference capabilities for semantic auditing. The OCR recognition module uses the RapidOCR engine to achieve local extraction of text from images. The audit rule base contains a standard chapter directory and business verification rules for drilling and completion reports, and supports dynamic expansion through natural language input. The local lightweight model parses the input into executable verification logic, which is then appended to the rule base, adapting to the personalized auditing needs of specific oilfields or well types without modifying the underlying code. All support components run on local devices, without accessing the external network or uploading to the public cloud, preventing the leakage of confidential information throughout the entire data flow chain.
[0132] The output layer includes the fusion module and the drilling and completion report review results. The fusion module receives the rule verification results from the rule verification module and the semantic review results from the semantic review module. It then categorizes the results according to structural integrity, data consistency, format standardization, and text semantic quality. It automatically numbers and lists abnormal issues and clearly identifies compliant content, ultimately generating a well-structured, complete, and highly readable drilling and completion report review result document.
[0133] The drilling and completion report review method provided in this application solves the technical problems of low efficiency of manual review, inability of traditional tools to adapt to complex business logic, inability to upload confidential documents to the cloud, high computing power threshold of local large models, and easy parsing failure or inference error of lightweight models due to excessively long text in existing technologies through bidirectional collaboration of document structured parsing, rigid verification of rule engine and semantic review of local lightweight model, full-process local offline operation and dynamic expansion of natural language rules.
[0134] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0135] Figure 2 Flowchart of the drilling and completion report review method provided for this application Figure 1 ,like Figure 2 As shown, the method in this embodiment can be specifically executed by a drilling and completion report intelligent review system deployed on a local computing device. Specifically, the method includes:
[0136] S201. Obtain the drilling and completion report document to be reviewed.
[0137] The drilling and completion report document to be reviewed is a standardized document that records technical data, construction techniques, equipment operating status, and results summary of the entire drilling and completion process in oil and gas extraction engineering, serving as the original data source for the review process. This document is usually prepared in Word format (.docx) and includes a cover, table of contents, multiple chapters of text, data tables, and figures.
[0138] In this step, the system retrieves the user-specified drilling and completion report document to be reviewed via a file reading interface. The system first performs environment and file validity checks to confirm the document exists and is formatted correctly, providing an input basis for subsequent parsing and review.
[0139] For example, the documents to be reviewed may come from archived documents of the oilfield data management department, results reports submitted by the construction unit, or well completion summaries prepared by the field engineer.
[0140] S202. Perform structured parsing on the drilling and completion report document to be reviewed to obtain structured text containing chapter-level information.
[0141] Structured parsing is used to convert unstructured or semi-structured Word document content into standardized data that is machine-readable, clearly hierarchical, and uniformly formatted. Its purpose is to eliminate the interference of original document formatting differences on subsequent review, providing standardized input for rule validation and semantic review. Chapter hierarchy information is used to identify the hierarchical relationships and boundary positions of each chapter in the document, serving as the basis for subsequent chapter-based grouping review and data extraction.
[0142] In this step, a Python script can be used to call the python-docx third-party library to traverse the entire content of the drilling and completion report document to be reviewed, identify the heading level information of each paragraph; calculate the heading level score of each paragraph based on the heading level information of each paragraph; determine the chapter boundaries based on the heading level scores; group the drilling and completion report document to be reviewed by chapter based on the chapter boundaries and extract the main text content to generate structured text.
[0143] Specifically, the system first determines whether the current paragraph is a cover page or a table of contents page. If not, it further determines whether it is a first-level, second-level, or third-level heading. A heading level weighted recognition algorithm is used to calculate the heading level score; the calculation formula is as follows:
[0144] Score = α·L1+β·L2+γ·L3
[0145] Here, L1, L2, and L3 represent the matching results for level 1, level 2, and level 3 headings, respectively, while α, β, and γ are the hierarchy weight coefficients, with α > β > γ. After determining the chapter boundaries using these scores, the main text paragraphs and table data are extracted by chapter group. The paragraph text and table content are then broken down and converted into a data structure. The scattered sections are then structurally organized and merged into a single chapter, ultimately generating structured text in Markdown format.
[0146] The above-mentioned structured processing simultaneously meets two core requirements: on the one hand, it provides standardized data with a unified format and clear boundaries for business rule verification, ensuring that verification logic such as data consistency, format standardization, and structural integrity can be executed stably; on the other hand, it provides standardized input with controllable length and clear structure for local lightweight model inference, solving the problem that lightweight models are prone to parsing failure or inference errors due to excessively long text.
[0147] For example, structured text can include elements such as chapter titles, body paragraphs, data tables, and subsection information, using Markdown format as an intermediate representation to facilitate unified processing by rule engines and lightweight models.
[0148] S203. Perform business rule validation on the structured text and obtain the rule validation results.
[0149] The business rule verification is used to quantitatively judge and logically check the data format, field integrity, cross-chapter consistency, and structural standardization of drilling and completion reports based on preset rigid audit rules. Its function is to achieve rigid audit of data accuracy, format compliance, and structural integrity. The rule verification results are used to summarize the judgment conclusions of each individual verification, providing a rigid audit basis for subsequent integrated output.
[0150] In this step, the business validation rules of the drilling and completion report can be converted into executable Python code by performing core methods such as non-empty checks, regular expression matching, keyword location and search, and multi-field association validation through the rule engine.
[0151] Specifically, key fields in the structured text (such as hash numbers, author information, etc.) are checked for non-emptiness using the formula Field ≠ empty set. If empty, key information is considered missing, resulting in a field integrity assessment. Key data in the structured text (such as hash numbers, numbers, units, dates, etc.) are checked for format pattern matching using the formula Text ∈RegPattern. If no match is found, the format is considered non-standard, resulting in a format standardization assessment. Standard engineering terms in the structured text are checked for keyword retrieval using the formula Keyword ⊂ Text. If no match is found, content is considered missing, resulting in a content integrity assessment. Cross-chapter related data in the structured text (such as well depth, directional drilling point, target point, etc., cross-table parameters) are checked for consistency using the formula |Data1 − Data2| ≤ ε. If the error exceeds ε, the data is considered inconsistent, resulting in a data consistency assessment. In addition, the actual chapter titles in the structured text are traversed and matched based on a preset standard chapter directory list, using a chapter matching degree and integrity scoring model.
[0152] S = (N match / N total ) × 100%
[0153] Where, N match N represents the number of standard chapters actually matched. total The total number of chapters in the standard template is used as the reference. When S < 100%, it is determined that a necessary chapter is missing, and the structural integrity judgment result is obtained. Finally, the field integrity judgment result, format standardization judgment result, content integrity judgment result, data consistency judgment result, and structural integrity judgment result are summarized to obtain the rule verification result.
[0154] For example, business rule verification can cover scenarios such as well number consistency check, well depth data continuity check, drilling completion data logic check, typesetting standard check, key process time consistency check between construction log and actual drilling data, cross-chapter data consistency check, and full-text unit and number format standardization check.
[0155] S204. Perform semantic auditing on the structured text to obtain the semantic auditing results.
[0156] Semantic review, based on natural language processing technology, examines the unstructured text content in drilling and completion reports for accuracy, logical coherence, completeness, and correctness. Its role is to cover flexible text quality reviews that are difficult to handle with fixed rules. The semantic review results are used to summarize the judgments of each individual semantic review, providing a flexible review basis for subsequent fusion output.
[0157] In this step, semantic review can be performed using a local lightweight language model, and at the same time, a content confidence judgment mechanism and local optical character recognition technology can be combined to conduct multi-dimensional review of text semantic quality, credibility and image consistency.
[0158] Specifically, firstly, target chapters suitable for semantic review (such as unstructured text chapters like construction technology summaries, special technology summaries, and appendix descriptions) are selected and concatenated into standardized input text. Then, a local lightweight language model (such as the qwen3.5:9b model deployed via the Ollam framework) is invoked chapter by chapter to perform semantic review. By setting a low-temperature parameter (temperature=0.2) and explicit negative constraint instructions, the focus is on checking the text's standardization, completeness, logical rationality, and typographical correctness, obtaining the semantic quality judgment results for each target chapter. Finally, the ratio of the effective text length of each target chapter to the total chapter length is calculated. Based on this ratio and the content quality coefficient λ, the content confidence level of each target chapter is determined using the following formula:
[0159] C = (L eff / L all ) × λ
[0160] Among them, L eff For the effective text length, L all λ represents the total length of the chapters, and λ is the content quality coefficient. A confidence score C is used to distinguish between vague, missing, and substandard text, avoiding model illusions, and yielding a semantic credibility assessment result. For key text information on images in the document, a local RapidOCR optical character recognition model is used to identify and extract the text information from the images. This information is then compared with the text content of the corresponding chapters to obtain an image consistency assessment result. Finally, the semantic quality assessment result, semantic credibility assessment result, and image consistency assessment result are summarized to obtain the semantic review result.
[0161] For example, semantic auditing can identify logical contradictions in construction technology summaries, missing content in descriptions of special technologies, textual errors in appendix descriptions, and inconsistencies between images and text.
[0162] S205, Integrating rule verification results with semantic audit results.
[0163] Among them, the integration is used to combine the results of rigid rule verification with the results of flexible semantic review. Its role is to achieve collaborative review with "basic rules as a safety net and AI review as a backup", ensuring that the review results have neither rule omissions nor semantic blind spots.
[0164] In this step, the rule verification results output by the rule verification module and the semantic review results output by the semantic review module can be uniformly aggregated through the fusion module, establishing a correlation mapping between the two types of results to form a complete intermediate review dataset.
[0165] Specifically, the fusion module receives rule verification results and semantic audit results, and performs deduplication and correlation on audit findings for the same chapter or the same field to ensure that hard data errors found by rule verification and soft text issues found by semantic audit can be presented in the same view, avoiding fragmentation of audit results.
[0166] For example, the merged intermediate dataset can include categorized indexes such as structural integrity audit items, data consistency verification items, format standardization issues items, text semantic quality assessment items, and image consistency comparison items.
[0167] S206. The rule verification results and semantic review results shall be classified and categorized in layers according to structural integrity, data consistency, format standardization, and text semantic quality.
[0168] Among them, hierarchical classification is used to organize the two types of audit results according to a unified dimension. Its function is to make the audit report structure clear and facilitate auditors to quickly locate the problem type.
[0169] In this step, the report generation module can be used to classify the integrated review results into four levels according to the four preset review dimensions: structural integrity, data consistency, format standardization, and text semantic quality.
[0170] Specifically, the structural integrity level includes architectural issues such as missing chapters and mismatched directories; the data consistency level includes logical issues such as cross-chapter data conflicts and numerical discrepancies; the formatting compliance level includes rule-related issues such as null values in fields, formatting errors, and inconsistent units; and the text semantic quality level includes semantic issues such as non-standard expressions, missing content, logical contradictions, and typographical errors. Each level is independently numbered, forming a tree-like structure for the review results.
[0171] For example, the audit results after hierarchical classification can be organized in a three-level directory format of "first-level dimension → second-level sub-item → specific problem description", which makes it easy for auditors to review the results layer by layer according to priority.
[0172] S207. Automatically number and list abnormal issues, and mark compliant content.
[0173] Among them, the list of abnormal issues is used to clearly present non-compliant items found in the audit in the form of numbers, while the compliance mark is used to clearly mark the content that has passed the audit. Its purpose is to improve the readability and traceability of the audit results.
[0174] In this step, the report generation module can traverse the review results at each level, automatically assign incremental numbers to the issues that are judged to be abnormal, and mark the content that is judged to be compliant with the "pass" or "compliant" mark.
[0175] Specifically, abnormal issues are numbered according to their level prefix plus a three-digit serial number (e.g., "Data Consistency-001", "Text Semantic Quality-005"). Each abnormal record includes a problem description, the relevant chapters, the relevant fields, the basis for judgment, and the suggested direction of correction. Compliant content is marked with a green icon or a "√" symbol, which contrasts sharply with abnormal issues and makes it easy to distinguish them quickly.
[0176] For example, the list of abnormal issues may include items such as "the well number is inconsistent between the cover and the first chapter of the main text (data consistency-001)" and "the construction technology summary lacks a description of the directional drilling point (text semantic quality-003)"; compliance indicators may include items such as "the cover information is complete (√)" and "the well depth data is continuous (√)".
[0177] S208, Output the drilling and completion report review results.
[0178] Among them, the drilling and completion report review results are used as a standardized document to summarize the review conclusions of the entire process. Its function is to realize the visualization, traceability and archiving of the review results.
[0179] In this step, the save_report function can be written to standardize and summarize the review results of the entire system process and output them in a documented form, ultimately generating a well-structured, complete, and highly readable well completion report review result document.
[0180] Specifically, the audit results document sequentially integrates the document structure audit results, rule engine verification issues, and AI supplementary audit opinions, and writes them into a text file in a fixed format, clearly presenting four categories of audit content: structural integrity, data consistency, format standardization, and text semantic quality; abnormal issues are automatically numbered and listed, and compliant content is clearly marked, realizing the visualization, traceability, and archiving of audit results.
[0181] For example, the audit results document can be output in the form of a text file or a structured report, including the structural integrity audit conclusion, data consistency verification details, a list of format standardization issues, text semantic quality assessment and image consistency comparison results, etc.
[0182] The drilling and completion report review method provided in this application, through the bidirectional collaboration of document structured parsing, rigid verification by the rule engine, and semantic review by the local lightweight model, as well as the full-process local offline operation, eliminates the need for manual item-by-item verification in the review of drilling and completion reports. This avoids the problems of low efficiency and easy omissions in traditional manual review, and also avoids the technical bottlenecks of traditional automated tools being unable to adapt to complex business logic and confidential documents being unable to be uploaded to the cloud. This achieves the effects of improving review efficiency and accuracy, reducing hardware dependence, and ensuring zero out-of-domain transmission of confidential data.
[0183] Figure 3 Flowchart of the drilling and completion report review method provided for this application Figure 2 ,like Figure 3 As shown, based on the above embodiments, the method provided in this embodiment may include the following steps:
[0184] S301. Traverse all contents of the drilling and completion report documents pending review.
[0185] The traversal is used to enumerate and access elements such as paragraphs, tables, and images in a Word document one by one. Its purpose is to ensure that no document content is missed, providing a complete raw data foundation for subsequent title recognition and chapter division.
[0186] In this step, a Python script can be used to call the python-docx third-party library to open the drilling and completion report document to be reviewed and read all the content of the document segment by segment.
[0187] Specifically, the system scans the document sequentially from the first paragraph, determining if the current paragraph is a cover page or table of contents. If not, it extracts the paragraph's text content, style attributes, and its position index within the document. For table elements, it reads the cell content row by row and column by column, recording the table's position. For image elements, it records their position index within the document for subsequent OCR recognition. During the traversal, the content of the cover page and table of contents is skipped to avoid interfering with subsequent chapter boundary determination.
[0188] For example, the traversal results may include data structures such as a list of paragraph texts, a list of paragraph style names, a table content matrix, and an image position index table.
[0189] S302. Identify the heading level information of each paragraph.
[0190] The heading level information is used to identify whether a paragraph belongs to a first-level heading, a second-level heading, or a third-level heading. Its function is to distinguish between the main body of the chapter and the main text, and to provide a classification basis for the subsequent weighted score calculation of heading levels.
[0191] In this step, the heading level information of each paragraph can be determined by parsing the paragraph style name or text format features.
[0192] Specifically, the system reads the style name of each paragraph (e.g., "Heading 1", "Heading 2", "Heading 3"). If the style name matches a level 1 heading style, the paragraph is marked as a level 1 heading; if it matches a level 2 heading style, it is marked as a level 2 heading; if it matches a level 3 heading style, it is marked as a level 3 heading; if the style name does not match any heading style, the paragraph is marked as body text. For paragraphs that do not use standard styles but have manual formatting, regular expressions can be used to match heading features (e.g., the prefix "1.", "1.1", "1.1.1") to assist in identification.
[0193] For example, heading level information can be tagged as enumeration values such as L1 (first-level heading), L2 (second-level heading), L3 (third-level heading), or Body (body text).
[0194] S303. Based on the heading level information of each paragraph, calculate the heading level score for each paragraph.
[0195] Among them, the heading level score is used to quantitatively represent the significance of paragraphs as chapter boundary markers. Its role is to combine heading level attributes and weight coefficients to provide a comparable numerical basis for chapter boundary determination.
[0196] In this step, a title level weighted recognition algorithm can be used to calculate the title level score for each paragraph obtained through the traversal.
[0197] Specifically, the calculation formula for the title level weighted recognition algorithm is as follows:
[0198] Score = α·L1+β·L2+γ·L3
[0199] Where L1, L2, and L3 represent the matching results for level 1, level 2, and level 3 headings (1 for a match, 0 for a miss), and α, β, and γ are the level weight coefficients, with α > β > γ. For each paragraph, a score is calculated by substituting its heading level information into the corresponding formula: level 1 heading paragraphs score α, level 2 heading paragraphs score β, level 3 heading paragraphs score γ, and body text paragraphs score 0.
[0200] For example, the hierarchy weight coefficients can be set to α=3, β=2, γ=1. In this case, the first-level heading paragraph scores 3, the second-level heading paragraph scores 2, and the third-level heading paragraph scores 1.
[0201] S304. Assign first-level weight to paragraphs whose heading level information is first-level heading.
[0202] The first-level weight is used to represent the highest priority of the first-level heading in the chapter boundary determination. Its role is to ensure the authoritative status of the first-level heading as the top-level chapter boundary marker and to avoid misjudging lower-level headings as the main chapter boundary.
[0203] In this step, the weight coefficient α corresponding to the first-level title can be set to the highest value in the title level weighted recognition algorithm.
[0204] Specifically, when calculating the heading level score, for paragraphs identified as first-level headings, let L1=1, L2=0, L3=0, and substitute them into the formula Score=α·L1+β·L2+γ·L3 to obtain the heading level score of the paragraph as α.
[0205] For example, a first-level heading can correspond to a top-level chapter heading such as "Chapter 1 Overview of Drilling Engineering", and its first-level weight α can be set to 3.
[0206] S305. Assign second-level weights to paragraphs whose heading level information is second-level headings.
[0207] The second-level weight is used to characterize the secondary priority of the second-level heading in the chapter boundary determination. Its function is to identify the sub-chapter boundaries under the first-level heading and realize the gradual refinement of the chapter level.
[0208] In this step, the weight coefficient β corresponding to the second-level heading can be set to the second-highest value in the heading level weighted recognition algorithm, and β < α can be satisfied.
[0209] Specifically, when calculating the heading level score, for paragraphs identified as second-level headings, let L1=0, L2=1, L3=0, and substitute them into the formula Score=α·L1+β·L2+γ·L3 to obtain the heading level score of the paragraph as β.
[0210] For example, a second-level heading can correspond to a sub-section heading such as "1.1 Drilling Equipment Selection", and its second-level weight β can be set to 2.
[0211] S306. Assign third-level weights to paragraphs whose heading level information is third-level headings.
[0212] The third-level weight is used to characterize the basic priority of the third-level heading in the chapter boundary determination. Its function is to identify the subdivisions of the second-level heading and support more refined chapter structure analysis.
[0213] In this step, the weight coefficient γ corresponding to the third-level heading can be set to the lowest value in the heading level weighted recognition algorithm, and γ < β < α can be satisfied.
[0214] Specifically, when calculating the heading level score, for paragraphs identified as third-level headings, let L1=0, L2=0, L3=1, and substitute them into the formula Score=α·L1+β·L2+γ·L3 to obtain the heading level score of the paragraph as γ.
[0215] For example, a third-level heading can correspond to a subdivided section heading such as "1.1.1 Drilling Rig Model Parameters", and its third-level weight γ can be set to 1.
[0216] S307. Determine chapter boundaries based on title level scores.
[0217] The chapter boundaries are used to identify the dividing positions between adjacent chapters. Their function is to divide continuous document content into independent chapter units based on the heading level score, so as to facilitate group review and data extraction.
[0218] In this step, chapter boundaries can be determined by using a preset scoring threshold or by detecting abrupt changes in scores between adjacent paragraphs.
[0219] Specifically, the system iterates through the heading level scores of each paragraph. When it detects that the current paragraph has a score greater than 0 (i.e., the paragraph is a heading) and the score is significantly higher than that of the adjacent paragraphs, it marks the position of the paragraph as the beginning boundary of the chapter. When the next heading paragraph is detected, the current position is marked as the end boundary of the previous chapter and the beginning boundary of the next chapter. For adjacent heading paragraphs with the same score, the nesting relationship is determined according to the level to ensure that the chapter levels do not overlap.
[0220] For example, the chapter boundary determination result may include fields such as chapter start index, chapter end index, chapter title text, and chapter level depth.
[0221] S308. Group the drilling and completion report documents to be reviewed by chapter based on chapter boundaries and extract the main text content.
[0222] Among them, grouping by chapter is used to divide the document content into independent chapter units based on the determined chapter boundaries, and extracting the main text content is used to obtain the substantive text and data within each chapter. Its function is to form a structured chapter dataset, providing organized input for subsequent business rule verification and semantic review.
[0223] In this step, document paragraphs can be sliced and grouped using chapter boundary indexes to extract the main text paragraphs and table data within each chapter.
[0224] Specifically, based on the chapter boundaries obtained from S307, the system categorizes paragraphs located between the two boundaries into the same chapter; for paragraphs within each chapter, it filters out the title paragraphs themselves and extracts the main text paragraphs; for tables within a chapter, it extracts the table content and converts it into a data structure (such as JSON or CSV format); and it organizes the scattered subsection content into a structured and merged chapter data packages.
[0225] For example, the extracted chapter data after grouping can include fields such as chapter title, chapter level, list of main text paragraphs, list of table data, and list of subsection information.
[0226] S309. Generate structured text.
[0227] Structured text is used to convert unstructured content in the original Word document into standardized data that is machine-readable, hierarchical, and uniformly formatted. Its role is to eliminate the interference of the original document's layout differences on subsequent review and to provide standardized input for rule verification and semantic review.
[0228] In this step, the extracted data, grouped by chapter, can be converted into structured text in Markdown format.
[0229] Specifically, the system extracts the titles, body paragraphs, data tables, and subsection information of each chapter in sequence, and combines and organizes them according to a fixed format: chapter titles are marked with the "#" symbol to indicate the level, body paragraphs remain unchanged, table data is converted into Markdown table syntax, and subsection information is organized in a nested hierarchical form; the discrete structured data is restored into continuous Markdown text to form a structured text containing complete chapter hierarchical information.
[0230] For example, the generated structured text may contain content such as "## Chapter 1 Overview of Drilling Engineering\n\nMain Text Paragraphs...\n\n| Well Depth (m) | Drilling Pressure (kN) | ...\n\n### 1.1 Drilling Equipment Selection\n..."
[0231] The technical solution in this embodiment, by traversing the entire document content and identifying paragraph heading hierarchy information, employs a heading hierarchy weighted recognition algorithm to assign differentiated weight coefficients to different levels. Based on the scores, it accurately determines chapter boundaries and extracts content by chapter grouping. This significantly improves the parsing accuracy of complex formatted documents. Chapter boundary determination has quantifiable numerical basis, avoiding the chapter misjudgment and boundary omission problems that are prone to occur when parsing based on fixed styles or regular expressions. Thus, the document structured parsing process has complete hierarchy recognition capabilities, accurate boundary determination results, and reliable chapter division quality. At the same time, by converting the original Word document into Markdown structured text, it provides standardized data with unified format and clear boundaries for subsequent business rule verification, and also provides standardized input with controllable length and clear structure for local lightweight model inference. This ensures that the document structured parsing results can stably support the strict execution of subsequent data consistency verification and semantic review, meeting the business requirements of drilling and completion report review for document structural integrity and review result accuracy.
[0232] Figure 4 Flowchart of the drilling and completion report review method provided for this application Figure 3 ,like Figure 4 As shown, based on the above embodiments, the method provided in this embodiment may include the following steps:
[0233] S401, Load structured text.
[0234] Among them, structured text is used to carry standardized chapter data after being converted by the document parsing module. Its role is to serve as the input data source for business rule verification, ensuring that the verification logic is executed based on data with uniform format and clear boundaries.
[0235] In this step, the Markdown-formatted structured text file generated by the document parsing module can be read and loaded into memory for use by the rule engine.
[0236] Specifically, the system reads a structured text file from a specified path, parses its chapter hierarchy, main paragraphs, data tables, and subsection information, and constructs a structured data object in memory, providing an accessible data foundation for subsequent various validation rules.
[0237] For example, structured text may contain content such as "## Chapter 1 Overview of Drilling Engineering\n\nMain Text Paragraphs...\n\n| Well Depth (m) | Drilling Pressure (kN) | ...".
[0238] S402. Perform a NOT NULL check on the key fields in the structured text.
[0239] Key fields are used to identify essential information items that must be filled in the drilling and completion report. Their purpose is to ensure the completeness of the necessary information in the report and prevent subsequent data verification or business analysis from failing due to missing key information. The NOT NULL check is used to check whether the field value is an empty set and is a fundamental method for integrity verification.
[0240] In this step, the rule engine can be used to check each of the predefined required fields in the structured text, with the condition that Field ≠ Empty set.
[0241] Specifically, the system iterates through the key fields in the structured text (such as hash, author, reviewer, completion date, etc.) and checks whether the field value is an empty string, None, or contains only blank characters. If the field value is empty, it is determined that key information is missing and the field integrity is recorded as abnormal. If the field value is not empty, it is determined that the field passes the integrity check.
[0242] For example, key fields may include well number, well name, writer, reviewer, well completion date, drilling team number, etc., and the non-empty judgment result can be marked as "pass" or "missing: well number is empty".
[0243] S403. Perform format pattern matching on key data in structured text.
[0244] Key data is used to identify numerical values or identifiers in drilling and completion reports that need to conform to specific format specifications. Its purpose is to ensure data format standardization, facilitating subsequent numerical calculations and cross-section comparisons. Format pattern matching, used to perform pattern validation on text through regular expressions, is the core method for standardization verification.
[0245] In this step, the rule engine can perform regular expression matching on the key data in the structured text, with the judgment formula being Text ∈ RegPattern.
[0246] Specifically, the system uses corresponding regular expression patterns to match different types of key data, such as hash symbols, numbers, units, and dates: hash symbols match preset hash symbol encoding rules (e.g., "^[AZ]{2}\d{4}"), numbers match numerical formats (e.g., "^\\d+\\.?\\d*"), units match standard engineering units (e.g., "m", "kN", "MPa"), and dates match date formats (e.g., "YYYY-MM-DD"). If the text does not match the regular expression pattern, it is determined that the format is not standard and the record shows an abnormal format.
[0247] For example, key data may include well number "XX-1234", well depth "3500.5m", date "2024-03-15", etc. The format pattern matching result can be marked as "pass" or "format error: well depth '3500.5' is missing units".
[0248] S404. Perform keyword retrieval on standard engineering terms in structured text.
[0249] Standard engineering terminology is used to identify industry-standard terms or key technical terms that should appear in drilling and completion reports. Its purpose is to ensure the professionalism and completeness of the report content and prevent the omission of technical content due to missing terminology. Keyword retrieval is used to search for the presence of standard terms in the text, serving as a method for verifying content completeness.
[0250] In this step, the rule engine can be used to perform keyword retrieval on standard engineering terms in the structured text, with the judgment formula being Keyword ⊂ Text.
[0251] Specifically, the system loads a pre-set standard engineering terminology database (such as "build-up point", "target point", "drilled layer", "cement return height", "cementing quality", etc.) and searches each term in each chapter of the structured text to see if it appears. If a term is not found in the corresponding chapter, it is determined that the content is missing and the record is abnormal. If a term is found, it is determined that the technical content corresponding to the term has been covered.
[0252] For example, standard engineering terms may include "start-up point", "target point", "wellbore structure", "cementing process", etc. Keyword search results can be marked as "hit" or "missing: 'start-up point' did not appear in Chapter 3".
[0253] S405. Perform a consistency comparison on cross-chapter related data in structured text.
[0254] Cross-chapter related data is used to identify the same key parameter that appears repeatedly in multiple chapters or tables of the drilling and completion report. Its purpose is to ensure the logical consistency of the data throughout the report and prevent the same parameter from appearing in different locations due to data entry errors or version confusion. Consistency comparison is used to verify the numerical differences of parameters across chapters, and is a method for verifying the logical consistency of data.
[0255] In this step, the rule engine can be used to compare the values of the same type of parameters distributed in different chapters or tables, and the judgment formula is |Data1 − Data2| ≤ ε.
[0256] Specifically, the system extracts key parameters (such as well depth, start-up point depth, target coordinates, and completed well depth) across chapters in the structured text and compares their values in different chapters or tables. If the absolute value of the difference between the two values exceeds the preset error threshold ε, the data is determined to be inconsistent and an abnormal data consistency is recorded. If it is within the error range, the data is determined to be consistent.
[0257] For example, cross-chapter related data may include a well depth recorded as "3500.5m" in Chapter 1 "Drilling Engineering Overview" and "3500.5m" in Chapter 5 "Completion Data Table". The error threshold ε can be set to 0.1m, and the comparison result can be marked as "consistent" or "inconsistent: the difference of 0.3m between the well depth of 3500.5m in Chapter 1 and the well depth of 3500.8m in Chapter 5 exceeds the threshold".
[0258] S406. Based on the preset standard chapter directory list, traverse and match the actual chapter titles in the structured text.
[0259] The preset standard chapter list defines the standard chapter structure required by the drilling and completion report specifications, serving as a benchmark template for document structure integrity verification. Traversal matching compares the actual identified chapter titles in the document with the preset standard chapters one by one, representing a method for structural integrity verification.
[0260] In this step, the actual chapter titles in the structured text can be verified by using a chapter matching degree and integrity scoring model through a rule engine.
[0261] Specifically, the system loads a pre-built list of standard chapter headings for drilling and completion reports (such as "Chapter 1: Overview of Drilling Engineering", "Chapter 2: Drilling Equipment and Tools", "Chapter 3: Drilling Technology", etc.), and matches the actual parsed chapter titles of the document with the standard list one by one; it then counts the number N of standard chapters that are actually matched. match Total number of chapters N in the standard template total Calculate the chapter matching degree S = (N match / N total ) × 100%; when S < 100%, it is determined that a necessary chapter is missing, and the record structure integrity is abnormal.
[0262] For example, the preset standard chapter list can contain 10 standard chapters, and the actual document matches 9 of them, then N match =9、N total =10, S=90%, the judgment result is "Missing necessary chapter: 'Chapter 10 Appendix' is missing".
[0263] S407. Summarize the results of field integrity judgment, format standardization judgment, content integrity judgment, data consistency judgment, and structural integrity judgment to obtain the rule verification results.
[0264] Among them, the rule verification results are used to summarize the individual judgment conclusions of the entire business rule verification process. Its role is to form a complete rigid audit report, providing a rule-level audit basis for subsequent integration with semantic audit results.
[0265] In this step, the judgment results generated from steps S402 to S406 can be uniformly aggregated and formatted through the rule engine.
[0266] Specifically, the system collects the results of field integrity assessment, format standardization assessment, content integrity assessment, data consistency assessment, and structural integrity assessment, and summarizes them according to a preset report template: each type of assessment result is classified and numbered, and the problem description, relevant chapters, relevant fields, and assessment basis are recorded; all types of results are integrated into a structured rule verification result dataset for the fusion module to call.
[0267] For example, the rule validation results may include summary information such as "Field integrity: 2 exceptions", "Format compliance: 1 exception", "Content integrity: 0 exceptions", "Data consistency: 1 exception", and "Structure integrity: 1 exception".
[0268] S408, Receive natural language rules input by the user.
[0269] Natural language rules are used to describe new business review requirements in human-readable language. Their purpose is to lower the barrier to entry for user-defined rules, enabling non-technical personnel to participate in rule maintenance. User input is used to receive descriptions of newly added or changed review rules from external sources.
[0270] In this step, the user can input the text of the new review rule through a natural language interactive interface.
[0271] Specifically, the system provides a text input box or command line interface to receive rules described by the user in natural language (such as "check if the well depth is greater than 3000 meters" or "ensure that each chapter contains safety prompts"), performs basic cleaning and preprocessing on the input text, and extracts key entities and constraints from the rule description.
[0272] For example, the natural language rule input by the user could be "If the well type is a horizontal well, then build-up point data must be included".
[0273] S409. Parse natural language rules into executable verification logic.
[0274] Among them, the executable verification logic is used to convert business rules described in human language into code or configuration that the rule engine can directly execute. Its role is to realize the automatic translation from natural language to machine language and support the zero-code dynamic expansion of rules.
[0275] In this step, the natural language rules can be semantically parsed and logically transformed by calling a local lightweight language model.
[0276] Specifically, the system sends the natural language rule text input by the user to a locally deployed lightweight language model (such as qwen3.5:9b). The model is guided to parse the validation objects, validation conditions and decision logic in the rules through preset prompt word templates. The model outputs structured executable validation logic (such as Python function fragments or rule configuration JSON). The system performs syntax validation and security checks on the output to ensure that the executable logic is free from injection risks and conforms to the rule engine interface specification.
[0277] For example, the natural language rule "check if the well depth is greater than 3000 meters" can be parsed into executable validation logic: {"field": "well depth", "operator": ">", "threshold": 3000, "unit": "m"}.
[0278] S410. Add the executable verification logic to the preset audit rule library.
[0279] The preset audit rule base stores all business verification rules, both built into the system and those extended by users, serving as the authoritative basis for the rule engine to perform audits. The append operation dynamically adds newly parsed rules to the rule base, enabling the rules to take effect immediately.
[0280] In this step, the executable verification logic generated by S409 can be written into the rule base file or the memory rule table through the rule management module.
[0281] Specifically, the system uniquely identifies and encodes the executable verification logic, checks for conflicts with existing rules (such as duplicate fields or logical contradictions); if there are no conflicts, the new rule is inserted into the corresponding category node of the rule base, and the rule base version number and effective timestamp are updated; the rule engine automatically loads the updated rule base in subsequent audit tasks, so that the new rule takes effect immediately.
[0282] For example, executable validation logic can be appended as a Python function to the list of extended rules in rule_engine.py, or as a JSON configuration file to rules.json.
[0283] The technical solution in this embodiment performs five types of rigid checks on structured text: non-empty check, format pattern matching, keyword retrieval, cross-chapter consistency comparison, and standard chapter directory traversal matching. The results are then aggregated to obtain rule verification results, enabling a comprehensive quantitative review of the data accuracy, format compliance, and structural integrity of drilling and completion reports. This avoids oversights and biases that are prone to occur in traditional manual reviews. Simultaneously, by receiving user-input natural language rules and parsing them into executable verification logic using a local lightweight language model, these rules are appended to the rule base. This allows for the maintenance of business rules without modifying the underlying code, and enables rapid response to personalized review needs for specific oilfields or well types. Therefore, the business rule verification process possesses complete rigid verification capabilities, a flexible rule expansion mechanism, and high rule maintenance efficiency. Furthermore, by integrating natural language rule parsing and rule base updates into a local offline environment, data security during rule expansion is ensured, meeting the rigid requirements of classified document management.
[0284] Figure 5 Flowchart of the drilling and completion report review method provided for this application Figure 4 ,like Figure 5 As shown, based on the above embodiments, the method provided in this embodiment may include the following steps:
[0285] S501. Select chapters suitable for semantic review.
[0286] Among them, the chapters suitable for semantic review are used to identify content units in drilling and completion reports that are mainly unstructured text descriptions and are suitable for natural language understanding and quality assessment. Their function is to exclude elements such as data tables, covers, and tables of contents that are not suitable for semantic review, concentrate computing resources on text quality review, and improve review efficiency and targeting.
[0287] In this step, the semantic auditing module can be used to filter target chapters from structured text, include unstructured text chapters in the semantic auditing scope, and exclude structured or fixed-format elements such as data tables, cover pages, and table of contents.
[0288] Specifically, the system iterates through each chapter in the structured text and determines whether a chapter belongs to the unstructured text type based on the chapter title keyword matching rules (such as "Construction Technology Summary", "Special Technology Summary", "Attachment Description", "Process Description", "Problem Analysis", etc.). If a chapter is mainly a paragraph description and contains technical discussions or experience summaries, it is marked as a target chapter suitable for semantic review. If a chapter is a pure data table, equipment list, or cover information, it is marked as skipping semantic review and only rigidly verified by the rule engine.
[0289] For example, chapters suitable for semantic review may include "Chapter 3 Summary of Construction Technology", "Chapter 5 Application of Featured Technologies", "Appendix 1 Technical Description", etc.; chapters unsuitable for semantic review may include "Cover", "Table of Contents", "Chapter 2 Drilling Equipment Parameter Table", etc.
[0290] S502. Organize the chapter content into a standard format.
[0291] The standard format is used to restore discrete structured chapter data into continuous, readable, and uniformly formatted plain text. Its function is to form a standardized input suitable for lightweight language model understanding and processing, and to eliminate the interference of Markdown tags or data structures on model parsing.
[0292] In this step, the structured chapter data obtained from the previous parsing can be reassembled into complete text using the build_chapter_text function.
[0293] Specifically, the function extracts the chapter title, main paragraphs, data tables, and subsection information of the target chapter in sequence, and combines and organizes them according to a fixed format: the chapter title begins with a prominent logo, the main paragraphs are arranged in a logical order, the table data is converted into text descriptions, and the subsection information is organized in a hierarchical indentation manner; the discrete structured data is restored into continuous text, forming standardized text materials suitable for lightweight model processing.
[0294] For example, the standard format output could be "[Chapter Title: Construction Technology Summary]\n\nParagraph 1...\n\nParagraph 2...\n\n[Subsection: Directional Drilling Technology]\n...".
[0295] S503, call the local lightweight model chapter by chapter to conduct semantic auditing.
[0296] The local lightweight model is used to perform natural language understanding and text quality assessment on local computing devices. Its function is to check the text's expressive standardization, content completeness, logical rationality, and typographical correctness, achieving flexible review that is difficult to cover with fixed rules. Semantic review is used to determine the quality of unstructured text.
[0297] In this step, standardized request parameters can be constructed by writing the call_ollama function to call the locally deployed lightweight large model service chapter by chapter.
[0298] Specifically, the function specifies the audit model (such as the qwen3.5:9b model deployed through the Ollam framework), context window length, generation parameters (temperature parameter temperature=0.2), and other configurations. It sends the standard format text compiled by S502 along with professional audit instructions (clearly defining the audit scope, requirements, and output format) to the model. The model analyzes the text chapter by chapter, focusing on checking whether the expression is standardized, the content is complete, the logic is reasonable, and the text is correct. It collects the audit comments returned by each chapter and forms a semantic quality judgment result.
[0299] For example, semantic auditing can identify problems such as "the construction technology summary lacks specific descriptions of well control measures" and "the paragraphs of the special technology summary are logically disordered, describing the results before the reasons."
[0300] S504. Calculate the ratio of the effective text length of each target chapter to the total length of the chapter.
[0301] Among them, the effective text length is used to represent the number of substantive content characters in the chapter, the total chapter length is used to represent the total number of characters in the chapter, and the ratio is used to quantify the richness and information density of the text content. Its function is to provide basic data for content confidence calculation and to identify empty or filler text.
[0302] In this step, the length of the original text of the target chapter can be calculated after cleaning it using the text statistics module.
[0303] Specifically, the system reads the complete text of the target chapter and counts the total number of characters to obtain the total chapter length L. all The effective text length L is obtained by filtering out whitespace characters, meaningless placeholders (such as templated phrases like "to be supplemented" and "see attachment"), repeated punctuation, and formatting marks, and then counting the remaining substantive content characters. eff ; Calculate the ratio R = L eff / L all .
[0304] For example, if the total length of a chapter is L all 500 characters, effective text length L eff If the total length of another chapter is 450 characters, the ratio R is 0.9; if the total length of another chapter is 500 characters but the effective text length is only 100 characters, the ratio R is 0.2, indicating that the content of that chapter is empty.
[0305] S505. Use the local optical character recognition model to extract text information from images in the drilling and completion report document to be reviewed.
[0306] Among them, the local optical character recognition model is used to recognize the text content in document images in a local offline environment. Its function is to extract key data or descriptive information from the image so that it can be compared with the text content of the corresponding chapter and discover inconsistencies between the image and the text.
[0307] In this step, the text can be extracted from the images in the document by calling the RapidOCR local optical character recognition engine.
[0308] Specifically, the system iterates through all image elements in the drilling and completion report document to be reviewed, and calls the RapidOCR model to perform text recognition for each image; the model returns the text content, confidence level, and location information in the image; the system records the correlation between the recognition results and the chapter in which the image is located, providing a data source for subsequent consistency comparison.
[0309] For example, in the document " Figure 3-1 The "well structure diagram" may contain the text label "start point: 1850m", which is extracted by RapidOCR to obtain the text information "start point: 1850m".
[0310] S506. Determine the content confidence level of each target chapter based on the ratio and content quality coefficient.
[0311] Content confidence is a quantitative indicator used to measure the reliability of semantic review results. Its function is to distinguish between vague, missing, and unqualified text, avoid misjudgments caused by model illusions, and improve the credibility of semantic review results.
[0312] In this step, the content confidence level of each chapter can be calculated using the content confidence level determination mechanism, based on the ratio obtained in S504 and the preset content quality coefficient.
[0313] Specifically, the calculation formula is as follows:
[0314] C = (L eff / L all ) × λ
[0315] Among them, L eff For the effective text length, L all λ represents the total length of the chapter, and λ is the content quality coefficient. The system presets the λ value based on the chapter type and industry standards (e.g., λ=1.0 for technical summary chapters, λ=0.8 for appendix description chapters); it substitutes the ratio calculated by S504 into the formula to obtain the content confidence score C; if C is lower than the preset threshold, the content quality of the chapter is judged to be unqualified or too vague, and its semantic review result weight is reduced or it is marked as requiring manual review.
[0316] For example, if a certain chapter has a ratio R=0.9 and a content quality coefficient λ=1.0, then the content confidence level C=0.9; if the ratio R=0.3 and λ=1.0, then C=0.3, and the chapter is judged to have vague content and may have missing content.
[0317] S507. Perform a consistency comparison between the text information and the text content of the corresponding chapter.
[0318] Among them, consistency comparison is used to verify whether the text information extracted from the image matches the text description of the same chapter. Its purpose is to discover problems such as contradictions between the text and image descriptions, incorrect data labeling, or improper image referencing.
[0319] In this step, the image text information extracted by S505 can be compared with the text content of the corresponding chapter through string matching or similarity calculation.
[0320] Specifically, the system locates the chapter where the image is located and extracts descriptive statements related to the image from the text of that chapter (such as "as..."). Figure 3-1 As shown, the slope is located at 1850m). The OCR recognition result of the image, "Slope: 1850m", is compared with the text description. If the values or descriptions are consistent, the image is considered to be consistent. If there is a difference (e.g., the text says 1850m but the image says 1800m), the image is considered to be inconsistent, and the image consistency judgment result is recorded.
[0321] For example, if the image text "Ascending point: 1850m" is consistent with the chapter text "Ascending point depth is 1850m", it is judged as "passed"; if the chapter text is "Ascending point depth is 1800m", it is judged as "inconsistent: the text 1800m contradicts the image 1850m".
[0322] S508. Summarize the semantic quality judgment results, semantic credibility judgment results, and image consistency judgment results to obtain the semantic audit results.
[0323] The semantic audit results are used to summarize the individual judgment conclusions of the entire semantic audit process. Their role is to form a complete flexible audit report, providing a semantic audit basis for subsequent integration with rule verification results.
[0324] In this step, the judgment results generated by steps S503, S506, and S507 can be uniformly aggregated and formatted through the semantic auditing module.
[0325] Specifically, the system collects semantic quality judgment results (standardization of expression, completeness of content, logical rationality, and correctness of text), semantic credibility judgment results (content confidence C and whether it is vague), and image consistency judgment results (whether the text and images are consistent). The results are then summarized according to a preset report template: each type of judgment result is classified and numbered, and the problem description, relevant chapters, relevant images, and judgment basis are recorded. All types of results are integrated into a structured semantic audit result dataset for use by the fusion module.
[0326] For example, the semantic audit results may include summary information such as "Semantic quality: 2 abnormalities", "Semantic credibility: 1 item needs to be reviewed", and "Image consistency: 1 inconsistency".
[0327] S509. Output semantic audit results.
[0328] The output semantic audit result is used to submit the final conclusion of the semantic audit module to the fusion module. Its function is to complete the closed loop of the semantic audit process, so that the flexible audit conclusion can participate in the generation of the final audit report.
[0329] In this step, the semantic audit result dataset can be sent to the fusion module through function return or message passing mechanisms.
[0330] Specifically, the system outputs the structured semantic audit results generated by S508 in a standard data format (such as JSON or dictionary objects), including audit category, issue list, confidence score and suggested correction direction; after receiving the results, the fusion module performs fusion processing with the rule verification results to generate the final drilling and completion report audit results.
[0331] For example, the output semantic audit results may contain structured data such as {"semantic_quality": [...], "confidence": [...], "image_consistency": [...]}.
[0332] The technical solution of this embodiment selects chapters suitable for semantic review and organizes them into a standard format. It then calls a local lightweight model to conduct semantic review chapter by chapter. Simultaneously, it calculates content confidence by combining the effective text length ratio and content quality coefficient to suppress model illusions. Furthermore, it calls a local optical character recognition model to extract image and text information for consistency comparison. This ensures that the unstructured text content in the drilling and completion report undergoes comprehensive and flexible quality review, accurately verifying the correspondence between images and text. This avoids the problems of traditional rule engines failing to identify semantic logic errors, missing content, and contradictions between images and text. Moreover, by running the entire process locally offline, it ensures that the image and text information in the document to be reviewed does not leave the domain. This satisfies the efficiency requirements of intelligent review while strictly adhering to the rigid requirements of classified document management. Therefore, the semantic review process possesses complete text quality assessment capabilities, a reliable confidence determination mechanism, and secure image-text consistency verification results.
[0333] Figure 6 A schematic diagram of the drilling and completion report review device provided in this application; as shown in the figure, the drilling and completion report review device 60 provided in this embodiment includes:
[0334] Module 601 is used to retrieve drilling and completion report documents pending review.
[0335] The parsing module 602 is used to perform structured parsing on the drilling and completion report document to be reviewed, and obtain structured text containing chapter and hierarchical information.
[0336] The rule validation module 603 is used to perform business rule validation on structured text and obtain the rule validation results.
[0337] The semantic auditing module 604 is used to perform semantic auditing on structured text and obtain the semantic auditing results.
[0338] The fusion module 605 is used to fuse rule verification results and semantic review results to generate drilling and completion report review results.
[0339] In one possible implementation, the target chapter shall include at least one or more of the following: a summary of construction techniques, a summary of special techniques, and explanatory notes.
[0340] In one possible implementation, the parsing module 602 is specifically used for:
[0341] The entire content of the drilling and completion report document to be reviewed is traversed to identify the heading level information of each paragraph. Based on the heading level information of each paragraph, the heading level score of each paragraph is calculated. Paragraphs with heading level information of level 1 are assigned first-level weight, paragraphs with heading level information of level 2 are assigned second-level weight, and paragraphs with heading level information of level 3 are assigned third-level weight. The first-level weight is greater than the second-level weight, and the second-level weight is greater than the third-level weight. The chapter boundaries are determined according to the heading level scores. Based on the chapter boundaries, the drilling and completion report document to be reviewed is grouped by chapter and the main text content is extracted to generate structured text.
[0342] In one possible implementation, the rule verification module 603 is specifically used for:
[0343] The following steps are performed to determine the integrity of a structured text: First, perform NOT NULL checks on key fields to obtain field integrity results. Second, perform format pattern matching on key data to obtain format conformity results. Third, perform keyword retrieval on standard engineering terms in the structured text to obtain content integrity results. Fourth, perform consistency comparison on cross-chapter related data in the structured text to obtain data consistency results. Fifth, perform traversal matching on actual chapter titles in the structured text based on a preset standard chapter directory list to obtain structural integrity results. Finally, summarize the field integrity results, format conformity results, content integrity results, data consistency results, and structural integrity results to obtain the rule validation results.
[0344] In one possible implementation, the semantic auditing module 604 is specifically used for:
[0345] The target chapters suitable for semantic review in the structured text are concatenated into standardized input text. Semantic review is performed chapter by chapter using a local lightweight language model to obtain the semantic quality judgment results for each target chapter. The ratio of the effective text length of each target chapter to the total length of the chapter is calculated. Based on the ratio and the content quality coefficient, the content confidence of each target chapter is determined to obtain the semantic credibility judgment results. The local optical character recognition model is used to extract the text information of the images in the drilling and completion report document to be reviewed. The text information is compared with the text content of the corresponding chapter to obtain the image consistency judgment results. The semantic quality judgment results, semantic credibility judgment results, and image consistency judgment results are summarized to obtain the semantic review results.
[0346] In one possible implementation, the device further includes:
[0347] The rule extension module is used to receive natural language rules input by the user, call the local lightweight language model to parse the natural language rules into executable verification logic, and append the executable verification logic to the preset audit rule library; the preset audit rule library is used to perform business rule verification on structured text.
[0348] In one possible implementation, the fusion module is specifically used for:
[0349] The rule verification results and semantic review results are categorized and hierarchically classified according to structural integrity, data consistency, format standardization, and text semantic quality. Abnormal issues are automatically numbered and listed, compliant content is marked, and the drilling and completion report review results are output.
[0350] The drilling and completion report review device provided in this embodiment can execute the technical solution provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0351] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Figure 7 As shown, the electronic device 70 provided in this embodiment includes at least one processor 701 and a memory 702. Optionally, the device 70 further includes a communication component 703. The processor 701, memory 702, and communication component 703 are connected via a bus 704.
[0352] In a specific implementation, at least one processor 701 executes computer execution instructions stored in memory 702, causing at least one processor 701 to perform the above-described method.
[0353] The specific implementation process of processor 701 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0354] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0355] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0356] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0357] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0358] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0359] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0360] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0361] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0362] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0363] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0364] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0365] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0366] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for reviewing drilling and completion reports, characterized in that, include: Obtain the drilling and completion report document pending review; The drilling and completion report document to be reviewed is parsed in a structured manner to obtain structured text containing chapter and hierarchical information; Perform business rule validation on the structured text to obtain rule validation results; and perform semantic audit on the structured text to obtain semantic audit results. The drilling and completion report review result is generated by combining the rule verification result and the semantic review result.
2. The method according to claim 1, characterized in that, The process of performing structured parsing on the drilling and completion report document to be reviewed yields structured text containing chapter-level information, including: Traverse the entire contents of the drilling and completion report document to be reviewed, and identify the heading level information of each paragraph; Based on the heading level information of each paragraph, the heading level score of each paragraph is calculated respectively; The chapter boundaries are determined based on the title level score. Based on the chapter boundaries, the drilling and completion report documents to be reviewed are grouped by chapter and the main text content is extracted to generate the structured text.
3. The method according to claim 2, characterized in that, The calculation of the heading level score for each paragraph based on the heading level information of each paragraph includes: Assign a first-level weight to paragraphs whose heading level information is a first-level heading; Assign a second-level weight to paragraphs whose heading level information is a second-level heading; Assign a third-level weight to paragraphs whose heading level information is a third-level heading; Wherein, the weight of the first level is greater than the weight of the second level, and the weight of the second level is greater than the weight of the third level.
4. The method according to claim 1, characterized in that, The step of performing business rule validation on the structured text to obtain the rule validation result includes: Perform a non-empty check on the key fields in the structured text to obtain the field integrity check result; Perform format pattern matching on the key data in the structured text to obtain the format standardization judgment result; Perform keyword retrieval on the standard engineering terms in the structured text to obtain the content completeness assessment result; A consistency comparison is performed on the cross-chapter related data in the structured text to obtain the data consistency determination result; Based on a preset standard chapter directory list, the actual chapter titles in the structured text are traversed and matched to obtain the structural integrity judgment result; The rule verification result is obtained by summarizing the field integrity judgment results, the format standardization judgment results, the content integrity judgment results, the data consistency judgment results, and the structural integrity judgment results.
5. The method according to claim 1, characterized in that, The step of performing semantic auditing on the structured text to obtain the semantic auditing result includes: The target chapters suitable for semantic review in the structured text are concatenated into the standardized input text, and the local lightweight language model is called to perform semantic review chapter by chapter to obtain the semantic quality judgment results of each target chapter. Calculate the ratio of the effective text length of each target chapter to the total length of the chapter, and determine the content confidence of each target chapter based on the ratio and the content quality coefficient to obtain the semantic credibility judgment result; The local optical character recognition model is called to extract the text information of the images in the drilling and completion report document to be reviewed, and the text information is compared with the text content of the corresponding chapter to obtain the image consistency judgment result. The semantic quality judgment results, the semantic credibility judgment results, and the image consistency judgment results are combined to obtain the semantic review results.
6. The method according to claim 1, characterized in that, The method further includes: Receive natural language rules from user input; The natural language rules are parsed into executable verification logic; The executable verification logic is added to the preset audit rule base; the preset audit rule base is used to perform business rule verification on the structured text shown.
7. The method according to claim 1, characterized in that, The process of integrating the rule verification result and the semantic review result to generate the drilling and completion report review result includes: The results of the rule verification and the results of the semantic review are categorized and classified in a hierarchical manner according to structural integrity, data consistency, format standardization, and text semantic quality. Abnormal issues are automatically numbered and listed, compliant content is marked, and the drilling and completion report review results are output.
8. A drilling and completion report review device, characterized in that, include: The acquisition module is used to retrieve drilling and completion report documents pending review. The parsing module is used to perform structured parsing on the drilling and completion report document to be reviewed, and obtain structured text containing chapter and hierarchical information; The rule verification module is used to perform business rule verification on the structured text and obtain the rule verification result; The semantic auditing module is used to perform semantic auditing on the structured text and obtain the semantic auditing result; The fusion module is used to fuse the rule verification results and the semantic review results to generate the drilling and completion report review results.
9. An electronic device, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is configured to execute a computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 7.