A method for pumping well operating condition diagnosis and report generation based on large language model

By constructing a pump well working condition diagnosis method based on a large language model, integrating historical data and generating structured reports, the problem of relying on professional knowledge on the pump well working condition diagnosis is solved, efficient and accurate working condition diagnosis and intelligent report generation are achieved, and the intelligent level of oil and gas field development is improved.

CN120217123BActive Publication Date: 2025-09-05CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510695662.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-05
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

In the prior art, the diagnosis of pumping wells depends on professional knowledge, the diagnosis efficiency is low and the document processing workload is large, making it difficult to achieve intelligent and standardized report generation.

Method used

Build a pump well condition diagnosis method based on large language models, integrate historical data through ETL tools, use the Agent module to call large language models and professional small model clusters, generate structured reports, combine multi-objective optimization functions and recommend measures and solutions, and use an intelligent layout engine to generate standardized reports.

Benefits of technology

Significantly improve the efficiency and accuracy of working conditions diagnosis, reduce the need for manual intervention, ensure the standardization and traceability of reports, and improve the intelligence level of oil and gas field development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217123B_ABST
    Figure CN120217123B_ABST
Patent Text Reader

Abstract

The present invention provides a method for diagnosing and reporting on the working conditions of a pumping unit well based on a large language model, relating to the technical field of oil and gas field development. Specifically, the method comprises: constructing a structured template based on oil and gas extraction industry standards; integrating historical working condition data sets and measure implementation archives of the oilfield operating area using an ETL tool to construct a measure corpus, including case working condition feature labels and measure effect labels, and encoding or standardizing the case working condition feature labels; inputting user instructions into an Agent module to output a working condition analysis diagram and corresponding fault type; inputting the current case working condition data and fault type into an improved RAG model, calling the structured template, and generating a case working condition diagnosis report. The technical solution of the present invention overcomes the problems in the prior art of pumping unit working condition diagnosis, which relies on professional knowledge and has a high workload in document processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of oil and gas field development, and in particular to a method for diagnosing the working condition of a pumping well and generating a report based on a large language model. Background Art

[0002] In oil and gas field development, pumping well condition diagnosis is a critical technical step in maintaining efficient oil and gas production. Traditional methods primarily rely on engineers manually analyzing production data such as dynamometer diagrams and current-load curves. This approach suffers from inherent drawbacks such as low diagnostic efficiency (analysis of a single well can take up to 4-8 hours) and a strong reliance on subjective experience. Intelligent improvement solutions often employ single machine learning models for diagnosis, but still face challenges such as difficulties integrating heterogeneous multi-source data, lacking collaborative decision-making mechanisms for models, and weak physical knowledge service capabilities. While current general-purpose big models offer advantages in natural language processing, they also suffer from significant drawbacks. Firstly, they lack the scientific computing capabilities required for oil and gas engineering, making them incapable of supporting specialized fluid dynamics models and physical and chemical mechanism analysis. Secondly, report generation is prone to issues such as non-standard terminology and the lack of standard chart templates, making it difficult to dynamically link historical data to generate diagnostic documents that meet industry standards. There is an urgent need to develop a comprehensive intelligent diagnostic system based on big model technology. By deeply integrating multimodal data, coupled mechanism models, and intelligent algorithms, and enabling the automated generation of standardized reports, this system can overcome the bottleneck in the current paradigm shift from empirical decision-making to intelligent closed-loop decision-making in oil and gas well condition diagnosis.

[0003] Therefore, there is a need for a method for diagnosing and reporting the working conditions of pumping wells based on a large language model, which can significantly reduce the dependence on professional knowledge and the workload of document processing, ensure the accuracy of working condition diagnosis, and effectively improve the intelligence level of oil and gas field development decision-making. Summary of the Invention

[0004] The main purpose of the present invention is to provide a method for diagnosing the working condition of a pumping unit well and generating a report based on a large language model, so as to solve the problem in the prior art that the working condition diagnosis of the pumping unit relies on professional knowledge and has a large workload of document processing.

[0005] To achieve the above objectives, the present invention provides a method for diagnosing the working condition of a pumping well and generating a report based on a large language model, which specifically includes the following steps:

[0006] S1, builds a structured template based on oil and gas extraction industry standards.

[0007] S2, using ETL tools to integrate historical operating condition data sets and measure implementation archives of the oilfield operating area, construct a measure corpus, including: case operating condition feature labels and measure effect labels, and encode or standardize the case operating condition feature labels.

[0008] S3: Input the user instructions into the Agent module. The Agent module calls the large language model and the professional small model cluster to output the working condition analysis diagram and the corresponding fault type.

[0009] S4, input the current case working condition data and fault type into the improved RAG model, call the structured template, and generate the case working condition diagnosis report.

[0010] Furthermore, the structured template in step S1 includes: a single well basic information table, a working condition dynamic monitoring module, a production efficiency evaluation module, a load safety analysis module and a measure decision recommendation module; wherein, the single well basic information table includes: well number, coordinate position and completion date, and the working condition dynamic monitoring module includes: an indicator diagram, a water content change trend diagram, a daily fluid production change trend diagram, an oil pressure and casing pressure change trend diagram, a load change trend diagram and an energy consumption parameter change trend diagram.

[0011] Furthermore, the case operating condition feature tags include: well depth, pump hanging depth, liquid production and water cut; the measure effect tags include oil increase, cost savings, operational risks, measure success rate, technical feasibility and safety. The measure effect tags are used to evaluate the comprehensive benefits and actual effects after the implementation of the measures.

[0012] Furthermore, step S2 specifically includes the following steps:

[0013] S2.1, build a dynamic feature coding system, use One-Hot encoding to process discrete variables in feature labels, and use Min-Max standardization to process continuous variables in feature labels.

[0014] S2.2, map numerical data and text data into a unified vector space for storage.

[0015] Furthermore, step S3 is specifically as follows:

[0016] The user instructions are input into the Agent module, which uses a large language model to understand the user input semantically, extract the semantic vector, and determine whether a report needs to be generated. If a report needs to be generated, the block and well number in the semantic vector are extracted to generate an SQL query statement. The SQL query statement is used to read the casing pressure, oil pressure, daily liquid production, daily oil production, water content, maximum load, and minimum load data from the database and input it into the professional small model cluster. The professional small model cluster outputs the operating condition analysis diagram and the corresponding operating condition fault type.

[0017] Furthermore, step S4 specifically includes the following steps:

[0018] S4.1, find the corresponding measure corpus according to the working condition fault type, calculate the cosine similarity between the working condition feature vector of the current case and the working condition feature vector in the measure corpus, and retain the case working conditions with a cosine similarity close to 1;

[0019] S4.2, build a time-dependent attenuation model and calculate the reference weight of the case working condition :

[0020] ;

[0021] in, is the initial weight, is the attenuation coefficient, is the time interval between case conditions, in months.

[0022] S4.3, based on the cosine similarity and the case condition reference weight, select several case conditions in the most recent period that are most similar to the current case condition.

[0023] Furthermore, step S4 further includes the following steps:

[0024] S4.4, use the characteristics of previous years' case working conditions and the labels of the measures' effects as input data, divide the input data into a training set and a test set, use the training set to train the random forest model, input the screened case working conditions into the trained random forest model, and output the success probability of each measure.

[0025] S4.5, for the cases where the success probability is in the top N, select several measures in the most recent period whose success probability is in the top N.

[0026] Furthermore, step S4 further includes the following steps:

[0027] S4.6, by multi-objective optimization function Calculate the comprehensive benefit value of each measure selected in step S4.7:

[0028] ;

[0029] in, 、 and is the weight coefficient, To take the maximum.

[0030] S4.7, the top five recommended solutions with the best comprehensive benefit values ​​will be used as the final measures.

[0031] S4.8, input the final action plan into the structured template and generate the case working condition diagnosis report.

[0032] Furthermore, the large language model includes: a model based on the transformer architecture; the professional small model cluster includes: an LSTM-based dynamometer diagram anomaly detection model and an operating condition data visualization model.

[0033] The present invention has the following beneficial effects:

[0034] 1. Through multi-model collaboration and intelligent retrieval technology, the efficiency and accuracy of working condition diagnosis are significantly improved, and the rapid identification and precise location of oil well problems can be achieved.

[0035] 2. Combine a dynamic knowledge base with a self-evolution mechanism to continuously optimize the quality of recommended measures and ensure the timeliness and engineering applicability of decision-making recommendations.

[0036] 3. Use an automated report generation engine to effectively unify the output standards of diagnostic results and ensure the standardization and traceability of technical documents.

[0037] 4. Build a full-process intelligent diagnostic system to significantly reduce the need for manual intervention and provide reliable technical support for digital operation and maintenance of oil and gas fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0039] Figure 1 A flow chart of a method for diagnosing the working condition of a pumping well and generating a report based on a large language model of the present invention is shown.

[0040] Figure 2 The figure shows the dynamometer diagram generated by the method provided by the present invention.

[0041] Figure 3 A comprehensive analysis diagram of production conditions generated using the method provided by the present invention is shown.

[0042] Figure 4 The figure shows the moisture content variation trend diagram generated by the method provided by the present invention.

[0043] Figure 5 The figure shows the daily liquid production trend chart generated by the method provided by the present invention.

[0044] Figure 6 The figure shows the oil pressure casing pressure variation trend diagram generated by the method provided by the present invention.

[0045] Figure 7 The maximum load variation trend diagram generated by the method provided by the present invention is shown.

[0046] Figure 8 The figure shows the minimum load variation trend diagram generated by the method provided by the present invention.

[0047] Figure 9 The figure shows the energy consumption parameter variation trend diagram generated by the method provided by the present invention. DETAILED DESCRIPTION

[0048] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0049] like Figure 1 The method for diagnosing the working condition of a pumping well and generating a report based on a large language model specifically includes the following steps:

[0050] S1, builds a structured template based on oil and gas extraction industry standards.

[0051] S2, using ETL tools to integrate historical operating condition data sets and measure implementation archives of the oilfield operating area, construct a measure corpus, including: case operating condition feature labels and measure effect labels, and encode or standardize the case operating condition feature labels.

[0052] S3: Input the user instructions into the Agent module. The Agent module calls the large language model and the professional small model cluster to output the working condition analysis diagram and the corresponding fault type.

[0053] S4, input the current case working condition data and fault type into the improved RAG model, call the structured template, and generate the case working condition diagnosis report.

[0054] Specifically, the structured template in step S1 includes: a single well basic information table, a working condition dynamic monitoring module, a production efficiency evaluation module, a load safety analysis module and a measure decision recommendation module; among them, the single well basic information table includes: well number, coordinate position and completion date, and the working condition dynamic monitoring module includes: indicator diagram, water content change trend diagram, daily fluid production change trend diagram, oil pressure and casing pressure change trend diagram, load change trend diagram and energy consumption parameter change trend diagram.

[0055] Specifically, based on industry standards such as "SY / T 6111-2018 Technical Specifications for Oil and Gas Field Production Dynamic Analysis", a five-layer modular reporting structure is established:

[0056] (1) Single well basic information table, the fields include: well number, well type, well category, coordinate location, well depth, completion date, casing structure, pump hanging depth, design stroke, pump diameter, pump depth, rod string combination, etc. 12 items.

[0057] (2) Dynamic working condition monitoring module: displays the following six types of graphical trend components: dynamometer (load-displacement), water content change trend chart, daily liquid production change trend chart, oil pressure / casing pressure change trend chart, gas load change trend chart, energy consumption parameter (current / power factor) change trend chart; some of the Excel data required to be read for the dynamic working condition monitoring module drawing are shown in Table 1.

[0058] (3) Production efficiency evaluation module: calculates parameters such as pump efficiency, system efficiency, and energy consumption per unit liquid volume.

[0059] (4) Load safety analysis module: Introduces the downhole rod dynamic load simulation and strength verification results.

[0060] (5) Measures and suggestions module: Outputs structured suggestions, including recommended plans, oil production potential prediction, operation cost estimation, and implementation risk level.

[0061] Table 1

[0062]

[0063] Specifically, the case operating condition feature labels include: well depth, pump hanging depth, liquid production and water cut; the measure effect labels include oil increase, cost savings, operational risks, measure success rate, technical feasibility and safety. The measure effect labels are used to evaluate the comprehensive benefits and actual effects after the implementation of the measures.

[0064] Specifically, step S2 includes the following steps:

[0065] S2.1, build a dynamic feature coding system, use One-Hot encoding to process discrete variables in feature labels, and use Min-Max standardization to process continuous variables in feature labels.

[0066] One-hot encoding is performed on the well type (vertical well / inclined well / horizontal well) and operating condition type (normal / gas lock / sand stuck) to generate a 12-dimensional feature vector.

[0067] Use formulas for continuous variables such as "water content" and "pump efficiency" Perform Min-Max normalization.

[0068] For text data, such as "low pump hang-up leads to decreased stroke utilization", semantic features are extracted through the BERT model.

[0069] S2.2, map numerical data and text data into a unified vector space.

[0070] Specifically, step S3 is as follows:

[0071] The user's instructions are input into the Agent module. The Agent module uses a large language model to understand the semantics of the user input, extract the semantic vector, and determine whether a report needs to be generated. If a report needs to be generated, the block and well number in the semantic vector are extracted to generate an SQL query statement. The SQL query statement is used to read the casing pressure, oil pressure, daily liquid production, daily oil production, water content, maximum load, and minimum load data from the database and input them into the professional small model cluster. The professional small model cluster outputs the working condition analysis diagram and the working condition fault type corresponding to the well. Among them, the working condition analysis diagram includes: indicator diagram (load-displacement), water content change trend diagram, daily liquid production change trend diagram, oil pressure / casing pressure change trend diagram, gas load change trend diagram, energy consumption parameter (current / power factor) change trend diagram, such as Figures 2 to 9 shown.

[0072] To ensure the accuracy of key decisions, a manual review interface is implemented at key points, such as "pump adjustment suggestions," allowing professionals to review and adjust model-generated recommendations. Furthermore, the system incorporates an anomaly detection mechanism. When a predicted result, such as oil increase, deviates by two or more times from the historical average, a manual review process is automatically triggered to further verify the model's predictions and initiate corrective actions.

[0073] A model performance monitoring dashboard was developed that integrates key indicators such as inference time, accuracy over the past 10 days, and GPU resource usage. The dashboard presents the model operation status in real time through a visual interface, providing data support for performance evaluation and optimization.

[0074] S4.1, find the corresponding measure corpus according to the working condition fault type, calculate the cosine similarity between the working condition feature vector of the current case and the working condition feature vector in the measure corpus, and retain the case working conditions with a cosine similarity close to 1.

[0075] S4.2, build a time-dependent attenuation model and calculate the reference weight of the case working condition :

[0076] ;

[0077] in, is the initial weight, which is set according to the oil-increasing effect of the measures, ranging from 0.8 to 1.2. is the attenuation coefficient, which is 0.05 in the early stage of oil field development and 0.15 in the late stage. is the time interval between case conditions, in months.

[0078] For example, the time-dependent decay function parameter settings in this case are shown in Table 2:

[0079] Table 2

[0080]

[0081] S4.3, based on the cosine similarity and the case condition reference weight, select several case conditions in the most recent period that are most similar to the current case condition.

[0082] Specifically, step S4 further includes the following steps:

[0083] S4.4, use the characteristics of previous years' case working conditions and the labels of the measures' effects as input data, divide the input data into a training set and a test set, use the training set to train the random forest model, input the screened case working conditions into the trained random forest model, and output the success probability of each measure.

[0084] S4.5, for the cases where the success probability is in the top N, select several measures in the most recent period whose success probability is in the top N.

[0085] Specifically, step S4 further includes the following steps:

[0086] S4.6, by multi-objective optimization function Calculate the comprehensive benefit value of each measure selected in step S4.7:

[0087] ;

[0088] in, 、 and is the weight coefficient, To take the maximum.

[0089] 、 、 is the strategy weight, which can be configured according to specific actual needs. The configuration parameters of this case are shown in Table 3. All variables are normalized before calculation.

[0090] Table 3

[0091]

[0092] S4.7: The top five recommended solutions based on comprehensive benefits are selected as the final action plan. The final action plan is presented through an interactive interface, including a 3D radar chart (showing technical feasibility, economics, and safety), a historical case comparison table, and a Gantt chart of the implementation plan (supporting export to PDF or embedding in a report).

[0093] S4.8, input the final action plan into the structured template and generate the case working condition diagnosis report.

[0094] Specifically, the large language model includes: a model based on the transformer architecture; the professional small model cluster includes: an LSTM-based dynamometer diagram anomaly detection model and an operating condition data visualization model.

[0095] Specifically, the diagnostic report scheme is typeset:

[0096] We've developed an intelligent layout engine that supports simultaneous paging of text and graphics, ensuring a well-organized and aesthetically pleasing content layout. Furthermore, the system automatically generates a multi-level directory, making it easy for users to quickly locate report content. Furthermore, chart sizes automatically adjust based on the page layout, achieving automated and intelligent layout and improving the efficiency and quality of report generation.

[0097] A digital watermark embedding module has been developed to embed digital watermarks during report generation, enhancing report traceability and security. A timestamp, accurate to the second, will be added to the header and footer to record the exact time the report was generated. A data source hash checksum will also be embedded to ensure data integrity and authenticity. Furthermore, the system version number will be embedded to facilitate subsequent version management and issue tracking.

[0098] A compliance verification module has been developed to ensure that report content adheres to industry standards and specifications. This module automatically checks the integrity of figure and chart numbers, ensuring the correct correspondence between figures and text. It also enforces unit consistency verification, automatically converting units to standard units (e.g., MPa, m³ / d, etc.). Furthermore, terminology is verified against SY / T 6111-2018 to ensure accurate use of professional terminology in reports, enhancing their professionalism and authority.

[0099] Specifically, the measure corpus is updated in real time:

[0100] The system automatically tracks key production data after the implementation of measures, including liquid volume, moisture content, energy consumption and other indicators within 30 days, 90 days and 180 days, to comprehensively evaluate the actual effect of the measures and provide data support for subsequent knowledge update and optimization. The parameters to be collected are shown in Table 4:

[0101] Table 4:

[0102]

[0103] Set the trigger conditions for updating the measure corpus. When the new measure's oil-increasing effect is better than the top three measures in history, the operating cost is reduced by more than 15% under the same operating conditions, or the risk level is reduced by 2 levels or more, the measure corpus update will be automatically initiated to ensure the timeliness and accuracy of the knowledge base.

[0104] Establish an incremental learning mechanism to periodically update the industry vocabulary of the large language model (LLM), the parameters of the small model, and the Faiss index database every month to continuously optimize model performance, enhance the intelligence level of the system, and promote the dynamic evolution of the knowledge system.

[0105] The present invention replaces traditional manual diagnosis with intelligent technology, significantly reducing the reliance on professional knowledge and the workload of document processing, while ensuring the accuracy of working condition diagnosis and effectively improving the intelligence level of oil and gas field development decision-making.

[0106] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the scope of protection of the present invention.

Claims

1. A method for diagnosing and reporting the working conditions of a pumping well based on a large language model, characterized in that: The specific steps include: S1, build a structured template based on oil and gas extraction industry standards; S2: Use ETL tools to integrate historical operating condition data sets and measure implementation archives of the oilfield operation area to build a measure corpus, including case operating condition feature labels and measure effect labels, and encode or standardize the case operating condition feature labels; S3: Input the user instructions into the Agent module, which calls the large language model and the specialized small model cluster to output the working condition analysis diagram and the corresponding fault type; S4, input the current case working condition data and fault type into the improved RAG model, call the structured template, and generate the case working condition diagnosis report; Step S4 specifically includes the following steps: S4.1, find the corresponding measure corpus according to the working condition fault type, calculate the cosine similarity between the working condition feature vector of the current case and the working condition feature vector in the measure corpus, and retain the case working conditions with a cosine similarity close to 1; S4.2, build a time-dependent attenuation model and calculate the reference weight of the case working condition : ; in, is the initial weight, is the attenuation coefficient, is the time interval of case working conditions, in months; S4.3, based on the cosine similarity and the reference weight of the case working condition, select several case working conditions in the most recent period that are most similar to the current case working condition; Step S4 also includes the following steps: S4.4: Use the characteristics of previous years' case conditions and labels of measure effects as input data, divide the input data into a training set and a test set, use the training set to train a random forest model, input the screened case conditions into the trained random forest model, and output the success probability of each measure; S4.5, calculate the success probability of the first N cases , select several measures in the most recent period whose success probability is in the top N; Step S4 also includes the following steps: S4.6, by multi-objective optimization function Calculate the comprehensive benefit value of each measure selected in step S4.5: ; in, 、 and is the weight coefficient, To take the maximum; S4.7, the top five recommended solutions with the best comprehensive benefit values ​​will be used as the final measures; S4.8, input the final action plan into the structured template and generate the case working condition diagnosis report.

2. The method for diagnosing and reporting the working condition of a pumping well based on a large language model according to claim 1, characterized in that: The structured template in step S1 includes: a single well basic information table, a working condition dynamic monitoring module, a production efficiency evaluation module, a load safety analysis module, and a measure decision recommendation module; among them, the single well basic information table includes: well number, coordinate location, and completion date, and the working condition dynamic monitoring module includes: an indicator diagram, a water content change trend diagram, a daily fluid production change trend diagram, an oil pressure and casing pressure change trend diagram, a load change trend diagram, and an energy consumption parameter change trend diagram.

3. The method for diagnosing and reporting the working condition of a pumping well based on a large language model according to claim 1, characterized in that: The case operating condition feature tags include: well depth, pump hanging depth, liquid production and water cut; the measure effect tags include oil increase, cost savings, operational risks, measure success rate, technical feasibility and safety. The measure effect tags are used to evaluate the comprehensive benefits and actual effects after the implementation of the measures.

4. The method for diagnosing and reporting the working condition of a pumping well based on a large language model according to claim 1, characterized in that: Step S2 specifically includes the following steps: S2.1, build a dynamic feature encoding system, use one-hot encoding to process discrete variables in feature labels, and use min-max standardization to process continuous variables in feature labels; S2.2, map numerical data and text data into a unified vector space for storage.

5. The method for diagnosing and reporting the working condition of a pumping well based on a large language model according to claim 1, characterized in that: Step S3 is specifically as follows: The user instructions are input into the Agent module, which uses a large language model to understand the user input semantically, extract the semantic vector, and determine whether a report needs to be generated. If a report needs to be generated, the block and well number in the semantic vector are extracted to generate an SQL query statement. The SQL query statement is used to read the casing pressure, oil pressure, daily liquid production, daily oil production, water content, maximum load, and minimum load data from the database and input it into the professional small model cluster. The professional small model cluster outputs the operating condition analysis diagram and the corresponding operating condition fault type.

6. The method for diagnosing and reporting the working condition of a pumping well based on a large language model according to claim 5, characterized in that: Large language models, including models based on the transformer architecture; professional small model clusters, including LSTM-based dynamometer diagram anomaly detection models and operating condition data visualization models.

Citation Information

Patent Citations

  • Large model-based data quality analysis report generation method, system and device and storage medium

    CN119357177A