A pumping unit well production condition analysis method based on prompt word adaptive generation
By constructing a prompt-based adaptive generation method, and utilizing a large language model in the analysis of pumping well production status, a closed-loop optimization of dynamically generated structured analysis results was achieved. This solved the problems of manual dependence and insufficient consistency in traditional analysis, and improved the quality and reliability of the analysis results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA UNIV OF PETROLEUM (EAST CHINA)
- Filing Date
- 2026-03-20
- Publication Date
- 2026-06-05
AI Technical Summary
Traditional analysis of the production status of pumping wells relies on human experience, which makes it difficult to unify the analysis criteria, trace the process of conclusion generation, and achieve insufficient consistency of results. Furthermore, it is heavily reliant on manually designed prompt templates, and the cost of template maintenance and updating is high, making it difficult to form a unified and quantifiable quality control mechanism for results.
A method based on adaptive generation of prompt words is constructed. Four sets of pre-designed meta-prompt word instructions work together on a large language model to form a data analysis, knowledge constraint, comprehensive analysis and quality assessment model, realize the dynamic generation of structured analysis results, and ensure quality compliance through closed-loop iterative optimization.
By integrating real-time data and professional knowledge, we can generate analytical results that are quantifiable in quality, consistent in structure, and executable in engineering, reducing reliance on manually designed prompt templates and improving the completeness, consistency, and verifiability of the analytical results.
Smart Images

Figure CN121882220B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of artificial intelligence and oil and gas field development engineering technology, specifically involving a method for analyzing the production status of pumping wells based on a large language model with adaptive cue word generation. Background Technology
[0002] Analysis of pumping well production status typically relies on multi-source data such as dynamometer cards, production dynamics, and equipment data. It also requires combining engineering mechanisms and empirical rules to comprehensively assess downhole supply and drainage relationships, equipment operating conditions, and energy efficiency. In traditional production management, such analyses are often completed by engineers based on reports and experience, leading to problems such as inconsistent analytical standards, difficulty in tracing the conclusion generation process, and insufficient consistency of results. This makes it difficult to meet the requirements of refined management for structured conclusion output and actionable recommendations.
[0003] Large Language Models (LLMs) have the ability to organize multi-source information into natural language interpretations. However, when general large language models are directly used for the analysis of pumping well production status, phenomena such as inferences that do not strictly rely on data evidence, improper application of engineering rules, and unstable output structure are prone to occur, resulting in fluctuations in the completeness, logical consistency, and executability of the analysis results. Consequently, the reliability, reproducibility, and auditability of the results are insufficient. To improve output quality, current practices typically rely on manual prompt engineering, where engineers design different prompt templates for different task types and data formats to ensure coverage of key analysis points, correct engineering logic, and stable expression structure. While this approach can improve output quality to some extent, different task types often correspond to different sets of key analysis points and engineering reasoning chains. Furthermore, the same task can exhibit characteristic differences under different well conditions and data formats. Engineers usually need to repeatedly design, maintain, and adjust prompt templates to address task and data differences. Consequently, this approach suffers from strong reliance on expert experience, high time costs for template maintenance and updates, and difficulties in migration and expansion. Moreover, it lacks adaptability to changes in tasks, differences in data characteristics, and knowledge updates, making it difficult to establish a unified and quantifiable result quality control mechanism.
[0004] Therefore, there is an urgent need for an automated method for analyzing the production status of pumping wells: by introducing structured production data and professional knowledge, it can reliably generate structured analysis results that cover key engineering features, are logically consistent and executable, without the need for manual design of prompt templates for different tasks. It can also quantitatively evaluate the quality of the results and provide feedback iterations to ensure the engineering applicability and consistency of the analysis conclusions. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, such as rigid static prompts, uncontrollable quality of analysis results, difficulty in integrating professional knowledge with real-time data, and lack of adaptive optimization for different tasks, this invention proposes a method for analyzing the production status of pumping wells based on adaptive prompt generation. The core idea of this method is to construct a closed-loop framework oriented towards achieving quality standards in pumping well production status analysis results. Through the dynamic generation and iterative updating of prompts, controllable constraints are achieved on the analysis process, thereby stably generating structured analysis results that meet quality standards. This invention utilizes the same basic large language model, and through four pre-designed and functionally differentiated meta-prompt instructions, functional instantiation is performed on the same basic large language model to form four collaborative functional models: a data analysis prompt generation model, a knowledge constraint prompt generation model, a comprehensive analysis model, and a quality assessment model. These four models work together to form a complete closed loop of "retrieval-generation-evaluation-optimization." This invention can stably output pumping well production status analysis results that are quantifiable in quality, consistent in structure, and executable in engineering, based on the integration of real-time data and professional knowledge. It is also adaptive to different task types, thereby reducing the reliance on manually designed prompt templates and improving the completeness, consistency, and verifiability of the analysis results.
[0006] The technical solution of the present invention is as follows:
[0007] A method for analyzing the production status of pumping wells based on adaptive cue word generation includes the following steps:
[0008] Step 1: Receive user requests, perform multi-source information retrieval, and obtain structured data input, professional knowledge fragments, and historical similar cases related to user requests;
[0009] Step 2: Based on the multi-source information retrieval results, the initial data analysis prompts and knowledge constraint prompts are generated in parallel using the data analysis prompt generation model and the knowledge constraint prompt generation model;
[0010] Step 3: Based on the comprehensive analysis model, apply the initial data analysis prompts and knowledge constraint prompts to conduct a structured analysis of the production status of the pumping well, and conduct a quality assessment of the structured analysis results based on the quality assessment model;
[0011] Step 4: If the quality assessment fails to meet the standards, the prompt words will be iteratively optimized based on the assessment feedback, and the analysis and assessment will be re-executed until a quality-compliant analysis result is generated.
[0012] Step 5: Construct the analysis task of achieving the quality assessment standard into a new case, perform similarity matching with the case library, add non-duplicate cases to the case library, and update the case library.
[0013] Furthermore, the specific process of step 1 is as follows:
[0014] Step 1.1: Data Retrieval: Receive user requests in natural language text format, parse and extract target well number, time range, and task type information; query the pumping unit well production database based on the information to obtain corresponding raw production data, and call a pre-set feature analysis function to process the raw production data, generating structured data input containing statistical feature values; the pumping unit well production database stores raw monitoring and acquisition data related to pumping unit well production, including dynamometer card data, production dynamic data, and equipment data; wherein, dynamometer card data includes suspension point load and displacement, pump end load and displacement; production dynamic data includes daily fluid production, wellhead pressure, bottom hole temperature, bottom hole flowing pressure, reservoir static pressure, and dynamic fluid level height; equipment data includes pump mounting depth, pump diameter, stroke, stroke frequency, torque, motor power, and power consumption;
[0015] Step 1.2: Knowledge Retrieval: The knowledge base stores expert reports and professional knowledge fragments formed by document processing related to the field of oil pumping unit production. Professional knowledge fragments relevant to user needs are retrieved from the knowledge base. First, keyword extraction and vectorization transformation are performed on user needs. Keyword retrieval and vector semantic retrieval are executed in parallel within the knowledge base to obtain keyword retrieval ranking lists and vector retrieval ranking lists, forming a preliminary ranking result. Then, the preliminary ranking result is re-ranked using a reciprocal ranking fusion algorithm to obtain a reciprocal ranking fusion score. Finally, based on the re-ranking result, the top performers with reciprocal ranking fusion scores exceeding a preset first threshold are returned. Individual pieces of professional knowledge form a set of professional knowledge fragments;
[0016] Step 1.3: Perform case retrieval: The case library stores historical cases whose structured analysis results meet the quality standards. Retrieve historical cases similar to user needs from the case library. First, vectorize the user needs, converting them into user need vectors. Then, traverse the case library, calculating the cosine similarity between the current user need vector and the user need vector in each historical case. Finally, generate a similarity ranking list based on the cosine similarity calculation results, and return the top cases with similarity higher than a preset second threshold. A set of historical similar cases is formed from 10 historical cases; if there are no historical cases with a similarity higher than the preset second threshold, an empty set of historical similar cases is returned.
[0017] Furthermore, the specific process of step 2 is as follows:
[0018] Step 2.1: Generate data analysis prompts: Based on user needs and the structured data input retrieved in Step 1.1, and referring to the historical similar case set retrieved in Step 1.3, the data analysis prompt generation model is invoked to generate data analysis prompts focusing on the analysis of existing data features; if the historical similar case set is empty, then only data analysis prompts are generated based on user needs and the structured data input retrieved in Step 1.1; the data analysis prompts are structured text composed of several analysis points, with predefined start and end identifiers as boundaries, and each analysis point begins with a predefined list symbol;
[0019] Step 2.2: Generate knowledge constraint prompts: Based on user needs and the previous search terms retrieved in Step 1.2 For each piece of professional knowledge, referring to the historical similar case set retrieved in step 1.3, the knowledge constraint prompt word generation model is invoked to generate knowledge constraint prompt words focusing on professional rule guidance; if the historical similar case set is empty, then only based on user needs and the previous cases retrieved in step 1.2... Each piece of professional knowledge generates a knowledge constraint prompt; the knowledge constraint prompt is a structured text composed of several professional rules, with predefined start and end identifiers as boundaries, and each rule starts with a predefined list symbol.
[0020] Furthermore, in step 2.1, the working principle of the data analysis prompt word generation model is as follows: Data analysis prompt words are generated by inputting preset meta-prompt word instructions into the basic large language model. The specific working process is as follows:
[0021] Roles and Inputs: The instruction model acts as a data analysis expert for the production status of pumping wells, receiving user requirements, structured data inputs, similar historical cases, and suggestions for improvement using prompts.
[0022] Core task logic: If no suggestions for improvement are received, the model is required to identify the task type and extract specific values and features from the data input; if the set of similar historical cases is not empty, suggestions containing multiple analysis points are generated by referring to the content of similar historical cases; if the set of similar historical cases is empty, suggestions containing multiple analysis points are directly generated based on the data input and task type; if suggestions for improvement are received, the model is required to adjust the provided suggestions to be optimized accordingly based on the specific type of suggestion.
[0023] Output and Constraints: The prompts output by the model must conform to a predefined format, and each analysis point must follow strict data constraint principles, that is, the analysis is only based on the data features that are clearly present in the data input.
[0024] Furthermore, in step 2.2, the working principle of the knowledge constraint prompt word generation model is as follows: knowledge constraint prompt words are generated by inputting preset meta-prompt word instructions into the basic large language model. The specific working process is as follows:
[0025] Roles and Inputs: The instruction model acts as a knowledge expert in the field of pumping well production status engineering, receiving user requirements, fragments of professional knowledge, similar historical cases, and suggestions for improvement based on prompts;
[0026] Core task logic: If no suggestions for improvement are received, the model is required to filter knowledge directly related to the task type from professional knowledge fragments; if the set of similar cases is not empty, multiple core professional rules are extracted and generated by referring to the rule structure of similar cases; if the set of similar cases is empty, multiple core professional rules are directly extracted and generated based on the professional knowledge fragments; if suggestions for improvement are received, the model is required to optimize the rules in the provided suggestions for improvement.
[0027] Output and Constraints: The prompts output by the model must conform to a predefined format, and the extracted rules must be engineering practice-oriented, focusing on engineering principles, operating procedures or empirical principles, and directly used to guide the analysis process of the task.
[0028] Furthermore, the specific process of step 3 is as follows:
[0029] Step 3.1: Generate Structured Analysis Results: Combining user needs with the structured data input retrieved in Step 1.1, and guided by the data analysis prompts generated in Step 2.1 and the knowledge constraint prompts generated in Step 2.2, the comprehensive analysis model is invoked to process the structured data input and generate structured analysis results. The structured analysis results are structured text organized according to preset hierarchical headings, including data display, professional analysis, and optimization suggestions.
[0030] Step 3.2: Generate Quality Assessment Results: Based on user needs and the structured analysis results generated in Step 3.1, the quality assessment model is invoked to evaluate the structured analysis results according to preset quantitative assessment indicators, and the quality assessment results are output. The quantitative assessment indicators include the completeness of data feature extraction, analytical practicality, and logical consistency. The quality assessment results are structured text, containing assessment conclusion fields separated by specific identifiers, data analysis prompt words and their improvement suggestion fields, and knowledge constraint prompt words and their improvement suggestion fields.
[0031] The evaluation conclusion field is a binary logic state. When the quality evaluation meets the standard, it is marked as "true" and each improvement suggestion field is confirmed as "quality meets the standard, no improvement is needed"; when the quality evaluation does not meet the standard, it is marked as "false" and each improvement suggestion field provides optimization suggestions for the corresponding prompt words.
[0032] Furthermore, in step 3.1, the working principle of the comprehensive analysis model is as follows: Structured analysis results are generated by inputting pre-set meta-prompt word instructions into the basic large language model. The specific working process is as follows:
[0033] Roles and Inputs: The instruction model acts as a comprehensive analysis expert for the production status of pumping wells, receiving user requirements, structured data input, data analysis prompts, and knowledge constraint prompts.
[0034] Analysis logic and structured output: The model is required to strictly follow the key points in the data analysis prompts and the rules in the knowledge constraint prompts, and to organize the output according to the preset template, which includes data processing and display, oil well production status analysis and optimization measures suggestions;
[0035] Process constraints: The instructions require that the analysis must be based on the actual values in the structured data input, and that each essential engineering feature covered in the data analysis prompts be specifically analyzed, and that the rules in the knowledge constraint prompts be applied for qualitative judgment.
[0036] Furthermore, in step 3.2, the working principle of the quality assessment model is as follows: Quality assessment results are generated by inputting pre-set meta-prompt word instructions into the basic large language model. The specific working process is as follows:
[0037] Roles and Inputs: The instruction model acts as a quality assessment expert, receiving user requirements, data analysis prompts, knowledge constraint prompts, and structured analysis results;
[0038] Quantitative evaluation of execution: The instruction model determines the task type based on user needs and calls the corresponding list of essential engineering features. Then, based on predefined quantitative evaluation indicators, it calculates and scores the data feature extraction completeness, analytical practicality, and logical consistency of the analysis results. Among them, the data feature extraction completeness statistics include the number of essential engineering features explicitly mentioned and used in the analysis; analytical practicality is comprehensively scored from three aspects: use of essential feature data, executability of conclusions, and rationality of professional logic, for a total of 100 points; logical consistency is comprehensively scored from four aspects: coverage of key analysis points, application of knowledge constraints, completeness of reasoning chain, and lack of internal contradictions, for a total of 100 points.
[0039] Achievement Judgment and Feedback Generation: The model is deemed to have met the standards if and only if the data feature extraction completeness is 100%, the analysis practicality score is ≥85, and the logical consistency score is ≥90. If the standards are not met, the model is required to generate specific optimization suggestions for the data analysis prompts and knowledge constraint prompts, respectively, and the suggestions must not require the addition of data that does not exist in the input data.
[0040] Output format: Evaluation results and recommendations must be output in a predefined format.
[0041] Furthermore, the specific process of step 4 is as follows:
[0042] Receive the quality assessment results generated in step 3.2. If the content of the assessment conclusion field is "true", the quality assessment meets the standard, the iteration terminates, and the structured analysis results that meet the quality standard are output. If the content of the assessment conclusion field is "false", the quality assessment does not meet the standard. The improvement suggestions for the data analysis prompts and the current version of the data analysis prompts are returned to the data analysis prompt generation model to generate optimized data analysis prompts. At the same time, the improvement suggestions for the knowledge constraint prompts and the current version of the knowledge constraint prompts are returned to the knowledge constraint prompt generation model to generate optimized knowledge constraint prompts.
[0043] Subsequently, steps 3 and 4 are re-executed based on the optimized data analysis prompts and optimized knowledge constraint prompts to form a closed-loop iterative optimization process until the quality assessment meets the standard or the preset maximum number of iterations is reached. If the standard is not met after reaching the maximum number of iterations, the current optimal result is output and manual intervention is prompted.
[0044] Furthermore, the specific process of step 5 is as follows:
[0045] Step 5.1: Construct a new structured case by combining the user requirements, final data analysis prompts, and knowledge constraint prompts related to this quality assessment achievement task.
[0046] Step 5.2: Traverse all historical cases in the case library, and calculate the cosine similarity between the new case and each historical case in three parts: user needs, data analysis prompts, and knowledge constraint prompts; then, perform a weighted sum of the three cosine similarities to obtain the comprehensive similarity score. :
[0047] ;
[0048] in, , , The weights of user needs, data analysis prompts, and knowledge constraint prompts are respectively assigned, and ; Cosine similarity to user needs; Cosine similarity of prompt words for data analysis; Cosine similarity of knowledge-constrained prompt words;
[0049] Step 5.3: Calculate the overall similarity score. With the preset duplicate detection threshold If a comparison is made, If it is, it is determined to be a new case and stored in the case database; if If the result is not found, it will be considered a duplicate case and will not be stored.
[0050] The beneficial technical effects brought about by this invention are as follows.
[0051] 1. Improved the adaptability and stability of pumping unit well production status analysis results: By generating instructions to constrain the analysis process based on specific tasks, real-time data and retrieval knowledge, the dependence on static prompt word templates is reduced, enabling stable output of pumping unit well production status analysis results that conform to the preset structure and engineering scope under different task types and different data input conditions.
[0052] 2. A closed-loop control mechanism for the quality of analysis results has been established: Through the closed loop of "retrieval-generation-evaluation-optimization", the quantitative quality assessment of structured analysis results is directly linked to the iterative update of the two types of prompt words, so that the analysis results can continuously approach and reach the preset quality standards under the drive of evaluation feedback, thereby ensuring the reliability, consistency and verifiability of the output results.
[0053] 3. It provides a balance between professionalism and automation: By injecting domain principles and rules through knowledge constraint prompts and closely adhering to real-time data characteristics through data analysis prompts, the automatically generated analysis results not only conform to professional standards but also have a solid data foundation, reducing reliance on human expert experience.
[0054] 4. Excellent Task Extensibility and Framework Generalization: This method first constructs and validates typical engineering analysis tasks such as oil well torque load analysis, downhole pump dynamometer diagram analysis, comparative analysis of dynamometer diagram index evolution, time-series prediction analysis of production indicators, dynamic analysis of oil well inflow, temperature and pressure field analysis, sucker rod string stress analysis, and single-well energy efficiency analysis. The core generalization capability of this framework lies in its universal "retrieval-generation-evaluation-optimization" closed-loop mechanism. For new analysis tasks, implementers only need to define the "essential engineering features" list for the task, while ensuring that structured data input that meets the requirements of the task can be retrieved and provided. Based on the above conditions, without changing the core process and model architecture, dedicated analysis prompts can be generated and continuously optimized for new tasks through the same adaptive mechanism, thereby achieving rapid expansion and analysis. Attached Figure Description
[0055] Figure 1 This is a diagram illustrating the overall architecture and workflow of a method for analyzing the production status of pumping wells based on adaptive prompt word generation, as described in this invention. Detailed Implementation
[0056] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0057] like Figure 1 As shown, the method of the present invention includes the following steps:
[0058] Step 1: Receive user requests and perform multi-source information retrieval to obtain structured data input, professional knowledge fragments, and historical similar cases related to the user requests. The core of this step is to receive user requests and then perform data retrieval, knowledge retrieval, and case retrieval in parallel to obtain structured data input, professional knowledge fragments, and historical similar cases related to the user requests. This provides complete basic data, professional knowledge, and historical experience support for subsequent analysis. Figure 1 The process involves three parallel retrieval steps. The specific steps are as follows:
[0059] Step 1.1: Data retrieval;
[0060] The system receives user requests in natural language text format, parses and extracts target well number, time range, and task type information; queries the pumping well production database based on the information to obtain the corresponding raw production data, and calls the preset feature analysis function to process the raw production data to generate structured data input containing statistical feature values.
[0061] The pumping well production database stores raw monitoring and acquisition data related to pumping well production, including dynamometer card data, production dynamic data, and equipment data. The dynamometer card data includes suspension point load and displacement, and pump end load and displacement. The production dynamic data includes daily fluid production, wellhead pressure, bottom hole temperature, bottom hole flowing pressure, reservoir static pressure, and dynamic fluid level. The equipment data includes pump mounting depth, pump diameter, stroke, stroke frequency, torque, motor power, and power consumption.
[0062] In this embodiment, the user inputs a specific analysis task description in natural language: "Analyze the changes in the dynamometer indicators of well XX in January 2025." Core information is extracted using natural language processing technology: target well number (XX), time range (January 2025), and task type (changes in dynamometer indicators), clarifying the search scope and core requirements. Based on the parsed target well number, time range, and task type, the dynamometer indicator data for well XX in January 2025 is queried from the pumping unit well production database, including maximum load, minimum load, daily fluid production, filling degree, and effective stroke. Subsequently, a pre-set feature analysis function library is invoked to process and transform the raw data, generating key statistical feature values, including average, maximum, minimum, standard deviation, and trend, forming structured data input.
[0063] Step 1.2: Knowledge Retrieval;
[0064] Professional knowledge fragments related to user needs are retrieved from the knowledge base to form a set of professional knowledge fragments. The knowledge base stores professional knowledge fragments formed after processing unstructured knowledge such as expert reports and literature related to the oil pumping unit production field. Specifically, user needs are first extracted and vectorized; keyword retrieval and vector semantic retrieval are then performed in parallel in the knowledge base to obtain preliminary ranking results, including a keyword retrieval ranking list and a vector retrieval ranking list. Then, the Reverse Ranking Fusion (RRF) algorithm is used to re-rank the preliminary ranking results to obtain the Reverse Ranking Fusion Score, calculated using the following formula:
[0065] (1);
[0066] in, For the first The inverse ranking of each candidate result is combined with the score; and Indicates the first The ranking of each candidate result in the ranking lists of keyword retrieval and vector semantic retrieval, respectively; This is a smoothing parameter, a constant, to avoid the problem of excessively high search scores due to excessively high rankings, thus affecting the results. If the... If a candidate result is not in a certain list, then the corresponding or It should be infinity.
[0067] Finally, based on the re-sorting results, the top-ranked candidates whose combined score exceeds a preset first threshold are returned. A collection of relevant professional knowledge fragments is formed by combining individual professional knowledge fragments.
[0068] In this embodiment, the extracted keywords are dynamometer indicators and changes, and the vectorization transformation adopts the TF-IDF model; Set the threshold to 60; preset the first threshold to 0.70; finally, return the top 5 knowledge fragments with the highest fusion scores to form a set of professional knowledge fragments. In this embodiment, the returned content covers the core professional knowledge such as the change law of the core indicators of the dynamometer card, the judgment standard of abnormal fluctuations, and the correlation rules between the indicator evolution and the downhole working conditions.
[0069] Step 1.3, Case Search;
[0070] Retrieve historical cases similar to user needs from the case library. The case library only stores historically successful cases (hereinafter referred to as historical cases) whose structured analysis results generated by this method meet the quality standards, and is not a case library pre-set by humans; each case is a structured record, including case identifier, user requirement text, data analysis prompt text, and knowledge constraint prompt text.
[0071] Specifically, firstly, user needs are vectorized into user need vectors. Then, the case library is traversed, and the similarity between the current user need vector and the user need vectors in each historical case is calculated using cosine similarity. The formula for calculating cosine similarity is:
[0072] (2);
[0073] in, for and Cosine similarity; This represents the current user demand vector; This represents a vector of user needs from historical cases. and They are respectively , In the Components in each dimension; This represents the total dimension of the vector;
[0074] A similarity ranking list is generated based on the cosine similarity calculation results, and the top results with similarity scores higher than a preset second threshold are returned. A set of historical similar cases is formed from 10 historical cases. If no historical case with a similarity higher than a preset second threshold exists, an empty set of historical similar cases is returned.
[0075] In this embodiment, the second threshold is preset to 0.70, and the top 3 historical cases with similarity higher than the threshold are returned. Specifically, the historical cases with similarity of 0.81 are matched: "Analysis of the changes in the dynamometer card of YY oil well in January 2025?", "Analysis of the evolution of the dynamometer card of XX oil well in November 2024?", and "What is the working condition of the dynamometer card of XX oil well in November 2024?". The returned historical cases contain structured information including user needs, data analysis prompts, and knowledge constraint prompts.
[0076] Step 2: Based on the multi-source information retrieval results, the initial data analysis prompt word generation model and knowledge constraint prompt word generation model are generated in parallel. The historical similar case set is an optional input for guiding with a small sample size. When the historical similar case set is empty, the data analysis prompt word generation model generates data analysis prompt words only based on structured data input, and the knowledge constraint prompt word generation model generates knowledge constraint prompt words only based on the set of professional knowledge fragments. The specific process is as follows:
[0077] Step 2.1: Data analysis prompt word generation;
[0078] Based on user needs and the structured data input retrieved in step 1.1, and referring to the historical similar case set retrieved in step 1.3, the data analysis prompt word generation model is invoked to generate data analysis prompt words focusing on the analysis of existing data features; if the historical similar case set is empty, then only data analysis prompt words are generated based on user needs and the structured data input retrieved in step 1.1.
[0079] The data analysis prompt generation model invoked here generates data analysis prompts by inputting preset meta-prompt instructions into the basic large language model. The data analysis prompt generation model invoked in step 2.1 of this invention, and the knowledge constraint prompt generation model, comprehensive analysis model, and quality assessment model invoked in subsequent steps 2.2, 3, and 4, achieve functional differentiation by inputting different preset meta-prompt instructions into the same basic large language model. The basic large language model is a general-purpose large language model based on the Transformer architecture, requiring no additional full-domain fine-tuning.
[0080] The specific content of the meta-prompt word instruction of the data analysis prompt word generation model is as follows:
[0081] (1) Role definition;
[0082] The instruction model acts as a "data analysis expert for analyzing the production status of pumping wells".
[0083] (2) Input information;
[0084] User requirements: A description of the specific analysis task proposed by the user in natural language; Structured data input: Structured data retrieved from the production database and processed through feature calculation, including specific numerical values and statistical features; Similar historical cases: Previous cases retrieved from the case library. Historical case studies; suggestion for improvement of prompt words (optional): specific text improvement suggestions for prompt words generated by the quality assessment model during the iterative optimization process, targeting data analysis prompt words.
[0085] (3) Core tasks;
[0086] If the suggestion for improvement is empty, perform the initial generation: the model is required to identify the task type in the user's needs; extract all available specific values and features from the structured data input; if the set of similar historical cases is not empty, refer to the structure and dimensions of the data analysis suggestion points in the similar historical cases; if the set of similar historical cases is empty, generate the analysis point structure based on the structured data input and user needs, and generate suggestion words containing 3 to 5 analysis points.
[0087] If the suggestion for improvement is not empty, perform iterative optimization: the model is required to make targeted adjustments to the data analysis suggestions from the previous round based on the type of suggestion, including adding key points, replacing generalized descriptions with analysis based on specific numerical values, or deleting key points involving non-existent data requirements.
[0088] (4) Output format;
[0089] The instructions mandate that the model output must strictly adhere to the following text format, including fixed boundary markers at the beginning and end:
[0090] <<<DATA_PROMPT_START> >>;
[0091] -Key Point 1: [Feature analysis based on specific numerical values explicitly present in the structured data input];
[0092] -Key Point 2: [Identify trends / anomalies based on clearly defined features in structured data input];
[0093] -Key Point 3: [Executable Analysis Directions Based on Existing Data];
[0094] -Key Point X: [···];
[0095] <<<DATA_PROMPT_END> >>;
[0096] (5) Core requirements;
[0097] Strict data constraints: Each analytical point must be based solely on fields, values, or features that actually exist in the structured data input. Statements such as "suggested supplementation," "requires monitoring," or "not provided," or assumptions about non-existent data, are strictly prohibited in the analytical points. Precise and quantitative language: Each analytical point should be concise, not exceeding 50 words, and should prioritize the use of specific numerical values and clear trend descriptions.
[0098] The data analysis prompts are structured text consisting of several analysis points. This text is bounded by predefined start and end identifiers, and each analysis point begins with a predefined list symbol. In this embodiment, the generated data analysis prompts are:
[0099] <<<DATA_PROMPT_START> >>;
[0100] - Key Point 1: The average fullness is 85.3%, showing a downward trend, with a standard deviation of 0.31 and relatively small fluctuations.
[0101] - Key Point 2: The average daily liquid production is 11.71 tons, with a maximum of 13.4 tons and a minimum of 9.92 tons, showing a downward trend.
[0102] - Key Point 3: The average maximum load is 34229.17 N, showing an upward trend, with a standard deviation of 230.63, indicating significant fluctuations.
[0103] <<<DATA_PROMPT_END> >>;
[0104] Step 2.2: Generation of knowledge constraint prompts;
[0105] Based on user needs and the results retrieved in step 1.2 For each piece of professional knowledge, referring to the historical similar case set retrieved in step 1.3, the knowledge constraint prompt word generation model is invoked to generate knowledge constraint prompt words focusing on professional rule guidance; if the historical similar case set is empty, then only based on user needs and the previous cases retrieved in step 1.2... Generate knowledge constraint prompts from each piece of professional knowledge.
[0106] The knowledge constraint prompt word generation model invoked here generates knowledge constraint prompt words by inputting pre-defined meta-prompt word instructions into the basic large language model. The specific content of the meta-prompt word instructions for the knowledge constraint prompt word generation model is as follows:
[0107] (1) Role definition;
[0108] The instruction model acts as a "knowledge expert in the field of oil and gas field development and artificial lift engineering".
[0109] (2) Input information;
[0110] User requirements: Description of the user's analysis task; Expert knowledge fragments: Retrieved from the knowledge base through hybrid retrieval. Each relevant piece of professional knowledge; similar historical cases: previous cases retrieved from the case database. One historical case; suggestion for improvement of prompt words (optional): specific improvement suggestions for prompt words based on knowledge constraints generated by the quality assessment model.
[0111] (3) Core tasks;
[0112] If the suggestion for improvement is empty, perform the initial generation: the model is required to select knowledge directly related to the user's task type from the professional knowledge fragments; if the set of similar historical cases is not empty, the rule structure and professional direction of the knowledge constraint suggestions in the similar historical cases are referenced; if the set of similar historical cases is empty, professional rules are directly extracted and generated based on the professional knowledge fragments, and 2 to 3 core professional rules that can directly guide the current task analysis process are extracted and generated.
[0113] If the suggestion for improvement is not empty, perform iterative optimization: the model is required to optimize the rules in the knowledge constraint suggestions from the previous round, so that the rules are closer to the core of the task and more operable.
[0114] (4) Output format;
[0115] The instructions mandate that the model output must strictly adhere to the following text format, including fixed boundary markers at the beginning and end:
[0116] <<<TEXT_PROMPT_START> >>;
[0117] - Rule 1: [Core professional rules directly related to task type];
[0118] - Rule 2: [Specific constraints based on engineering principles];
[0119] - Rule X: [···];
[0120] <<<TEXT_PROMPT_END> >>;
[0121] (5) Core requirements;
[0122] Task-oriented: Each rule must directly serve the specific task type determined by the user's needs, constraining or guiding the direction of data analysis; Engineering practice-oriented: Rules should be based on on-site operating procedures, industry standards, or expert experience, providing principled guidance that can be directly applied to engineering analysis, avoiding complex theoretical formulas or algorithm descriptions; Content filtering: Proactively exclude academic, purely theoretical, or content with low relevance to current engineering analysis practices.
[0123] The knowledge constraint prompts are structured texts composed of several professional rules. These texts are bounded by predefined start and end identifiers, and each rule begins with a predefined list symbol. In this embodiment, the generated knowledge constraint prompts are:
[0124] <<<TEXT_PROMPT_START> >>;
[0125] Rule 1: Compare the indicator diagram analysis data from January 2025 with the previous data, and pay close attention to the changing trends of key parameters such as filling degree, liquid production, and effective stroke.
[0126] Rule 2: Based on the fluctuation of data characteristics, judge the downhole working conditions, analyze the correlation between the extreme changes of maximum / minimum load and the decrease in filling degree, and investigate potential abnormal working conditions.
[0127] <<<TEXT_PROMPT_END> >>;
[0128] Step 3: Based on the comprehensive analysis model, apply the initial data analysis prompts and knowledge constraint prompts to conduct a structured analysis of the production status of the pumping well, and evaluate the quality of the structured analysis results based on the quality assessment model; this step is the core analysis and evaluation process of the method, such as... Figure 1 As shown, it includes two stages: structured analysis and quality assessment. The specific process is as follows:
[0129] Step 3.1: Generation of structured analysis results;
[0130] Combining user needs with the structured data input retrieved in step 1.1, and guided by the data analysis prompts generated in step 2.1 and the knowledge constraint prompts generated in step 2.2, the comprehensive analysis model is invoked to process the structured data input and generate structured analysis results. The structured analysis results are structured text organized according to preset hierarchical headings, including data display, professional analysis, and optimization suggestions.
[0131] The comprehensive analysis model invoked here generates structured analysis results by inputting pre-defined meta-prompt word instructions into the underlying large language model. The specific content of the meta-prompt word instructions in the comprehensive analysis model is as follows:
[0132] (1) Role definition;
[0133] The instruction model acts as a "comprehensive analysis expert of the production status of oil wells".
[0134] (2) Input information;
[0135] User requirements: Description of the user's analysis task; Structured data input: Structured oil well production data and feature values; Data analysis prompts: Prompts output by the data analysis prompt generation model, containing specific analysis points; Knowledge constraint prompts: Prompts output by the knowledge constraint prompt generation model, containing professional rules.
[0136] (3) Core tasks;
[0137] The model is required to strictly follow the analysis points in the data analysis prompts and the professional rules in the knowledge constraint prompts, and combine them with the actual values in the structured data input to complete the structured analysis and generate the final report.
[0138] (4) Output format;
[0139] The instruction forces the model to organize the output according to the following three-part structure, and embeds specific analysis content:
[0140] I. Data organization and presentation;
[0141] [List the key parameters extracted from the structured data input in Markdown table format, including parameter name, value / range, and unit];
[0142] II. Analysis of Oil Well Production Status;
[0143] [Based on actual numerical values in structured data input, professional analysis is conducted by combining key points of data analysis prompts and rules of knowledge constraint prompts. Each essential engineering feature must be analyzed individually or in combination, including trend analysis and risk assessment.]
[0144] III. Recommendations for Optimization Measures;
[0145] [Based on the foregoing analysis, specific and feasible engineering recommendations are proposed];
[0146] (5) Core requirements;
[0147] Authentic Data Basis: All analytical statements must be derived from actual values in structured data inputs, and data fabrication or assumptions are prohibited; Comprehensive Coverage of Key Points and Rules: The analysis must cover all key points in the data analysis prompts and reasonably apply all rules in the knowledge constraint prompts for qualitative analysis; Natural and Coherent Logic: The analysis report should naturally integrate the prompt requirements into the professional discourse, avoiding stiff expressions such as "according to point X" or "according to rule Y".
[0148] Step 3.2: Generation of quality assessment results;
[0149] Based on user requirements and the structured analysis results generated in step 3.1, the quality assessment model is invoked to evaluate the structured analysis results according to preset quantitative evaluation indicators, and the quality assessment results are output. The quantitative evaluation indicators include data feature extraction completeness assessment, analytical usability assessment, and logical consistency assessment.
[0150] The quality assessment model invoked here generates quality assessment results by inputting pre-defined meta-prompt word instructions into the underlying large language model. The specific content of the meta-prompt word instructions in the quality assessment model is as follows:
[0151] (1) Role definition;
[0152] The instruction model acts as an "expert in analyzing the production status of pumping wells and optimizing prompts."
[0153] (2) Input information;
[0154] User requirements: Description of the user's analysis task; Data analysis prompts: Data analysis prompts to be evaluated; Knowledge constraint prompts: Knowledge constraint prompts to be evaluated; Structured analysis results: Structured report generated by the comprehensive analysis model.
[0155] (3) Core tasks;
[0156] The structured analysis results are evaluated for quality based on preset quantitative evaluation indicators to determine whether they meet the standards; if they do not meet the standards, specific improvement suggestions are generated for data analysis prompts and knowledge constraint prompts.
[0157] (4) Quantitative evaluation indicators;
[0158] Data feature extraction completeness (target: 100%):
[0159] First, determine the task type based on user requirements and then retrieve the corresponding "List of Essential Engineering Features". Calculate the percentage of essential engineering features explicitly mentioned and used in the structured analysis results, representing the total number of features. The formula is:
[0160] (3);
[0161] in, For completeness; This represents the number of features extracted. The total number of engineering features required for the corresponding task;
[0162] Analysis of practicality (target: ≥85 points):
[0163] Weighted scoring based on three sub-dimensions:
[0164] a. Use of essential feature data (50 points): Evaluate whether the actual values in the structured data input were used for each essential engineering feature in the structured analysis results;
[0165] b. Feasibility of the conclusions (30 points): Evaluate whether the proposed optimization suggestions are clear, specific, and operable on-site;
[0166] c. Professional logical rationality (20 points): Evaluate whether the analysis process is professional and reasonable, and whether the qualitative application of the rules in the knowledge constraint prompts is appropriate;
[0167] (Key principle: The analysis results should not be required to provide data support for the rules that is not present in the structured data input.)
[0168] Logical consistency (target: ≥90 points):
[0169] Scoring is based on four sub-dimensions (25 points each):
[0170] a. The analysis covers all key points of the data analysis prompts.
[0171] b. Analyze and apply knowledge constraint prompt rules appropriately.
[0172] c. The reasoning chain from data to problem to suggestion is complete.
[0173] d. The report contains no internal contradictions or conflicts.
[0174] The formatted output will be assessed for quality based on the above quantitative evaluation indicators. The quality is considered satisfactory only if the data feature extraction completeness is 100%, the total score for analytical practicality is ≥85, and the total score for logical consistency is ≥90. If any one of these indicators is not met, specific and actionable improvement suggestions must be generated for that indicator. These suggestions must point to the optimization direction of both data analysis prompts and knowledge constraint prompts, and must not include suggestions such as "XX data needs to be supplemented," which are impossible to implement under the current data conditions.
[0175] The method pre-defines a list of essential engineering features for different analysis tasks to guide prompt word generation and quality assessment. The core content of this list is shown in Table 1 below:
[0176] Table 1 List of Required Engineering Characteristic Values
[0177] .
[0178] (5) Output format;
[0179] The instruction mandates that evaluation results must strictly adhere to the following text output format:
[0180] <<<DISTINGUISH_START> >>;
[0181] true / false;
[0182] <<<DISTINGUISH_END> >>;
[0183] <<<DATA_PROMPT_START> >>;
[0184] Data analysis prompt: [Full content];
[0185] <<<DATA_PROMPT_END> >>;
[0186] <<<DATA_PROMPT_ADVICE_START> >>;
[0187] Data analysis suggestion improvement suggestions: [Specific suggestions];
[0188] <<<DATA_PROMPT_ADVICE_END> >>;
[0189] <<<TEXT_PROMPT_START> >>;
[0190] Knowledge constraint prompt: [Full content];
[0191] <<<TEXT_PROMPT_END> >>;
[0192] <<<TEXT_PROMPT_ADVICE_START> >>;
[0193] Suggestions for improving knowledge constraint prompts: [Specific suggestions];
[0194] <<<TEXT_PROMPT_ADVICE_END> >>;
[0195] If the assessment meets the standards, then <<<DISTINGUISH_START> >> and <<<DISTINGUISH_END> Output true between >> and fill in "Quality meets the standard, no improvement needed" in the improvement suggestion section; if it does not meet the standard, output false and fill in the specific improvement suggestion text.
[0196] The quality assessment results are structured text, containing assessment conclusion fields separated by specific identifiers, data analysis prompt words and their improvement suggestion fields, and knowledge constraint prompt words and their improvement suggestion fields.
[0197] The evaluation conclusion field is a binary logic state. When the quality evaluation meets the standard, it is marked as "true" and each improvement suggestion field is confirmed as "quality meets the standard, no improvement is needed"; when the quality evaluation does not meet the standard, it is marked as "false" and each improvement suggestion field provides optimization suggestions for the corresponding prompt words.
[0198] The initial evaluation results generated in this embodiment according to the above format are shown below:
[0199] <<<DISTINGUISH_START> >>;
[0200] false;
[0201] <<<DISTINGUISH_END> >>;
[0202] <<<DATA_PROMPT_START> >>;
[0203] #Data Analysis Hints#:
[0204] - Key Point 1: The average fullness is 85.3%, showing a downward trend, with a standard deviation of 0.31 and relatively small fluctuations.
[0205] - Key Point 2: The average daily liquid production is 11.71 tons, with a maximum of 13.4 tons and a minimum of 9.92 tons, showing a downward trend.
[0206] - Key Point 3: The average maximum load is 34229.17 N, showing an upward trend, with a standard deviation of 230.63, indicating significant fluctuations.
[0207] <<<DATA_PROMPT_END> >>;
[0208] <<<DATA_PROMPT_ADVICE_START> >>;
[0209] Improvement Suggestions: Based on the task type "Comparative Analysis of Oil Well Dynamometer Indicator Evolution," the essential feature list includes five items: maximum load, daily fluid production, filling degree, effective stroke, and minimum load. However, the current data analysis prompts lack relevant data for "minimum load" and "effective stroke," resulting in incompleteness. It is recommended to supplement the statistical information for these two data points, for example: "Key Point 4: The mean minimum load is XX N, showing a trend of XX, standard deviation XX, and fluctuation XX." and "Key Point 5: The mean effective stroke is XX meters, showing a trend of XX, standard deviation XX, and fluctuation XX." This will ensure that all essential features are extracted and used for analysis.
[0210] <<<DATA_PROMPT_ADVICE_END> >>;
[0211] <<<TEXT_PROMPT_START> >>;
[0212] #Knowledge constraint prompts#:
[0213] Rule 1: Compare the indicator diagram analysis data from January 2025 with the previous data, and pay close attention to the changing trends of key parameters such as filling degree, liquid production, and effective stroke.
[0214] Rule 2: Based on the fluctuation of data characteristics, judge the downhole working conditions, analyze the correlation between the extreme changes of maximum / minimum load and the decrease in filling degree, and investigate potential abnormal working conditions.
[0215] <<<TEXT_PROMPT_END> >>;
[0216] <<<TEXT_PROMPT_ADVICE_START> >>;
[0217] Improvement suggestions: The current analysis results apply Rule 2 too generally, without clearly defining how specific numerical indicators map to abnormal downhole conditions. It is recommended to add judgment criteria based on specific data combinations to the knowledge constraint prompts. For example, refine "Rule 2" to: "Analyze operating conditions based on specific numerical values: Focus on whether 'decreased filling degree and increased load difference' indicates insufficient fluid supply or pump leakage," to improve the rationality of professional logic under visual conditions.
[0218] <<<TEXT_PROMPT_ADVICE_END> >>;
[0219] Step 4: If the quality assessment fails to meet the standards, the prompt words are iteratively optimized based on the assessment feedback, and the analysis and assessment are re-executed until a quality-compliant analysis result is generated; the specific process is as follows:
[0220] Receive the quality assessment results generated in step 3.2. If the assessment conclusion field is "true", the quality assessment is considered successful, the iteration terminates, and the structured analysis results at this point are output. If the assessment conclusion field is "false", the quality assessment is considered unsuccessful. Improvement suggestions for the data analysis prompts and the current version of the data analysis prompts are returned to the data analysis prompt generation model to generate optimized data analysis prompts. Simultaneously, improvement suggestions for the knowledge constraint prompts and the current version of the knowledge constraint prompts are returned to the knowledge constraint prompt generation model to generate optimized knowledge constraint prompts.
[0221] Subsequently, steps 3 and 4 are re-executed based on the optimized data analysis prompts and optimized knowledge constraint prompts to form a closed-loop iterative optimization process until the quality assessment meets the standard or the preset maximum number of iterations is reached. The maximum number of iterations is 5. If the standard is not met after 5 iterations, the current optimal result is output and manual intervention is prompted.
[0222] In this embodiment, in the initial evaluation result, the evaluation conclusion field is set to "false," indicating that the quality evaluation did not meet the standards. Improvement suggestions for the data analysis prompts and the current version of the data analysis prompts are then returned to the data analysis prompt generation model to generate optimized data analysis prompts. Simultaneously, improvement suggestions for the knowledge constraint prompts and the current version of the knowledge constraint prompts are returned to the knowledge constraint prompt generation model to generate optimized knowledge constraint prompts.
[0223] Subsequently, steps 3 and 4 are re-executed based on the optimized data analysis prompts and optimized knowledge constraint prompts to generate the second structured analysis results and evaluation results.
[0224] In this embodiment, the second generated evaluation result is shown below:
[0225] <<<DISTINGUISH_START> >>;
[0226] true;
[0227] <<<DISTINGUISH_END> >>;
[0228] <<<DATA_PROMPT_START> >>;
[0229] #Data Analysis Hints#:
[0230] - Key Point 1: The average fullness is 85.3%, showing a downward trend, with a standard deviation of 0.31 and relatively small fluctuations.
[0231] - Key Point 2: The average daily liquid production is 11.71 tons, with a maximum of 13.4 tons and a minimum of 9.92 tons, showing a downward trend.
[0232] - Key Point 3: The average maximum load is 34229.17 N, showing an upward trend, with a standard deviation of 230.63, indicating significant fluctuations.
[0233] - Key Point 4: The minimum load average is 19074.68N, showing a downward trend, with a standard deviation of 163.61 and relatively small fluctuations.
[0234] - Key Point 5: The average effective stroke is 4.02 meters, showing a downward trend, with a standard deviation of 0.04 and relatively small fluctuations.
[0235] <<<DATA_PROMPT_END> >>;
[0236] <<<DATA_PROMPT_ADVICE_START> >>;
[0237] Suggested improvement: Quality meets standards, no improvement required;
[0238] <<<DATA_PROMPT_ADVICE_END> >>;
[0239] <<<TEXT_PROMPT_START> >>;
[0240] #Knowledge constraint prompts#:
[0241] Rule 1: Compare the indicator diagram analysis data from January 2025 with the previous data, and pay close attention to the changing trends of key parameters such as filling degree, liquid production, and effective stroke.
[0242] - Rule 2: Based on pure numerical characteristics, the working condition is judged: If the filling degree and the effective stroke decrease simultaneously, and the maximum load increases and the minimum load decreases (the load difference increases), it is necessary to combine these quantitative characteristics to focus on investigating abnormal situations such as insufficient liquid supply or pump and valve leakage.
[0243] <<<TEXT_PROMPT_END> >>;
[0244] <<<TEXT_PROMPT_ADVICE_START> >>;
[0245] Suggested improvement: Quality meets standards, no improvement required;
[0246] <<<TEXT_PROMPT_ADVICE_END> >>;
[0247] In this example, if the evaluation conclusion field is set to "true," indicating that the generated analysis results meet the quality standards, the iteration terminates, and the structured analysis results that meet the quality assessment standards are output. The final structured analysis results output in this embodiment are shown below:
[0248] I. Data organization and presentation;
[0249] The final data processing results of the structured analysis are shown in Table 2:
[0250] Table 2. Data processing results of the final structured analysis output.
[0251] .
[0252] II. Analysis of Oil Well Production Status;
[0253] 1. Essential feature analysis;
[0254] - Filling degree: The average value is 85.3%, showing a downward trend, with a standard deviation of only 0.31, indicating small fluctuations. The decrease in filling degree may reflect reduced pump efficiency or weakened formation fluid supply capacity.
[0255] - Daily fluid production: The average daily fluid production was 11.71 tons, with a maximum of 13.4 tons and a minimum of 9.92 tons, showing an overall downward trend. This is consistent with the trend in well fullness, further indicating a decline in well production capacity.
[0256] - Maximum load: The average value is 34229.17 N, showing an upward trend, and the standard deviation reaches 230.63, indicating significant fluctuations. The increase in maximum load may be related to changes in the stress state of the sucker rod string or increased downhole frictional resistance.
[0257] -Minimum load: The mean is 19074.68 N, showing a decreasing trend, with a standard deviation of 163.61, indicating relatively small fluctuations. Combined with the changes in maximum load, it is speculated that this may be due to reduced downward resistance of the sucker rod string or increased leakage within the pump.
[0258] - Effective stroke: The mean is 4.02 meters, showing a decreasing trend, with a standard deviation of only 0.04, indicating small fluctuations. A decrease in effective stroke usually means increased stroke loss, which may be related to elastic deformation of the sucker rod string or reduced pump efficiency.
[0259] 2. Application of professional rules;
[0260] - Comparing the dynamometer data from January 2025 with previous data, the focus is on the coordinated changes in key parameters such as fill level, production volume, and effective stroke. This month's data shows a continuous downward trend in both fill level and effective stroke, directly reflected in the decline in daily production volume. This indicates that downhole pump efficiency is currently decreasing, and system reactive power loss and stroke loss are increasing.
[0261] - Current operating data shows that not only is the filling degree and effective stroke decreasing simultaneously, but the maximum load is also showing an upward trend while the minimum load is showing a downward trend, resulting in a significant increase in the load difference between the up and down directions. This multi-parameter linkage anomaly in the height indication well may indicate abnormal operating conditions such as insufficient fluid supply (the dynamic fluid level drops, leading to an increase in dynamic load) or pump and valve leakage (the leakage of traveling or fixed valves leads to a reduction in the effective lift volume).
[0262] 3. Trend analysis and risk assessment;
[0263] Based on the above analysis, the Xi14-9-5 oil well will exhibit the following trend in January 2025:
[0264] - The decline in oil well production capacity is mainly reflected in the decrease in fill level and daily fluid production.
[0265] - The pumping system is operating abnormally, characterized by an increase in maximum load, a decrease in minimum load, and a shortening of the effective stroke.
[0266] - There is a potential risk of pump leakage, which needs to be further confirmed by examining the dynamometer card pattern. If the above trend continues, it may lead to further reduction in pump efficiency, increased energy consumption, and a decline in production, affecting the economic benefits of the oil well.
[0267] III. Recommendations for Optimization Measures;
[0268] 1. Specific optimizations based on essential features;
[0269] - Improve filling degree: Improve the formation fluid supply capacity by adjusting the operating parameters of the pumping unit (such as reducing the number of strokes) or implementing production enhancement measures (such as injecting water to replenish formation energy), thereby improving the filling degree.
[0270] - Control load fluctuations: Inspect the sucker rod string to check for uneven wear or corrosion; at the same time, optimize the balance of the pumping unit to reduce load fluctuations.
[0271] -Restore effective stroke: Check the wear of internal pump components and replace pump valves or seals if necessary to reduce stroke loss.
[0272] 2. Principled recommendations based on professional rules;
[0273] - Regularly monitor changes in the indicator diagram to promptly detect and address gas interference or pump leakage issues.
[0274] - By comparing historical data, we can conduct in-depth analysis of the root causes of changes in pump efficiency and stroke loss, and formulate targeted improvement measures.
[0275] 3. Feasible on-site operating procedures;
[0276] -Data Acquisition and Analysis: Record indicator diagram data daily, focusing on observing the changing trends of fill degree, maximum load, minimum load, and effective stroke;
[0277] - Equipment inspection and maintenance: Arrange for technicians to inspect the sucker rod string and pump, paying particular attention to uneven wear, corrosion, and pump and valve sealing performance.
[0278] - Parameter optimization and adjustment: Adjust the pumping unit stroke rate and balance appropriately according to the actual production situation to ensure stable system operation.
[0279] - Implementation of production enhancement measures: If the formation's fluid supply capacity is insufficient, water injection or other production enhancement measures can be considered to increase the oil well's production capacity.
[0280] - The above measures can effectively improve the production status of the West 14-9-5 oil well, extend the service life of the equipment, and enhance economic benefits.
[0281] Step 5: Construct the analysis task of achieving quality assessment standards into a new case, perform similarity matching with the case library, and add non-duplicate cases to the case library to update the case library; the specific process is as follows:
[0282] Step 5.1: Construct a new structured case by combining the user requirements, final data analysis prompts, and knowledge constraint prompts related to this quality assessment achievement task.
[0283] Step 5.2: Traverse all historical cases in the case library, and calculate the cosine similarity between the new case and each historical case in three parts: user needs, data analysis prompts, and knowledge constraint prompts; then, perform a weighted sum of the three cosine similarities to obtain the comprehensive similarity score. :
[0284] (4);
[0285] in, , , The weights of user needs, data analysis prompts, and knowledge constraint prompts are respectively assigned, and ; Cosine similarity to user needs; Cosine similarity of prompt words for data analysis; Cosine similarity of knowledge constraint prompts.
[0286] Step 5.3: Calculate the overall similarity score. With the preset duplicate detection threshold If a comparison is made, If it is, it is determined to be a new case and stored in the case database; if If the result is not found, it will be considered a duplicate case and will not be stored.
[0287] In this embodiment, after step 4 outputs the achievement analysis results, the core information of this achievement task is packaged into the following structured new case:
[0288] ["<<"<USERS_QUESTION_START> >>{Analyze the changes in the dynamometer card indicators of oil well XX in January 2025.}<<<USERS_QUESTION_END> >>", "Data analysis prompt: <<"<DATA_PROMPT_START> >>\n- Key Point 1: The average filling degree is 85.3%, showing a downward trend, with a standard deviation of 0.31 and relatively small fluctuations.\n- Key Point 2: The average daily liquid production is 11.71 tons, with a maximum of 13.4 tons and a minimum of 9.92 tons, showing a downward trend.\n- Key Point 3: The average maximum load is 34229.17 N, showing an upward trend, with a standard deviation of 230.63 and significant fluctuations.\n- Key Point 4: The average minimum load is 19074.68 N, showing a downward trend, with a standard deviation of 163.61 and relatively small fluctuations.\n- Key Point 5: The average effective stroke is 4.02 meters, showing a downward trend, with a standard deviation of 0.04 and relatively small fluctuations.\n<<<DATA_PROMPT_END> >>", Knowledge constraint prompt: <<<TEXT_PROMPT_START> >>\n- Rule 1: Compare the indicator diagram analysis data from January 2025 with previous data, focusing on the changing trends of key parameters such as filling degree, liquid production, and effective stroke. \n- Rule 2: Analyze operating conditions based on purely numerical characteristics: If the filling degree and effective stroke decrease simultaneously, and the maximum load increases while the minimum load decreases (load difference increases), these quantitative characteristics should be used to investigate abnormalities such as insufficient liquid supply or pump / valve leakage. \n<<<TEXT_PROMPT_END> >>"];
[0289] In this embodiment, the repeated determination threshold is... =0.55; =0.4, =0.3, =0.3; =0.81, =0.56, =0.50, weighted to obtain the comprehensive similarity score. =0.64, indicating that the overall similarity exceeds the duplicate judgment threshold of 0.55, and is judged as a duplicate case, which will not be added to the database to avoid redundancy in the case database.
[0290] Steps 2, 3, and 4 in the method of this invention constitute the core closed loop of prompt word generation and optimization, achieving the following key objectives:
[0291] Parallel generation and specialized focus: Through independent generation models, data analysis prompts focusing on data features and knowledge constraint prompts focusing on domain rules are generated separately, realizing the decoupling and targeted integration of data and knowledge;
[0292] Quantitative assessment and precise feedback: The analysis results are rigorously evaluated through preset, quantifiable assessment indicators, and the root causes of non-compliance are transformed into specific improvement suggestions for the prompt words themselves;
[0293] Closed-loop iterative optimization: Improvement suggestions are fed back to the prompt word generation model to drive it to optimize the prompt words in the next round, thus forming an optimization loop that continuously iterates and fine-tunes the intermediate products, namely prompt words, with the goal of improving the quality of the final analysis results.
[0294] Case accumulation and few-sample guidance: User needs and prompt word combinations for the target task are constructed into historical cases and stored in the database; subsequently, similar cases are retrieved for new needs and input into the prompt word generation model in the form of few-sample examples to guide the generation of two types of prompt words, making the prompt words more in line with the target requirements, thereby improving the first generation hit rate and reducing the number of iterations.
[0295] To demonstrate the applicability and scalability of the method of this invention in various production analysis tasks of pumping units, complete output examples for three different analysis tasks are presented below, involving wellbore temperature and pressure field analysis, surface pumping unit load and torque analysis, and sucker rod string stress analysis, respectively. These examples are all automatically generated and iteratively optimized using this method to obtain high-quality analysis results, further verifying the generalization ability of the framework in diverse engineering tasks.
[0296] Example 1: Wellbore temperature and pressure field analysis;
[0297] User requirement: "Based on the production data of oil well West 47-3-2 in January 2025, analyze the current pressure and temperature field distribution within the wellbore?"
[0298] The structured analysis results generated by the method of this invention are shown below:
[0299] I. Data organization and presentation;
[0300] The data processing results of the generated structured analysis are shown in Table 3:
[0301] Table 3. Data processing results of the generated structured analysis results
[0302] .
[0303] II. Analysis of Oil Well Production Status;
[0304] 1. Essential feature analysis;
[0305] - Pressure field distribution: The bottom hole flowing pressure is 2.57 MPa, and the pump inlet pressure is 1.81 MPa, indicating that the pressure inside the wellbore gradually decreases with increasing depth. The pressure difference from the bottom hole to the pump inlet is 0.76 MPa, which conforms to the pressure gradient distribution law. According to Rule 1, the pressure gradient is calculated based on the fluid density in the wellbore. This pressure difference is within a reasonable range and no anomalies are observed.
[0306] - Dynamic fluid level position: The dynamic fluid level is located at a well depth of 662.89 m, which is higher than the pump depth of 844.0 m, indicating sufficient fluid supply. However, the relative position of the dynamic fluid level and the pump depth needs to be monitored to ensure pump efficiency. If the fluid level drops further, it may affect the fluid supply to the pump inlet.
[0307] - Temperature field distribution: The temperature field ranges from 4.0 to 844.0 m in depth, corresponding to a temperature range of 32.28 to 51.4℃. The temperature near the pump depth reaches 51.4℃, which is considered a high-temperature region. By fitting the temperature logging data segment by segment, the linear increase in temperature with depth was verified, and no abnormal temperature change intervals were found.
[0308] 2. Application of professional rules;
[0309] - Pressure gradient verification: The pressure gradient distribution in the wellbore was calculated using the formula ΔP=ρgΔh. The difference between the bottom hole flowing pressure and the pump inlet pressure (0.76 MPa) is consistent with the design range, indicating that the pressure field distribution in the wellbore is normal.
[0310] - Temperature change rate analysis: By fitting the temperature logging data segment by segment, the calculated temperature change rate was approximately 0.023 ℃ / m, showing a linear increasing trend, and no significant abnormal temperature change range was found. However, the temperature near the pump depth was relatively high (51.4 ℃), and the potential impact of high temperature on equipment performance and fluid properties needs to be considered.
[0311] 3. Trend analysis and risk assessment;
[0312] - Pressure field trend: The current pressure field distribution is stable, but the pressure difference from the bottom of the well to the pump inlet is large, which may indicate that there is a certain flow resistance or gas interference in the wellbore, and the flow state in the wellbore needs to be further monitored.
[0313] - Temperature risk: Temperatures near the pump depth reach 51.4 ℃. Long-term operation may lead to material aging or changes in fluid viscosity, thus affecting pump efficiency and lifespan. Regular checks of the pump and related equipment's temperature resistance are recommended.
[0314] III. Recommendations for Optimization Measures;
[0315] 1. Specific optimizations based on essential features;
[0316] - Adjust pump depth: Consider moving the pump depth upwards appropriately to reduce the impact of high temperatures near the pump depth on the equipment, while ensuring a reasonable distance between the dynamic liquid surface and the pump depth to avoid insufficient liquid supply.
[0317] - Optimize liquid supply conditions: Monitor the changing trend of the dynamic liquid level and ensure that it is always a certain distance above the pump depth to avoid a decrease in pump efficiency due to insufficient liquid supply.
[0318] 2. Principled recommendations based on professional rules;
[0319] - Pressure gradient management: Continue to use the formula ΔP=ρgΔh to periodically check the wellbore pressure gradient distribution to ensure that the difference between the bottom hole flowing pressure and the pump inlet pressure is within the design range.
[0320] -Temperature monitoring and protection: Strengthen real-time monitoring of temperature near the pump depth, and use high-temperature resistant materials or cooling devices when necessary to reduce the impact of high temperature on equipment and fluid properties.
[0321] 3. Feasible on-site operating procedures;
[0322] - Regularly measure the pressure distribution inside the wellbore, paying particular attention to whether the difference between the bottom hole flowing pressure and the pump inlet pressure exceeds the design range.
[0323] - Check the dynamic liquid level position, record its changing trend, ensure sufficient liquid supply and avoid the liquid level from being too low.
[0324] - Install a temperature sensor near the pump depth to monitor temperature changes in real time. If the temperature continues to rise, take cooling measures.
[0325] - Perform regular maintenance on pumps and related equipment, focusing on checking the sealing and mechanical performance of equipment under high-temperature conditions, and replace aging parts in a timely manner.
[0326] The results fully cover the 10 essential engineering features required for thermo-baric field analysis, including bottom hole flowing pressure, pump depth, pump inlet pressure, and pump outlet pressure. Based on real data and combined with professional rules, the thermo-baric field distribution was systematically analyzed.
[0327] Example 2: Load torque analysis of a surface oil pumping unit;
[0328] User requirement: "Analyze the current load torque of the surface pumping unit in oil well West 47-3-2?"
[0329] The generated structured analysis results are shown below:
[0330] I. Data organization and presentation;
[0331] The data processing results of the generated structured analysis are shown in Table 4:
[0332] Table 4. Data processing results of the generated structured analysis results
[0333] .
[0334] II. Analysis of Oil Well Production Status;
[0335] 1. Essential feature analysis;
[0336] The net torque value of 30.98 N·m indicates a significant imbalance in the load on the pumping unit. This value directly reflects the asymmetrical load experienced by the pumping unit during operation, which may lead to decreased equipment efficiency and accelerated wear of mechanical components.
[0337] - Excessive power output during the upstroke may indicate reduced pump efficiency or rod friction issues. This is typically related to downhole conditions such as pump leakage, gas interference, or rod wear.
[0338] 2. Application of professional rules;
[0339] - The net torque value has exceeded the reasonable range of balance. It is preliminarily judged that it may be close to or exceed the rated torque bearing capacity of the equipment. Further verification of the equipment specifications is required.
[0340] - Based on the changes in downhole operating conditions and pump efficiency, the current torque fluctuations exhibit abnormal characteristics, which may increase the risk of equipment operation, including problems such as motor overload and gear wear in the gearbox.
[0341] 3. Trend analysis and risk assessment;
[0342] - If the current imbalance in the pumping unit persists for an extended period, it will lead to a significant increase in energy consumption and accelerate equipment aging. Furthermore, excessive work done during the upstroke may indicate a gradual decrease in downhole pump efficiency, requiring vigilance regarding the potential risk of production decline.
[0343] Torque curve analysis shows that optimizing balance is a top priority; otherwise, equipment wear and maintenance costs will increase significantly.
[0344] III. Recommendations for Optimization Measures;
[0345] Based on the above analysis, specific and feasible engineering recommendations are proposed:
[0346] 1. Specific optimizations based on essential features;
[0347] - Adjust the position or weight of the pumping unit's balance block to balance the load during the up and down strokes, thereby reducing the net torque value to a reasonable range (generally recommended to be controlled within ±10 kN·m).
[0348] - Regularly monitor the trend of net torque changes to promptly identify and address new imbalance issues.
[0349] 2. Principled recommendations based on professional rules;
[0350] - Verify the pumping unit model and its rated torque parameters to ensure that the current operating conditions do not exceed the design limits.
[0351] - Conduct a comprehensive inspection of the downhole conditions, focusing on pump efficiency, rod string condition, and wellbore fluid properties, and rule out the possibility of pump leakage or gas interference.
[0352] 3. Feasible on-site operating procedures;
[0353] - Step 1: After stopping the machine, adjust the position of the balance block and gradually test the operating effect under different configurations until the optimal balance is achieved.
[0354] - Step 2: Verify the adjusted pump efficiency and rod stress through dynamometer diagram testing to ensure that the improvement effect meets expectations.
[0355] - Step 3: Record energy consumption data before and after the adjustment, compare and analyze the energy-saving effect, and use it as a reference for subsequent management.
[0356] The above analysis accurately extracted three essential features: net torque, load, and operating conditions. Based on real data and professional rules, targeted suggestions were made for abnormal torque fluctuations.
[0357] Example 3: Force analysis of the sucker rod string;
[0358] User requirement: "Analyze and calculate the stress on the sucker rod string based on the current sucker rod assembly in well West 48-7-2."
[0359] The generated structured analysis results are shown below:
[0360] I. Data organization and presentation;
[0361] The data processing results of the generated structured analysis are shown in Table 5:
[0362] Table 5. Data processing results of the generated structured analysis.
[0363] .
[0364] II. Analysis of Oil Well Production Status;
[0365] 1. Essential feature analysis;
[0366] - Maximum lateral force: The maximum lateral force is 621.54 N, occurring at a depth of 336 m. The rod at this location experiences concentrated stress, possibly due to localized vibration or uneven wear during pumping. Special attention should be paid to the risk of fatigue damage at this depth.
[0367] - Stress ratio distribution: The stress ratio ranges from 0.0801 to 0.3999, with an average of 0.3071, and the overall distribution is relatively uniform. According to the safety threshold requirements, the current stress ratio is within a reasonable range, indicating that the overall safety of the pole is high.
[0368] - Lateral force fluctuation: The standard deviation of the lateral force is 225.74 N, which is relatively large, reflecting uneven stress on the rod string. Combined with the depth distribution characteristics, this may indicate unreasonable pumping parameters or significant fluid disturbance within the wellbore.
[0369] 2. Application of professional rules;
[0370] The material and diameter of the sucker rod determine its elastic modulus and cross-sectional area. By combining this with segmented calculations of the static load and buoyancy effects based on the fluid density in the wellbore, the source of the lateral force can be further clarified. Current data shows that the maximum lateral force is concentrated at 336m, suggesting that there may be wellbore eccentricity or abrupt changes in fluid density near this depth.
[0371] - Stroke and stroke parameters have a significant impact on inertial force and vibration load. The current average stress ratio is low (0.3071), indicating that the total stress after superposition is not exceeded. However, the large fluctuation in lateral force suggests that the stroke and stroke parameters need to be optimized to reduce the impact of vibration load on the rod.
[0372] 3. Trend analysis and risk assessment;
[0373] - The current force distribution of the pole exhibits obvious local concentration, especially the high peak value of the lateral force at 336m. Long-term operation may lead to local fatigue damage or even the risk of fracture.
[0374] Although the stress ratio distribution is within a safe range, the large fluctuation in lateral force indicates that the overall stress on the rod is uneven, which may exacerbate the problem of uneven wear. It is recommended to take measures to improve the stress distribution as soon as possible.
[0375] - If pumping parameters are not adjusted or the wellbore environment is not optimized in time, early rod string failure or downhole malfunction may occur in the future.
[0376] III. Recommendations for Optimization Measures;
[0377] 1. Specific optimizations based on essential features;
[0378] - Regarding the issue of lateral force concentration: It is recommended to focus on monitoring at 336m, and to install a centralizer or adjust the sucker rod assembly to reduce localized force concentration.
[0379] - Reduce lateral force fluctuations: By optimizing the pumping unit stroke and stroke parameters, the influence of inertial forces and vibration loads is reduced, thereby improving the force distribution of the rod column.
[0380] 2. Principled recommendations based on professional rules;
[0381] - Re-examine the effects of static load and buoyancy based on the fluid density distribution in the wellbore to ensure that the sucker rod design meets the actual working conditions.
[0382] -Based on the material and diameter characteristics of the sucker rod, regularly check the elastic deformation and fatigue state of the rod post to avoid safety hazards caused by material performance degradation.
[0383] 3. Feasible on-site operating procedures;
[0384] - Install centralizer: Install centralizer near 336m to reduce uneven wear between rod string and wellbore, and alleviate the problem of lateral force concentration.
[0385] - Adjust pumping parameters: appropriately reduce the number of pumping strokes and optimize the stroke length to reduce the impact of inertial forces and vibration loads on the rod column.
[0386] -Strengthen monitoring: Regularly collect stress data on poles, focusing on the trend of lateral force changes at 336m, and promptly identify potential risks.
[0387] - Maintain the wellbore environment: Clean any foreign objects or deposits that may be present in the wellbore to ensure smooth fluid flow and reduce additional disturbance to the rod string.
[0388] The above suggestions are based on existing data analysis. When implementing them, the plan should be further refined according to the actual situation on site.
[0389] This analysis fully utilizes two essential features: lateral force and stress ratio. Based on real data, it systematically evaluates the stress state of the pole and proposes optimization measures.
[0390] The three examples above demonstrate that the method of the present invention can dynamically generate prompts that are highly adapted to different types of analysis tasks and output analysis results that meet professional requirements, thus possessing good task scalability and practicality.
[0391] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for analyzing the production status of pumping wells based on adaptive cue word generation, characterized in that, Includes the following steps: Step 1: Receive user requests, perform multi-source information retrieval, and obtain structured data input, professional knowledge fragments, and historical similar cases related to user requests; Step 1.1: Data Retrieval: Receive user requests in natural language text format, parse and extract target well number, time range, and task type information; query the pumping unit well production database based on the information to obtain corresponding raw production data, and call a pre-set feature analysis function to process the raw production data, generating structured data input containing statistical feature values; the pumping unit well production database stores raw monitoring and acquisition data related to pumping unit well production, including dynamometer card data, production dynamic data, and equipment data; wherein, dynamometer card data includes suspension point load and displacement, pump end load and displacement; production dynamic data includes daily fluid production, wellhead pressure, bottom hole temperature, bottom hole flowing pressure, reservoir static pressure, and dynamic fluid level height; equipment data includes pump mounting depth, pump diameter, stroke, stroke frequency, torque, motor power, and power consumption; Step 1.2: Knowledge Retrieval: The knowledge base stores expert reports and professional knowledge fragments formed by document processing related to the field of oil pumping unit production. Professional knowledge fragments relevant to user needs are retrieved from the knowledge base. First, keyword extraction and vectorization transformation are performed on user needs. Keyword retrieval and vector semantic retrieval are executed in parallel within the knowledge base to obtain keyword retrieval ranking lists and vector retrieval ranking lists, forming a preliminary ranking result. Then, the preliminary ranking result is re-ranked using a reciprocal ranking fusion algorithm to obtain a reciprocal ranking fusion score. Finally, based on the re-ranking result, the top performers with reciprocal ranking fusion scores exceeding a preset first threshold are returned. Individual pieces of professional knowledge form a set of professional knowledge fragments; Step 1.3: Perform case retrieval: The case library stores historical cases whose structured analysis results meet the quality standards. Retrieve historical cases similar to user needs from the case library. First, vectorize the user needs, converting them into user need vectors. Then, traverse the case library, calculating the cosine similarity between the current user need vector and the user need vector in each historical case. Finally, generate a similarity ranking list based on the cosine similarity calculation results, and return the top cases with similarity higher than a preset second threshold. A set of historical similar cases is formed from 10 historical cases; if there are no historical cases with a similarity higher than the preset second threshold, an empty set of historical similar cases is returned. Step 2: Based on the multi-source information retrieval results, the initial data analysis prompts and knowledge constraint prompts are generated in parallel using the data analysis prompt generation model and the knowledge constraint prompt generation model; Step 2.1: Generate data analysis prompts: Based on user needs and the structured data input retrieved in Step 1.1, and referring to the historical similar case set retrieved in Step 1.3, the data analysis prompt generation model is invoked to generate data analysis prompts focusing on the analysis of existing data features; if the historical similar case set is empty, then only data analysis prompts are generated based on user needs and the structured data input retrieved in Step 1.1; the data analysis prompts are structured text composed of several analysis points, with predefined start and end identifiers as boundaries, and each analysis point begins with a predefined list symbol; Step 2.2: Generate knowledge constraint prompts: Based on user needs and the previous search terms retrieved in Step 1.2 For each piece of professional knowledge, referring to the historical similar case set retrieved in step 1.3, the knowledge constraint prompt word generation model is invoked to generate knowledge constraint prompt words focusing on professional rule guidance; if the historical similar case set is empty, then only based on user needs and the previous cases retrieved in step 1.2... Each piece of professional knowledge generates a knowledge constraint prompt; the knowledge constraint prompt is a structured text composed of several professional rules, with predefined start and end identifiers as boundaries, and each rule starts with a predefined list symbol; Step 3: Based on the comprehensive analysis model, apply the initial data analysis prompts and knowledge constraint prompts to conduct a structured analysis of the production status of the pumping well, and conduct a quality assessment of the structured analysis results based on the quality assessment model; Step 4: If the quality assessment fails to meet the standards, the prompt words will be iteratively optimized based on the assessment feedback, and the analysis and assessment will be re-executed until a quality-compliant analysis result is generated. Step 5: Construct the analysis task of achieving the quality assessment standard into a new case, perform similarity matching with the case library, add non-duplicate cases to the case library, and update the case library.
2. The method for analyzing the production status of pumping wells based on adaptive cue word generation according to claim 1, characterized in that, In step 2.1, the working principle of the data analysis prompt word generation model is as follows: Data analysis prompt words are generated by inputting pre-set meta-prompt word instructions into the basic large language model. The specific working process is as follows: Roles and Inputs: The instruction model acts as a data analysis expert for the production status of pumping wells, receiving user requirements, structured data input, similar historical cases, and suggestions for improvement using prompts. Core task logic: If no suggestions for improvement are received, the model is required to identify the task type and extract specific values and features from the data input; If the set of similar historical cases is not empty, then prompt words containing multiple analysis points are generated by referring to the content of similar historical cases; if the set of similar historical cases is empty, then prompt words containing multiple analysis points are directly generated based on the data input and task type. If suggestions for improvement are received, the model is required to adjust the provided suggestions for improvement according to the specific type of suggestion. Output and Constraints: The prompts output by the model must conform to a predefined format, and each analysis point must follow strict data constraint principles, that is, the analysis is only based on the data features that are clearly present in the data input.
3. The method for analyzing the production status of pumping wells based on adaptive cue word generation according to claim 2, characterized in that, In step 2.2, the working principle of the knowledge constraint prompt word generation model is as follows: knowledge constraint prompt words are generated by inputting pre-set meta-prompt word instructions into the basic large language model. The specific working process is as follows: Roles and Inputs: The instruction model acts as a knowledge expert in the field of pumping well production status engineering, receiving user requirements, fragments of professional knowledge, similar historical cases, and suggestions for improvement based on prompts; Core task logic: If no suggestions for improvement are received, the model is required to filter knowledge directly related to the task type from professional knowledge fragments; If the set of similar cases is not empty, multiple core professional rules are extracted and generated by referring to the rule structure of similar cases; if the set of similar cases is empty, multiple core professional rules are directly extracted and generated based on the professional knowledge fragments; if suggestions for improvement of prompt words are received, the model is required to optimize the rules in the provided prompt words to be optimized according to the suggestions. Output and Constraints: The prompts output by the model must conform to a predefined format, and the extracted rules must be engineering practice-oriented, focusing on engineering principles, operating procedures or empirical principles, and directly used to guide the analysis process of the task.
4. The method for analyzing the production status of pumping wells based on adaptive cue word generation according to claim 3, characterized in that, The specific process of step 3 is as follows: Step 3.1: Generate Structured Analysis Results: Combining user needs with the structured data input retrieved in Step 1.1, and guided by the data analysis prompts generated in Step 2.1 and the knowledge constraint prompts generated in Step 2.2, the comprehensive analysis model is invoked to process the structured data input and generate structured analysis results. The structured analysis results are structured text organized according to preset hierarchical headings, including data display, professional analysis, and optimization suggestions. Step 3.2: Generate Quality Assessment Results: Based on user needs and the structured analysis results generated in Step 3.1, the quality assessment model is invoked to evaluate the structured analysis results according to preset quantitative assessment indicators, and the quality assessment results are output. The quantitative assessment indicators include the completeness of data feature extraction, analytical practicality, and logical consistency. The quality assessment results are structured text, containing assessment conclusion fields separated by specific identifiers, data analysis prompt words and their improvement suggestion fields, and knowledge constraint prompt words and their improvement suggestion fields. The evaluation conclusion field is a binary logic state. When the quality evaluation meets the standard, it is marked as "true" and each improvement suggestion field is confirmed as "quality meets the standard, no improvement is needed"; when the quality evaluation does not meet the standard, it is marked as "false" and each improvement suggestion field provides optimization suggestions for the corresponding prompt words.
5. The method for analyzing the production status of pumping wells based on adaptive cue word generation according to claim 4, characterized in that, In step 3.1, the working principle of the comprehensive analysis model is as follows: Structured analysis results are generated by inputting pre-set meta-prompt word instructions into the basic large language model. The specific working process is as follows: Roles and Inputs: The instruction model acts as a comprehensive analysis expert for the production status of pumping wells, receiving user requirements, structured data input, data analysis prompts, and knowledge constraint prompts. Analysis logic and structured output: The model is required to strictly follow the key points in the data analysis prompts and the rules in the knowledge constraint prompts, and to organize the output according to the preset template, which includes data processing and display, oil well production status analysis and optimization measures suggestions; Process constraints: The instructions require that the analysis must be based on the actual values in the structured data input, and that each essential engineering feature covered in the data analysis prompts be specifically analyzed, and that the rules in the knowledge constraint prompts be applied for qualitative judgment.
6. The method for analyzing the production status of pumping wells based on adaptive cue word generation according to claim 5, characterized in that, In step 3.2, the working principle of the quality assessment model is as follows: Quality assessment results are generated by inputting pre-set meta-prompt word instructions into the basic large language model. The specific working process is as follows: Roles and Inputs: The instruction model acts as a quality assessment expert, receiving user requirements, data analysis prompts, knowledge constraint prompts, and structured analysis results; Quantitative evaluation of execution: The instruction model determines the task type based on user needs and calls the corresponding list of essential engineering features. Then, based on predefined quantitative evaluation indicators, it calculates and scores the data feature extraction completeness, analytical practicality, and logical consistency of the analysis results. Among them, the data feature extraction completeness statistics include the number of essential engineering features explicitly mentioned and used in the analysis; analytical practicality is comprehensively scored from three aspects: use of essential feature data, executability of conclusions, and rationality of professional logic, for a total of 100 points; logical consistency is comprehensively scored from four aspects: coverage of key analysis points, application of knowledge constraints, completeness of reasoning chain, and lack of internal contradictions, for a total of 100 points. Achievement Judgment and Feedback Generation: The model is deemed to have met the standards if and only if the data feature extraction completeness is 100%, the analysis practicality score is ≥85, and the logical consistency score is ≥90. If the standards are not met, the model is required to generate specific optimization suggestions for the data analysis prompts and knowledge constraint prompts, respectively, and the suggestions must not require the addition of data that does not exist in the input data. Output format: Evaluation results and recommendations must be output in a predefined format.
7. The method for analyzing the production status of pumping wells based on adaptive cue word generation according to claim 6, characterized in that, The specific process of step 4 is as follows: Receive the quality assessment results generated in step 3.
2. If the content of the assessment conclusion field is "true", then the quality assessment meets the standard, terminate the iteration, and output the structured analysis results at this time. If the evaluation conclusion field is "false", the quality evaluation has not met the standard. The improvement suggestions for the data analysis prompts and the current version of the data analysis prompts will be returned to the data analysis prompt generation model to generate optimized data analysis prompts. At the same time, the improvement suggestions for the knowledge constraint prompts and the current version of the knowledge constraint prompts will be returned to the knowledge constraint prompt generation model to generate optimized knowledge constraint prompts. Subsequently, steps 3 and 4 are re-executed based on the optimized data analysis prompts and optimized knowledge constraint prompts to form a closed-loop iterative optimization process until the quality assessment meets the standard or the preset maximum number of iterations is reached. If the standard is not met after reaching the maximum number of iterations, the current optimal result is output and manual intervention is prompted.
8. The method for analyzing the production status of pumping wells based on adaptive cue word generation according to claim 1, characterized in that, The specific process of step 5 is as follows: Step 5.1: Construct a new structured case by combining the user requirements, final data analysis prompts, and knowledge constraint prompts related to this quality assessment achievement task. Step 5.2: Traverse all historical cases in the case library, and calculate the cosine similarity between the new case and each historical case in three parts: user needs, data analysis prompts, and knowledge constraint prompts; then, perform a weighted sum of the three cosine similarities to obtain the comprehensive similarity score. : ; in, , , The weights of user needs, data analysis prompts, and knowledge constraint prompts are respectively assigned, and ; Cosine similarity to user needs; Cosine similarity of prompt words for data analysis; Cosine similarity of knowledge-constrained prompt words; Step 5.3: Calculate the overall similarity score. With the preset duplicate detection threshold If a comparison is made, If it is, it is determined to be a new case and stored in the case database; if If the result is not found, it will be considered a duplicate case and will not be stored.