A large model reservoir measure recommendation method based on hierarchical reward preference optimization

CN122654877APending Publication Date: 2026-08-28QINGDAO UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611133713.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0008]针对现有技术中存在的上述技术问题,本发明提出了一种基于层次化门控奖励偏好优化的基于层次化奖励偏好优化的大模型油藏措施推荐方法,解决了现有技术中大语言模型用于油藏措施推荐时存在的只关注最终措施标签、缺少生产矛盾诊断约束、难以识别伪高分样本、输出逻辑一致性不足以及小样本训练稳定性不足的问题

Benefits of technology

[0043] 1. This invention breaks down reservoir measure recommendation into multiple stages: "parameter analysis, contradiction diagnosis, measure recommendation, and consistency verification," and uses a hierarchical gating reward function for joint constraints, enabling the model to learn a more complete "parameter-contradiction-measure" causal chain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122654877A_ABST
    Figure CN122654877A_ABST
Patent Text Reader

Abstract

The application discloses a large model oil reservoir measure recommendation method based on hierarchical reward preference optimization and belongs to the technical field of oil and gas field development and artificial intelligence. The application is used for solving the problems that a large language model lacks causal reasoning constraints in oil reservoir measure recommendation, is prone to produce pseudo-high-score responses and is unstable in small sample training. The method acquires production dynamic data to construct an oil reservoir dynamic analysis sample library; constructs a supervised fine-tuning sample training reference model; constructs a hierarchical gating reward function containing a contradiction diagnosis reward, a measure recommendation reward, a consistency constraint reward and a hard constraint gate; the measure recommendation reward is constrained by the contradiction diagnosis gate; candidate responses are generated and preferred and unpreferred preferences are constructed; a multi-reference preference optimization algorithm is used to train the model; and finally, a structured recommendation result is output. The application learns a "parameter-contradiction-measure" causal decision chain by the hierarchical gating reward constraint model, and improves the recommendation accuracy and logical consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of oil and gas field development and artificial intelligence technology, specifically involving a large-scale reservoir measure recommendation method based on hierarchical reward preference optimization. Background Technology

[0002] In the development of oil and gas fields, reservoir dynamic analysis is a core technical aspect for maintaining efficient oil and gas production. Field technicians typically need to integrate multi-source information (such as daily oil production, daily fluid production, water cut, water-oil ratio, pressure, water injection volume, injection-production response, historical action records, etc.) to determine the main production problems existing in the target well or well group, and propose engineering measures accordingly (such as profile modification, water shut-off, acidizing to remove blockages, fracturing, perforation repair, fluid extraction and pump replacement, major overhaul and maintenance, maintaining the status quo, or strengthening monitoring, etc.).

[0003] Traditional reservoir protection recommendations rely primarily on manual experience and operational rules, resulting in low analytical efficiency, inconsistent evaluation standards, difficulty in accumulating expert experience, and weak capacity for handling large-scale well operations. With the advancement of oilfield digitalization, machine learning has been applied to individual tasks such as production forecasting, water cut prediction, operational condition identification, injection-production connectivity analysis, and measure effectiveness analysis. However, existing models mostly serve independent prediction or classification tasks, failing to construct a complete causal decision-making chain from production parameters to development conflicts and then to engineering measures. Furthermore, they lack structured reasoning processes, making it difficult for engineers to verify and trace the source of information.

[0004] In recent years, large language models, with their powerful natural language understanding, knowledge organization, and reasoning generation capabilities, have provided new technical approaches for reservoir dynamic analysis and policy recommendations. Supervised fine-tuning can guide large language models to output data analysis, contradiction diagnosis, and policy recommendations in a standardized, structured format. However, relying solely on supervised fine-tuning has the following limitations;

[0005] First, the model tends to learn shortcuts, directly memorizing high-frequency measure labels, lacking a complete contradiction reasoning process based on production parameters. Second, there is a disconnect between the reasoning logic and the output results; the model may diagnose logical errors but the measure labels may be correct by chance, and this "pseudo-high score" response can mislead performance judgments. Third, traditional preference optimization methods only evaluate the overall output, making it difficult to separately constrain multiple reasoning elements such as parameter basis, contradiction identification, measure matching, and causal logic consistency. Fourth, reservoir labeled samples are scarce and class imbalanced, making reinforcement learning training prone to policy drift and posing a risk of overoptimization in later stages.

[0006] Existing intelligent analysis solutions for oil and gas wells primarily rely on enhanced retrieval, similar case retrieval, or the invocation of professional simulation models to achieve automatic report generation. Their core objective is merely to improve diagnostic output efficiency, without guiding the model to learn the causal decision-making chain of "parameters-contradictions-measures" during the training phase. Currently, various reinforcement learning and preference optimization frameworks lack hierarchical reward mechanisms adapted to reservoir mechanisms, making it impossible to finely evaluate the quality of segmented reasoning or effectively filter and penalize irregular predictions that do not meet engineering practice requirements.

[0007] Therefore, there is an urgent need to propose a hierarchical gating reward preference optimization method for reservoir measure recommendation tasks, which transforms reservoir dynamic analysis rules, contradiction diagnosis logic, measure adaptation relationship and hard constraint penalties into trainable reward signals, enabling large language models to recommend measures based on the correct identification of production contradictions. Summary of the Invention

[0008] To address the aforementioned technical problems in existing technologies, this invention proposes a large-scale reservoir measure recommendation method based on hierarchical gating reward preference optimization. This method solves the problems in existing technologies where large language models used for reservoir measure recommendation only focus on the final measure label, lack production contradiction diagnosis constraints, have difficulty identifying pseudo-high-scoring samples, have insufficient output logic consistency, and have insufficient training stability with small samples.

[0009] To achieve the above objectives, the present invention adopts the following technical solution.

[0010] A method for recommending reservoir measures in a large model based on hierarchical reward preference optimization includes the following steps;

[0011] S1. Obtain the production dynamic data of the target oil well or well group, clean the production dynamic data, extract features and organize labels, and construct a reservoir dynamic analysis sample library. The reservoir dynamic analysis sample library includes key production parameters, parameter change trends, production contradiction labels and measure type labels.

[0012] S2. Based on the reservoir dynamic analysis sample library, construct a supervised fine-tuning sample, so that the expected output of each supervised fine-tuning sample includes at least a data analysis segment, a contradiction diagnosis segment, and a measure recommendation segment. Use the supervised fine-tuning sample to train a basic large language model to obtain a reference model with reservoir dynamic analysis output format.

[0013] S3. Construct a hierarchical gating reward function for reservoir measure recommendation tasks. The hierarchical gating reward function includes at least a conflict diagnosis reward, a measure recommendation reward, a consistency constraint reward, and a hard constraint gating, wherein the measure recommendation reward is constrained by the gating threshold of the conflict diagnosis reward.

[0014] S4. Generate multiple candidate responses for the same input sample, calculate the reward score of each candidate response using the hierarchical gating reward function, and construct a preference pair consisting of the preferred response and the inferior response based on the reward score, reward gap and hard constraint gating results.

[0015] S5. Based on the preference pairs, the reference model is trained using a reward-guided preference optimization algorithm to obtain a reservoir measures recommendation model;

[0016] S6. Input the production dynamic data of the oil well or well group to be analyzed into the reservoir measure recommendation model, and output a structured result including data analysis, production contradiction diagnosis, measure recommendation and recommendation basis.

[0017] Preferably, the production dynamic data in step S1 includes one or more of the following: daily oil production, daily liquid production, daily water production, water cut, water-oil ratio, bottom hole pressure, oil pressure, casing pressure, water injection volume, and injection-production response parameters.

[0018] The key production parameters include one or more of the following: water cut, water cut variation trend, water-oil ratio, daily oil production variation trend, daily liquid production variation trend, pressure variation trend, injection-production response characteristics, production decline rate, and water cut surge magnitude.

[0019] The production contradiction labels include one or more of the following: normal production, advantageous channel or ineffective circulation, ineffective water injection or insufficient formation energy, water injection surge or flooding, well blockage or contamination, insufficient fluid supply, and excessively rapid decline in production capacity.

[0020] The action type labels include one or more of the following: shutting in or suspending production, fracturing, perforation repair, acidizing to remove blockages, profile control, water shut-off, overhaul or maintenance operations, fluid extraction or pump replacement, maintaining the status quo or enhanced monitoring.

[0021] Preferably, the structured results in step S6 include;

[0022] The data analysis section is used to list key production parameters, parameter change trends, and abnormal characteristics;

[0023] The contradiction diagnosis section is used to identify the main types of production contradictions in the current oil well or well group and the basis for their diagnosis.

[0024] The recommended measures section provides the types of engineering measures that match the main production contradictions, along with the reasons for the recommendations, applicable conditions, and risk warnings.

[0025] Preferably, the total reward of the hierarchical gating reward function in step S3 is obtained by weighting the contradiction diagnosis reward and the measure recommendation reward through hard constraint gating, and then superimposing the consistency constraint reward; the measure recommendation reward is constrained by the contradiction diagnosis gating, which enables the measure recommendation reward when the contradiction diagnosis reward is greater than a preset threshold, and disables the measure recommendation reward otherwise.

[0026] Preferably, the contradiction diagnosis reward is obtained by multiplying the contradiction type identification score by the parameter-contradiction causality consistency factor; the contradiction type identification score is used to evaluate the degree of consistency between the production contradiction type output by the model and the standard production contradiction label; the parameter-contradiction causality consistency factor is used to evaluate whether the key production parameters referenced or analyzed by the model support its output production contradiction diagnosis.

[0027] The parameter-contradictory causal consistency factor is obtained by weighted summation of three factors: key parameter coverage, key trend judgment accuracy, and key value or interval rationality. The key parameter coverage represents the ratio between the number of key parameters mentioned by the model and the number of key parameters required for the contradiction. The key trend judgment accuracy represents the correct proportion of key parameter trend judgments. The key value or interval rationality represents whether the key value or interval referenced by the model is reasonable.

[0028] Preferably, the measure recommendation reward is obtained by multiplying the contradiction diagnosis gating function, the contradiction-measure causal consistency factor, and the measure type matching score; the contradiction diagnosis gating function takes a value of 1 when the contradiction diagnosis reward is greater than a preset threshold and a value of 0 when the contradiction diagnosis reward is less than or equal to the preset threshold, so as to constrain the measure recommendation reward by the contradiction diagnosis result; the contradiction-measure causal consistency factor is used to evaluate whether the recommended measures are applicable to the diagnosed production contradiction; the measure type matching score is used to evaluate the degree of consistency between the recommended measure type and the standard measure type label.

[0029] Preferably, the consistency constraint reward is a negative penalty, determined by the sum of the products of the penalty weights of various consistency violation rules and the number of occurrences; the consistency violation rules include one or more of the following: parameter-contradiction inconsistency, contradiction-measure inconsistency, data-conclusion inconsistency, trend judgment error, measure application condition error, and engineering feasibility error;

[0030] The hard constraint gating determines its value based on whether the consistency constraint reward is lower than the hard constraint threshold: when the consistency constraint reward is lower than the hard constraint threshold, the value is 0, which sets the contradiction diagnosis reward and the measure recommendation reward to zero; when the consistency constraint reward is not lower than the hard constraint threshold, the value is 1, which retains the contradiction diagnosis reward and the measure recommendation reward.

[0031] Preferably, the method of constructing preference pairs in step S4 includes: generating no less than two candidate responses for the same input sample, calculating the total reward for each candidate response, determining the candidate response with the highest reward score or that meets the preset high score condition as the preferred response, determining the candidate response with the lowest reward score, that is lower than the preset quantile or that triggers hard constraint gating as the undesired response, and calculating the reward gap between the preferred response and the undesired response.

[0032] The candidate responses include one or more of the following: responses generated by sampling from a reference model, responses rewritten from historical expert answers, responses generated by rule perturbation, and responses generated by different decoding parameters.

[0033] Preferably, the reward-guided preference optimization algorithm in step S5 adopts multi-reference preference optimization, and a fixed reference model and a dynamic reference model are set during the training process; the fixed reference model keeps the initial reference strategy unchanged, and the dynamic reference model is updated with the training process; the training checkpoint is determined comprehensively based on the dynamic reference preference interval, the verification loss and the preference discrimination accuracy; a course-based reward scheduling mechanism is set during the training process so that the reward guidance intensity gradually intervenes from weak to strong with the training progress.

[0034] Furthermore, this invention also mentions a large-scale reservoir measure recommendation system based on hierarchical reward preference optimization, comprising:

[0035] The data processing module is used to acquire dynamic production data of oil wells or well groups and generate key production parameters, parameter change trends, production conflict labels, and measure type labels.

[0036] The supervised fine-tuning module is used to construct structured supervised fine-tuning samples and train the basic large language model to obtain the reference model;

[0037] The reward calculation module is used to calculate the contradiction diagnosis reward, measure recommendation reward, consistency constraint reward, hard constraint gating result and total reward of the candidate response based on the hierarchical gating reward function;

[0038] The preference pair construction module is used to construct preferred and unpreferred response preference pairs based on the total reward of candidate responses, the reward gap, and the hard constraint gating results.

[0039] The preference optimization training module is used to perform reward-guided preference optimization training on the reference model based on the preference pairs.

[0040] The checkpoint selection module is used to select training checkpoints based on the dynamic reference preference interval, validation loss, and preference discrimination accuracy.

[0041] The inference output module is used to output data analysis, production conflict diagnosis, and recommended measures results based on the production dynamic data of the oil well or well group to be analyzed.

[0042] The beneficial technical effects brought about by this invention;

[0043] 1. This invention breaks down reservoir measure recommendation into multiple stages: "parameter analysis, contradiction diagnosis, measure recommendation, and consistency verification," and uses a hierarchical gating reward function for joint constraints, enabling the model to learn a more complete "parameter-contradiction-measure" causal chain.

[0044] 2. This invention uses the contradiction diagnosis reward as a pre-gating of the measure recommendation reward, which can suppress pseudo-high-scoring responses where the contradiction diagnosis is wrong but the measure label is accidentally hit, thereby improving the engineering credibility of the measure recommendation.

[0045] 3. This invention identifies parameter-contradiction inconsistencies, contradiction-measure inconsistencies, trend judgment errors, and engineering feasibility errors through consistency constraints and hard constraint gating, making the model output more consistent with the reservoir development mechanism.

[0046] 4. This invention constructs preference pairs by reward gap and introduces reward weighting, reward gap regularization and hard constraint penalty in preference optimization, so that the model can obtain more stable preference learning effect under small sample reservoir data conditions.

[0047] 5. This invention employs a multi-reference preference optimization strategy that combines fixed and dynamic references, which can reduce the risk of policy drift in the training of large language models in small-sample reinforcement learning in oil reservoirs.

[0048] 6. The present invention outputs structured results, which facilitates engineers to review key parameters, diagnostic basis and applicable conditions of measures, and is conducive to the accumulation of experience in dynamic analysis of oilfields and the screening of measures for batch well groups. Attached Figure Description

[0049] Figure 1 The flowchart illustrates the reservoir conflict diagnosis and multi-reference preference optimization method provided in this embodiment of the invention.

[0050] Figure 2 The diagram shows the structure of the reservoir hierarchical gating reward function provided in this embodiment of the invention.

[0051] Figure 3 The training and validation loss curves provided for embodiments of the present invention.

[0052] Figure 4 This diagram illustrates the effect of separating the preferred and undesired responses as provided in an embodiment of the present invention.

[0053] Figure 5 The preference-reward gap distribution diagram provided for embodiments of the present invention.

[0054] Figure 6 This is an optimal checkpoint selection diagram provided for embodiments of the present invention.

[0055] Figure 7 The mean value chart of the distinguishing ability of the gated reward sub-items provided in the embodiments of the present invention.

[0056] Figure 8 This is a diagram of core evaluation indicators for reservoir operations provided in an embodiment of the present invention. Detailed Implementation

[0057] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0058] Example 1: Recommended methods for reservoir measures.

[0059] This embodiment provides a method for recommending reservoir measures in a large model based on hierarchical reward preference optimization, such as... Figure 1 As shown in the figure, this diagram illustrates the complete technical process from data source, rule base, reference model to the final structured output result, including core steps such as dynamic sample library construction, supervised fine-tuning, hierarchical gating reward function, preference pair construction, multi-reference preference optimization training, and measure recommendation output; specifically, it includes the following steps;

[0060] S1. Obtain production dynamic data of the target oil well or well group and construct a reservoir dynamic analysis sample library.

[0061] This step focuses on a target oil well or well group in a specific oilfield block, collecting its production dynamic data. The production dynamic data is then cleaned, its features extracted, and tagged to construct a reservoir dynamic analysis sample library. This library includes key production parameters, parameter change trends, production conflict tags, and measure type tags.

[0062] Production dynamic data includes one or more of the following: daily oil production, daily fluid production, daily water production, water cut, water-oil ratio, bottom hole pressure, oil pressure, casing pressure, water injection volume, and injection-production response parameters.

[0063] The key production parameters include one or more of the following: water cut, water cut variation trend, water-oil ratio, daily oil production variation trend, daily liquid production variation trend, pressure variation trend, injection-production response characteristics, production decline rate, and water cut surge magnitude.

[0064] The production contradiction labels include one or more of the following: normal production, advantageous channel or ineffective circulation, ineffective water injection or insufficient formation energy, water injection surge or flooding, well blockage or contamination, insufficient fluid supply, and excessively rapid decline in production capacity.

[0065] The action type labels include one or more of the following: shutting in or suspending production, fracturing, perforation repair, acidizing to remove blockages, profile control, water shut-off, overhaul or maintenance operations, fluid extraction or pump replacement, maintaining the status quo or enhanced monitoring.

[0066] Step S1 further includes the following processing procedures;

[0067] S1.1 Unify the well names, times, units of measurement, and field names in different data sources to ensure data consistency.

[0068] S1.2. Remove, impute, or mark missing values, outliers, and data that clearly do not conform to common sense in engineering. For example, remove abnormal data such as negative daily oil production or data that exceeds a reasonable range.

[0069] S1.3 Calculate key characteristics such as water cut change trend, water-oil ratio change trend, daily oil production change trend, daily liquid production change trend, pressure change trend, injection-production response characteristics, production decline rate, and water cut surge magnitude.

[0070] S1.4 Generate production conflict labels and measure type labels based on reservoir engineering rules and expert annotations. Production conflict labels include one or more of the following: normal production, dominant channel or ineffective circulation, ineffective water injection or insufficient formation energy, water injection surge or water flooding, wellbore blockage or contamination, insufficient fluid supply, and excessively rapid production decline. Measure type labels include one or more of the following: well shut-in or production suspension, fracturing, perforation repair, acidizing to remove blockages, profile control, water shut-off, major workover or maintenance, fluid extraction or pump replacement, maintaining the status quo or enhanced monitoring.

[0071] S1.5 Organize key production parameters, trend characteristics, production contradiction labels, and measure type labels into structured samples.

[0072] In one specific implementation, production conflict labels include one or more of the following: normal production, dominant channel or ineffective circulation, ineffective water injection or insufficient formation energy, water injection surge or water flooding, wellbore blockage or contamination, insufficient fluid supply, and excessively rapid production decline; measure type labels include one or more of the following: well shut-in or production suspension, fracturing, perforation repair, acidizing to remove blockages, profile control, water shut-off, major workover or maintenance, fluid extraction or pump replacement, maintaining the status quo or enhanced monitoring. For different reservoir types and different development stages, conflict labels, measure labels, and parameter thresholds can be adjusted according to the characteristics of the block development.

[0073] S2. Based on the reservoir dynamic analysis sample library, construct a supervised fine-tuning sample and train a reference model.

[0074] The structured samples obtained in step S1 are transformed into supervised fine-tuning samples. Each supervised fine-tuning sample includes input text and expected output text. The input text describes the production dynamics, key production parameters, and trend characteristics of the target well or well group. The expected output text adopts a uniform format and includes at least three parts: data analysis, contradiction diagnosis, and recommended measures.

[0075] In one specific implementation, the output text is expected to be organized according to the following structure;

[0076] [Data Analysis] List key parameters such as water cut, daily oil production, daily liquid production, and pressure, along with their changing trends, and describe any abnormal characteristics;

[0077] [Contradiction Diagnosis] Identify the main types of production contradictions and explain the correspondence between these contradictions and key parameters;

[0078] [Recommended Measures] Provide the type of recommended measures and explain the compatibility between the measures and the diagnosed production contradictions, applicable conditions, and precautions.

[0079] Through supervised fine-tuning in step S2, the basic large language model learns reservoir-related terminology, structured output formats, and fundamental dynamic analysis logic. The model obtained in this stage serves as a reference model for subsequent preference optimization stages, rather than being the final model used for measure recommendation.

[0080] S3. Construct a hierarchical gated reward function for reservoir measure recommendation tasks.

[0081] This embodiment constructs a hierarchical gated reward function for reservoir measure recommendation tasks. For example... Figure 2 As shown in the figure, the process of calculating the contradiction diagnosis reward, the measure recommendation reward, and the consistency constraint reward after the candidate response is parsed is illustrated. The reward is then hierarchically controlled through diagnostic gating and hard constraint gating, and finally outputs the total reward for response ranking and training weights.

[0082] The reward function includes a conflict diagnosis reward, a measure recommendation reward, a consistency constraint reward, and a hard constraint gating. Its core constraint is that the measure recommendation reward cannot be activated independently of the production conflict diagnosis; it is only activated when the conflict diagnosis reaches a preset credibility level.

[0083] The total reward of the hierarchical gating reward function is:

[0084] ;

[0085] in, Rewards for diagnosing conflicts Recommendation and reward for measures Rewards for logical consistency constraints It is a hard constraint gating, , and This is the weighting coefficient. Preferably, , , Those skilled in the art can adjust the above weights according to oilfield type, sample size, and distribution of measures.

[0086] The formula for calculating the reward for conflict diagnosis is as follows: .

[0087] in, The contradiction type identification score is used to evaluate the consistency between the production contradiction type output by the model and the standard production contradiction label; The parameter-contradictory causal consistency factor is used to evaluate whether the key production parameters referenced or analyzed by the model support the production contradiction diagnosis it outputs.

[0088] The parameter—the contradictory causal consistency factor—is calculated in the following manner;

[0089] ;

[0090] in, For key parameter coverage; To assess the accuracy of key trend judgments; For the rationality of key values ​​or ranges; , and These are the weighting coefficients, and .

[0091] Preferably, , , .

[0092] Taking sudden water injection or flooding as examples, candidate responses should focus on key characteristics such as rapid increase in water cut, decrease in oil production, and increase in water-oil ratio; if the candidate response only provides conclusions but does not cite supporting parameters, then... This design reduces the likelihood of parameter information being treated as an independent reward item, but rather as a supporting factor for contradiction diagnosis, thus preventing the model from obtaining high scores by simply piling up parameter names.

[0093] The formula for calculating the incentive for the recommended measures is as follows;

[0094] ;

[0095] in, For contradiction diagnosis gating function, when Greater than the preset threshold The value is 1 if the condition is met, and 0 otherwise. The contradiction-measure causal consistency factor is used to evaluate whether the recommended measures are applicable to the diagnosed production contradictions. The measure type matching score is used to evaluate the degree of consistency between the recommended measure type and the standard measure type label. Preferably, the value is 0.3.

[0096] The core of this design lies in the requirement that measure rewards must be subject to a contradiction diagnosis gating system. When the contradiction diagnosis output by the model fails to reach a preset confidence level, the contradiction diagnosis gating function is closed, preventing the measure recommendation reward from taking effect, thereby suppressing candidate responses where the contradiction diagnosis is incorrect but the measure type happens to match.

[0097] Consistency constraints reward negative penalties.

[0098] ;

[0099] in, For the first Penalty weights for class consistency violations For the first The number of times a consistency violation rule occurs; the consistency violation rule includes one or more of the following: parameter-contradiction inconsistency, contradiction-measure inconsistency, data-conclusion inconsistency, trend judgment error, measure application condition error, and engineering feasibility error.

[0100] Specifically, engineering logic violations include inconsistencies between parameter trends and contradiction diagnosis, inconsistencies between production contradictions and recommended measures, recommending strong intervention measures for normal production samples, recommending acidification to resolve blockages when there is a lack of evidence of blockage, and interpreting samples with sudden increases in water content as insufficient liquid supply.

[0101] Hard constraint gating is;

[0102] ;

[0103] in, The hard constraint threshold is set to 0 when a candidate response triggers a serious engineering logic violation, thus reducing the contradiction diagnosis reward and measure recommendation reward to zero and preventing the model from obtaining false high scores due to local label hits.

[0104] S4. Generate multiple candidate responses for the same input sample and construct preference pairs.

[0105] Multiple candidate responses are generated for the same input sample. Candidate responses can come from sampled outputs of the reference model under different decoding parameters, rewritten text of historical expert answers, or responses with different error patterns constructed through rule perturbations. Candidate responses can include high-quality responses, medium-quality responses, low-quality responses, and empty responses.

[0106] A high-quality response is one in which the key parameters, contradiction diagnosis, and recommended measures are consistent with the annotations and engineering rules; a medium-quality response is one in which the key parameter analysis is basically correct but there are deviations in the contradiction diagnosis or the adaptation of measures; a low-quality response is one in which there are obvious errors in the key parameters, contradiction diagnosis, and recommended measures; and a vague response is one in which only general observations or maintaining the status quo are output, lacking sufficient parameter basis.

[0107] The total reward and reward sub-items for each candidate response are calculated using the hierarchical gating reward function in step S3. For the same input sample, responses with higher reward scores and no hard constraint violations are identified as preferred responses, while responses with lower reward scores or hard constraint violations are identified as inferior responses, and the reward gap is calculated.

[0108] reward_gap = R total (chosen) - R total (rejected);

[0109] When the reward gap (reward_gap) is greater than a preset threshold, the corresponding preferred response and unpreferred response are constructed as a preference pair; when the reward gap (reward_gap) is too small or there are labeling conflicts among samples, the training weight of the preference pair is reduced or it is removed.

[0110] The preference pairs obtained in the above manner not only include response ranking information, but also reward gap, hard constraint violation indicators, and reward sub-item information.

[0111] S5. Based on the preference pairs, the reference model is trained using a reward-guided preference optimization algorithm.

[0112] Starting with the reference model obtained in step S2, the current policy model is trained using reward-guided multi-reference preference optimization. The training objective is to increase the probability of generating preferred responses and decrease the probability of generating inferior responses by maintaining the model's language generation capability and the stability of its structured output.

[0113] In one specific implementation, the training loss includes a basic preference optimization loss, a reward gap weighting term, a reward gap regularization term, and a hard constraint penalty term. For preference pairs with large reward gaps, their training weights are increased; for preference pairs whose inferior responses trigger hard constraint gating, additional penalties are applied to keep the model away from engineering-unacceptable outputs.

[0114] To reduce policy drift during training with small reservoir samples, a fixed reference model and a dynamic reference model are set up during training. The fixed reference model keeps the initial reference policy unchanged and is used to measure the separation gain of the current model relative to the initial reference model; the dynamic reference model is updated during the training process and is used to constrain the preference interval and training stability during the current policy update process.

[0115] In the early stages of training, reward signals may contain noise, and imposing strong reward constraints too early can easily lead to unstable model output. Therefore, a curriculum-based reward scheduling mechanism is set up during training, so that the intensity of reward guidance gradually increases from weak to strong as training progresses. In the early stages of training, the focus is on format learning and preference direction learning, while in the middle and later stages of training, reward gap weighting and hard constraint penalties are gradually increased.

[0116] Checkpoint selection: Training checkpoints are determined comprehensively based on the dynamic reference preference interval, validation loss, and preference discrimination accuracy, avoiding overfitting checkpoint selection solely based on training loss or a fixed reference interval. The fixed reference preference interval is used to measure the separation gain relative to the initial reference model, while the dynamic interval and validation loss are used to help determine generalization stability.

[0117] S6. Input the production dynamic data of the oil well or well group to be analyzed into the reservoir measure recommendation model and output the structured results.

[0118] The production dynamic data of the oil well or well group to be analyzed is input into the reservoir measure recommendation model, and the output is a structured result that includes data analysis, production contradiction diagnosis, measure recommendation and recommendation basis.

[0119] The production dynamic data of the oil well or well group to be analyzed are input into the trained large language model, and the model outputs structured recommendation results. These structured recommendation results include key parameter analysis results, production contradiction diagnosis results, recommended measure types, applicable conditions for the measures, and risk warnings.

[0120] For example, when the input sample shows a rapid increase in water cut, a significant decrease in daily oil production, an increase in the water-oil ratio, and abnormal injection-production response, the model lists the above parameters and trends in the key parameter analysis section; in the production contradiction diagnosis section, it judges that there is a risk of water injection surge or flooding; in the measure recommendation section, it provides suggestions for measures such as limited injection, profile adjustment, water shut-off, or controlled water injection, and explains the basis and applicable conditions for the recommendations.

[0121] For example, when the input sample shows a sudden drop in daily production, relatively stable water cut, and pressure changes that do not support the judgment of insufficient formation energy, the model tends to judge wellbore blockage or near-well contamination in the production contradiction diagnosis section, and prioritizes acidizing to unblock, pump inspection or major maintenance in the measure recommendation section, rather than simply recommending profile control and water shut-off.

[0122] Model Evaluation: During the model validation phase, samples are divided into training and test sets at an 8:2 ratio. Evaluation metrics include contradiction Macro-F1, measure Top-1 accuracy, end-to-end exact match rate, and constraint violation rate.

[0123] Among them, the contradiction Macro-F1 is used to evaluate the ability to classify production contradictions and reduce the impact of class imbalance; the measure Top-1 accuracy is used to evaluate whether the primary measure type matches the expert annotation; the end-to-end exact matching rate is used to evaluate whether the contradiction diagnosis, measure recommendation and hard constraint verification simultaneously meet the requirements; and the constraint violation rate is used to evaluate whether there are any violations in the model output that are unacceptable in engineering logic.

[0124] Example 2: Reservoir Measures Recommendation System.

[0125] This embodiment provides a large language model reservoir measure recommendation system based on hierarchical gating reward preference optimization, including a data processing module, a supervised fine-tuning module, a reward calculation module, a preference pair construction module, a preference optimization training module, a checkpoint selection module, and an inference output module.

[0126] The data processing module is used to acquire dynamic production data of oil wells or well groups and generate key production parameters, parameter change trends, production conflict labels, and measure type labels.

[0127] The supervised fine-tuning module is used to construct structured supervised fine-tuning samples and train the basic large language model to obtain the reference model;

[0128] The reward calculation module is used to calculate the contradiction diagnosis reward, measure recommendation reward, consistency constraint reward, hard constraint gating result and total reward of the candidate response based on the hierarchical gating reward function;

[0129] The preference pair construction module is used to construct preferred and unpreferred response preference pairs based on the total reward of candidate responses, the reward gap, and the hard constraint gating results.

[0130] The preference optimization training module is used to perform reward-guided preference optimization training on the reference model based on the preference pairs.

[0131] The checkpoint selection module is used to select training checkpoints based on the dynamic reference preference interval, validation loss, and preference discrimination accuracy.

[0132] The inference output module is used to output data analysis, production conflict diagnosis, and recommended measures results based on the production dynamic data of the oil well or well group to be analyzed.

[0133] The aforementioned system can be deployed on oilfield production data platforms, reservoir dynamic analysis platforms, or independent large language model inference services. For oilfield intranet environments where production data cannot be uploaded, model training and inference can both be deployed on local private servers. For scenarios requiring cross-block reuse, the reward function rule base, contradiction label system, and measure label system can be maintained as configuration files.

[0134] Example 3: Experimental verification and effect analysis.

[0135] To verify the effectiveness of the present invention, this embodiment was experimentally verified using actual reservoir data.

[0136] Dataset and experimental setup.

[0137] This embodiment uses 67 basic well samples to construct a reservoir dynamic analysis dataset, and divides it into a training set and a test set in an 8:2 ratio, with the training set consisting of 53 basic well samples and the test set consisting of 14 basic well samples. Based on multiple candidate responses under the same basic well sample, a hierarchical gated reward function is used to construct 524 reward preference samples. The 524 samples represent the number of preference pairs, which is not the same as the number of basic well samples.

[0138] During training, the training loss, validation loss, implicit reward scores for preferred and inferior responses, preference-to-reward gap, validation reward interval at each checkpoint, and statistics on reward sub-items are recorded.

[0139] Training process analysis.

[0140] Figure 3 This demonstrates how the training loss and validation loss change with the number of training steps. From Figure 3 It can be seen that the training loss and validation loss gradually decrease and stabilize as the number of training steps increases, and the model continues to converge during the training process.

[0141] Figure 4 The study demonstrates the reward separation effect between the preferred and undesirable responses. The reward score for the preferred response is 5.469, while the reward score for the undesirable response is 1.059, showing a clear gap between the two, indicating that the model can effectively distinguish between the preferred and undesirable responses.

[0142] Figure 5 This illustrates the distribution of the preference-reward gap. From... Figure 5 It can be seen that the average reward gap between preference pairs is 3.37, and the reward gap is mainly distributed in the positive region, indicating that the preferred response has a stable reward advantage over the inferior response, and the preference pairs constructed by the hierarchical gating reward function have significant quality differences.

[0143] Checkpoint selection.

[0144] like Figure 6 As shown, the training checkpoint is determined comprehensively based on the dynamic reference preference interval, validation loss, and preference discrimination accuracy. In the specific implementation described above, the training and validation loss curves, as well as the optimal checkpoint selection curve, all point to the 100th checkpoint. The validation reward interval near the 100th checkpoint reaches the main peak of this training round, and the validation loss does not show abnormal divergence. Therefore, the 100th checkpoint is selected as the optimal model checkpoint.

[0145] Analysis of the ability to differentiate between reward sub-items.

[0146] Figure 7 This demonstrates the ability of each reward sub-item—including conflict type matching, parameter-conflict consistency, conflict diagnosis reward, diagnosis gating opening rate, conflict-measure consistency, measure type matching, and measure recommendation reward—to distinguish between preferred and inferior responses. Figure 7 It can be seen that the average total reward for the preferred response is 0.94, while the average total reward for the undesirable response is -2.43. Sub-items such as contradiction type matching score, parameter-contradiction causal consistency, contradiction diagnosis reward, diagnosis gating opening rate, contradiction-measure causal consistency, measure type matching score, and measure recommendation reward all show that the preferred response is higher than the undesirable response. This result indicates that the hierarchical gating reward function can effectively distinguish response quality from multiple stages, including contradiction diagnosis, measure adaptation, and consistency constraints.

[0147] Evaluation of core business metrics.

[0148] Figure 8 The presentation showcases four core business evaluation metrics for the trained model on the test set: the macro-average F1 score for contradictions, the accuracy of the preferred measure, the end-to-end exact match rate, and the constraint violation rate. The macro-average F1 score for contradictions is 100%, the accuracy of the preferred measure is 100%, the end-to-end exact match rate is 97%, and the constraint violation rate is 3%. Specifically, the macro-average F1 score for contradictions evaluates the model's ability to classify production contradictions and mitigate the impact of class imbalance; the accuracy of the preferred measure evaluates whether the primary measure type matches expert annotations; the end-to-end exact match rate evaluates whether contradiction diagnosis, measure recommendation, and hard constraint verification simultaneously meet requirements; and the constraint violation rate evaluates whether the model output contains any logically unacceptable violations.

[0149] Therefore, the evaluation results not only reflect whether the model training loss has been reduced, but also whether the model has formed a verifiable reservoir "parameter-contradiction-measure" reasoning chain.

[0150] This invention can be applied to dynamic analysis of oilfield development, recommendation of single-well measures, screening of well group management schemes, intelligent matching of measures databases, and reservoir development auxiliary decision-making systems. This method can utilize existing production dynamic data, expert labels, and historical measures records to construct training samples. In small-sample scenarios, it improves the engineering applicability of large language models through hierarchical gating rewards and multi-reference preference optimization.

[0151] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A method for recommending reservoir measures in a large-scale model based on hierarchical reward preference optimization, characterized in that, Includes the following steps: S1. Obtain the production dynamic data of the target oil well or well group, clean the production dynamic data, extract features and organize labels, and construct a reservoir dynamic analysis sample library. The reservoir dynamic analysis sample library includes key production parameters, parameter change trends, production contradiction labels and measure type labels. S2. Construct supervised fine-tuning samples based on the reservoir dynamic analysis sample library, so that the expected output of each supervised fine-tuning sample includes at least a data analysis segment, a contradiction diagnosis segment, and a measure recommendation segment. Use the supervised fine-tuning samples to train the basic large language model and obtain a reference model with the reservoir dynamic analysis output format. S3. Construct a hierarchical gating reward function for reservoir measure recommendation tasks. The hierarchical gating reward function includes at least a conflict diagnosis reward, a measure recommendation reward, a consistency constraint reward, and a hard constraint gating, wherein the measure recommendation reward is constrained by the gating threshold of the conflict diagnosis reward. S4. Generate multiple candidate responses for the same input sample, calculate the reward score of each candidate response using a hierarchical gating reward function, and construct a preference pair consisting of the preferred response and the unpreferred response based on the reward score, reward gap and hard constraint gating results. S5. Based on preference pairs, a reward-guided preference optimization algorithm is used to train the reference model to obtain a reservoir measures recommendation model; S6. Input the production dynamic data of the oil well or well group to be analyzed into the reservoir measure recommendation model, and output a structured result that includes data analysis, production contradiction diagnosis, measure recommendation and recommendation basis.

2. The method for recommending reservoir measures based on hierarchical reward preference optimization in a large model, as described in claim 1, is characterized in that... The production dynamic data in step S1 includes one or more of the following: daily oil production, daily fluid production, daily water production, water cut, water-oil ratio, bottom hole pressure, oil pressure, casing pressure, water injection volume, and injection-production response parameters. Key production parameters include one or more of the following: water cut, water cut trend, water-oil ratio, daily oil production trend, daily liquid production trend, pressure trend, injection-production response characteristics, production decline rate, and water cut surge magnitude. Production contradiction labels include one or more of the following: normal production, advantageous channel or ineffective circulation, ineffective water injection or insufficient formation energy, water injection surge or flooding, well blockage or contamination, insufficient fluid supply, and excessively rapid decline in production capacity. The action type label includes one or more of the following: shutting in the well or suspending production, fracturing, perforation repair, acidizing to remove blockages, profile control, water shut-off, overhaul or maintenance operations, fluid extraction or pump replacement, maintaining the status quo or enhanced monitoring.

3. The method for recommending reservoir measures based on hierarchical reward preference optimization in a large model, as described in claim 1, is characterized in that... The structured results in step S6 include: The data analysis section is used to list key production parameters, parameter change trends, and abnormal characteristics; The contradiction diagnosis section is used to identify the main types of production contradictions in the current oil well or well group and the basis for their diagnosis. The recommended measures section provides the types of engineering measures that match the main production contradictions, along with the reasons for the recommendations, applicable conditions, and risk warnings.

4. The method for recommending reservoir measures based on hierarchical reward preference optimization in a large model, as described in claim 1, is characterized in that... In step S3, the total reward of the hierarchical gating reward function is obtained by weighting the contradiction diagnosis reward and the measure recommendation reward through hard constraint gating, and then superimposing the consistency constraint reward. The measure recommendation reward is constrained by the contradiction diagnosis gating. The contradiction diagnosis gating opens the measure recommendation reward when the contradiction diagnosis reward is greater than the preset threshold, and closes the measure recommendation reward otherwise.

5. The method for recommending reservoir measures based on hierarchical reward preference optimization in a large model, as described in claim 4, is characterized in that... The contradiction diagnosis reward is obtained by multiplying the contradiction type identification score by the parameter-contradiction causal consistency factor; The contradiction type identification score is used to evaluate the consistency between the production contradiction type output by the model and the standard production contradiction label; the parameter-contradiction causality consistency factor is used to evaluate whether the key production parameters referenced or analyzed by the model support its output production contradiction diagnosis. The parameter-contradictory causal consistency factor is obtained by weighted summation of three factors: coverage of key parameters, accuracy of key trend judgment, and rationality of key values ​​or intervals. Key parameter coverage represents the ratio between the number of key parameters mentioned in the model and the number of key parameters required to resolve the contradiction. The accuracy of key trend judgment indicates the proportion of correct trend judgments for key parameters; the rationality of key values ​​or intervals indicates whether the key values ​​or intervals referenced by the model are reasonable.

6. The method for recommending reservoir measures based on hierarchical reward preference optimization in a large model, as described in claim 4, is characterized in that... The reward for recommended measures is obtained by multiplying the contradiction diagnosis gating function, the contradiction-measure causal consistency factor, and the measure type matching score. The contradiction diagnosis gating function takes a value of 1 when the contradiction diagnosis reward is greater than a preset threshold and a value of 0 when the contradiction diagnosis reward is less than or equal to the preset threshold, so as to constrain the reward for recommended measures by the contradiction diagnosis result. The contradiction-measure causal consistency factor is used to evaluate whether the recommended measures are applicable to the diagnosed production contradiction. The measure type matching score is used to evaluate the degree of consistency between the recommended measure type and the standard measure type label.

7. The method for recommending reservoir measures based on hierarchical reward preference optimization in a large model according to claim 4, characterized in that, The consistency constraint reward is a negative penalty, determined by the sum of the products of the penalty weights of various consistency violation rules and the number of occurrences. The consistency violation rules include one or more of the following: parameter-contradiction inconsistency, contradiction-measure inconsistency, data-conclusion inconsistency, trend judgment error, measure application condition error, and engineering feasibility error. The hard constraint gating determines its value based on whether the consistency constraint reward is lower than the hard constraint threshold: when the consistency constraint reward is lower than the hard constraint threshold, the value is 0, which sets the contradiction diagnosis reward and measure recommendation reward to zero; when the consistency constraint reward is not lower than the hard constraint threshold, the value is 1, which retains the contradiction diagnosis reward and measure recommendation reward.

8. The method for recommending reservoir measures based on hierarchical reward preference optimization in a large model according to claim 1, characterized in that, The method of constructing preference pairs in step S4 includes: generating no less than two candidate responses for the same input sample, calculating the total reward for each candidate response, determining the candidate response with the highest reward score or that meets the preset high score condition as the preferred response, determining the candidate response with the lowest reward score, that is lower than the preset quantile or that triggers hard constraint gating as the undesired response, and calculating the reward gap between the preferred response and the undesired response. Candidate responses include one or more of the following: responses generated by sampling from the reference model, responses rewritten from historical expert answers, responses generated by rule perturbation, and responses generated by different decoding parameters.

9. The method for recommending reservoir measures based on hierarchical reward preference optimization in a large model according to claim 1, characterized in that, In step S5, the reward-guided preference optimization algorithm adopts multi-reference preference optimization. During training, a fixed reference model and a dynamic reference model are set. The fixed reference model keeps the initial reference policy unchanged, while the dynamic reference model is updated during the training process. Training checkpoints are determined based on a combination of dynamic reference preference intervals, validation loss, and preference discrimination accuracy. A course-based reward scheduling mechanism is set up during training to gradually increase the intensity of reward guidance as the training progresses.

10. A large-scale reservoir measure recommendation system based on hierarchical reward preference optimization, characterized in that, include: The data processing module is used to acquire dynamic production data of oil wells or well groups and generate key production parameters, parameter change trends, production conflict labels, and measure type labels. The supervised fine-tuning module is used to construct structured supervised fine-tuning samples and train the basic large language model to obtain the reference model; The reward calculation module is used to calculate the contradiction diagnosis reward, measure recommendation reward, consistency constraint reward, hard constraint gating result and total reward of the candidate response based on the hierarchical gating reward function; The preference pair construction module is used to construct preferred and unpreferred response preference pairs based on the total reward of candidate responses, the reward gap, and the hard constraint gating results. The preference optimization training module is used to perform reward-guided preference optimization training on the reference model based on the preference pairs. The checkpoint selection module is used to select training checkpoints based on the dynamic reference preference interval, validation loss, and preference discrimination accuracy. The inference output module is used to output data analysis, production conflict diagnosis, and recommended measures results based on the production dynamic data of the oil well or well group to be analyzed.