Well drilling risk processing method and device based on large language model
By employing a drilling risk processing method based on a large language model, and utilizing drilling state parameters and an external knowledge base for similarity matching, the real-time performance and cross-domain knowledge fusion issues in existing drilling risk detection technologies are resolved, enabling accurate detection and rapid processing of drilling risks.
Patent Information
- Application Number
- CN202510839780.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-11-11
AI Technical Summary
Existing drilling risk detection methods rely on human experience or fixed rules, resulting in poor real-time performance and weak cross-domain knowledge integration capabilities, leading to large errors and strong delays in drilling risk detection and early warning.
A drilling risk management method based on a large language model is adopted. By acquiring drilling status parameters, drilling status text is constructed. Similarity matching is performed using a pre-set external knowledge base and a large language model to generate prompt words, determine drilling risks, and provide explanations of risk causes and handling measures.
It enables accurate and efficient detection of drilling risks, as well as rapid and detailed explanation of risk causes and mitigation measures, ensuring drilling safety.
Smart Images

Figure CN120931065A_ABST
Abstract
Description
Technical Field
[0001] This manual belongs to the field of drilling development technology, and in particular relates to drilling risk management methods and devices based on large language models. Background Technology
[0002] As shallow oil and gas resources are gradually depleted, oil and gas exploration and development is now expanding from conventional shallow oil and gas resources to deeper formations and deep-sea areas. However, drilling in deep formations and deep-sea areas involves complex geological conditions and highly uncertain formation pressures, which can easily lead to drilling risks such as overflows and lost circulation, affecting drilling safety.
[0003] Existing methods mostly rely on human experience or numerical models based on fixed rules for drilling risk detection. However, these methods suffer from drawbacks such as poor real-time performance, weak cross-domain knowledge integration capabilities, and insufficient interpretability. This leads to errors and a certain degree of lag when used for drilling risk detection and early warning in specific applications.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This manual provides a drilling risk management method and apparatus based on a large language model, which can effectively utilize the large language model to accurately and efficiently detect the existence of drilling risks, and can also quickly and precisely determine the corresponding risk cause explanations and / or risk management measures, and can be well adapted to drilling operation scenarios.
[0006] This specification provides a drilling risk management method based on a large language model, including:
[0007] Obtain the current drilling status parameters of the target well;
[0008] Based on the current drilling status parameters of the target well, determine the matching status perception data for the current drilling status of the target well; wherein, the status perception data includes at least: geological attribute information, wellbore flow information, and change information of key engineering parameters;
[0009] Based on the state-aware data, the current drilling status text of the target well is constructed;
[0010] Based on the current drilling status text of the target well and a preset external knowledge base, target knowledge data that matches the current drilling status of the target well is determined through similarity matching. The preset external knowledge base is constructed according to preset construction rules, using drilling domain knowledge information and knowledge information of the target area to which the target well belongs, and is associated with the target well. The preset external knowledge base includes multiple knowledge data.
[0011] Based on the target knowledge data and the current drilling status text of the target well, generate target prompt words for the current drilling status of the target well;
[0012] By using a pre-defined large language model and processing target prompts, it can determine whether there is a drilling risk in the target well, and if so, determine the corresponding risk cause explanation and / or risk management measures.
[0013] In one embodiment, the geological attribute information includes at least: formation lithology, formation pressure, and rock mechanical parameters; the wellbore flow information includes at least: drilling fluid density, viscosity, rheological properties, and solid content; and the key engineering parameters include at least: hook height, drilling pressure, torque, rotational speed, standpipe pressure, pump flush, casing pressure, total tank volume, outlet flow rate, inlet flow rate, outlet density, outlet density, and gas measurement value.
[0014] In one embodiment, based on the current drilling status text of the target well and a preset external knowledge base, target knowledge data matching the current drilling status of the target well is determined through similarity matching, including:
[0015] Convert the current drilling status text of the target well into the corresponding current drilling status vector of the target well;
[0016] Using the current drilling status vector of the target well, a similarity match is performed on a preset external knowledge base to determine a preset number of candidate knowledge data with the highest matching degree.
[0017] Based on the candidate knowledge data, the target knowledge data is determined.
[0018] In one embodiment, the target cue words include at least: a first cue word related to drilling risk detection;
[0019] Accordingly, the method further includes:
[0020] The current drilling status text of the target well is transformed into a semantic vector using a pre-defined semantic processing model to obtain the corresponding first intermediate vector.
[0021] Based on the first retrieval enhancement rule, the first key feature in the current drilling status text of the target well and the weight coefficient of the first key feature are determined.
[0022] Based on the first key feature and its weight coefficient, the first intermediate vector is adjusted to obtain the first drilling state vector for drilling risk detection of the current target well.
[0023] Using the first drilling state vector, the first knowledge data is determined by similarity matching with a preset external knowledge base;
[0024] Based on the first knowledge data and the current drilling status text of the target well, a first prompt word is constructed for the target well; wherein, the first prompt word is used to input into a preset large language model for reasoning to determine whether there is a drilling risk in the target well.
[0025] In one embodiment, when it is determined that there is a drilling risk in the target well, the target cue word further includes: a second cue word related to the explanation of the cause of the drilling risk and a third cue word related to the risk management measures for the drilling risk.
[0026] In one embodiment, after obtaining the corresponding first intermediate vector, the method further includes:
[0027] Based on the second retrieval enhancement rule, the second key feature in the current drilling status text of the target well and the weight coefficient of the second key feature are determined.
[0028] Based on the second key feature and its weight coefficient, the second intermediate vector is adjusted to obtain the second drilling state vector for explaining the risk causes of the current target well.
[0029] Using the second drilling state vector, the second knowledge data is determined by similarity matching with a preset external knowledge base;
[0030] Based on the second knowledge data and the current drilling status text of the target well, construct a second prompt word for the target well;
[0031] By using a pre-defined large language model and processing the second cue word, the risk causes of drilling risks for the target well are determined.
[0032] In one embodiment, after determining the cause of the drilling risk for the target well, the method further includes:
[0033] Based on the third retrieval enhancement rule, the current drilling status text of the target well and the explanation of the risk reasons are fused to obtain the corresponding joint text; and the joint text is processed by semantic vector transformation using a preset semantic processing model to obtain the corresponding third intermediate vector.
[0034] Based on the third retrieval enhancement rule, the third key feature in the joint text and the weight coefficient of the third key feature are determined;
[0035] Based on the third key feature and its weight coefficient, the third intermediate vector is adjusted to obtain the third drilling state vector for risk management measures for the current target well.
[0036] Using the third drilling state vector, the third knowledge data is determined by similarity matching with a pre-set external knowledge base;
[0037] Based on the aforementioned third knowledge data, explanations of risk causes, and the current drilling status text of the target well, a third prompt word is constructed for the target well;
[0038] By using a pre-defined large language model and processing third-party prompts, the system determines the appropriate measures to address drilling risks in the target well.
[0039] In one embodiment, the method further includes:
[0040] We will acquire and construct a basic general knowledge base for the field of drilling engineering based on classic drilling engineering theories, industry standards, and typical case documents.
[0041] Obtain geological background data of the target area where the target well is located, as well as drilling records of adjacent wells already drilled in the target area;
[0042] Based on the geological background data of the target area and the drilling records of adjacent wells of the target well in the target area, the geological-engineering data information of the target area is extracted.
[0043] By integrating geological and engineering data information of the target area into the basic general knowledge base, a preset external knowledge base associated with the target well is obtained.
[0044] This specification also provides a drilling risk management device based on a large language model, including:
[0045] The acquisition module is used to obtain the current drilling status parameters of the target well;
[0046] The first determining module is used to determine matching state perception data about the current drilling state of the target well based on the current drilling state parameters of the target well; wherein, the state perception data includes at least: geological attribute information, wellbore flow information, and change information of key engineering parameters;
[0047] The construction module is used to construct the current drilling status text of the target well based on the status perception data;
[0048] The second determining module is used to determine target knowledge data that matches the current drilling status of the target well based on the current drilling status text of the target well and a preset external knowledge base through similarity matching; wherein, the preset external knowledge base is a knowledge base constructed according to preset construction rules, using drilling domain knowledge information and knowledge information of the target area to which the target well belongs, and is associated with the target well; the preset external knowledge base includes multiple knowledge data.
[0049] The generation module is used to generate target prompt words for the current drilling status of the target well based on the target knowledge data and the current drilling status text of the target well.
[0050] The third determination module is used to determine whether there is a drilling risk in the target well by processing target prompt words using a preset large language model, and if it is determined that there is a drilling risk in the target well, to determine the corresponding risk cause explanation and / or risk treatment measures.
[0051] This specification also provides a computer program product comprising a computer program that, when executed by a processor, implements the relevant steps of the drilling risk handling method based on a large language model.
[0052] Based on the drilling risk management method and apparatus based on a large language model provided in this specification, before specific implementation, a preset external knowledge base associated with the target well can be constructed by combining drilling domain knowledge information and the knowledge information of the target area to which the target well belongs, according to preset construction rules; and a preset large language model adapted to the drilling construction scenario can be trained. During specific implementation, the current drilling status parameters of the target well can be acquired and used to determine matching status perception data; then, based on this status perception data, the current drilling status text of the target well can be constructed; based on the current drilling status text of the target well and the preset external knowledge base, target knowledge data matching the current drilling status of the target well can be determined through similarity matching; based on the target knowledge data and the current drilling status text of the target well, target prompt words for the current drilling status of the target well can be generated; using the preset large language model to process the target prompt words, it can be determined whether there is a drilling risk in the target well, and if it is determined that there is a drilling risk, the corresponding risk cause explanation and / or risk management measures can be further determined. This allows for the effective use of large language models, fully integrating the characteristics of drilling scenarios, to accurately and efficiently detect whether there are drilling risks in the current target well. Furthermore, it enables the rapid and precise identification of risk causes and / or risk mitigation measures with high reference value, thereby eliminating the corresponding drilling risks in a timely manner and ensuring the safety of drilling operations in the target well. Attached Figure Description
[0053] To more clearly illustrate the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. The drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1This is a flowchart illustrating a drilling risk management method based on a large language model, provided in one embodiment of this specification.
[0055] Figure 2 This is a schematic diagram illustrating an embodiment of the drilling risk handling method based on a large language model provided in this specification, applied in a scenario example.
[0056] Figure 3 This is a schematic diagram illustrating an embodiment of the drilling risk handling method based on a large language model provided in this specification, applied in a scenario example.
[0057] Figure 4 This is a schematic diagram illustrating an embodiment of the drilling risk handling method based on a large language model provided in this specification, applied in a scenario example.
[0058] Figure 5 This is a schematic diagram illustrating an embodiment of the drilling risk handling method based on a large language model provided in this specification, applied in a scenario example.
[0059] Figure 6 This is a schematic diagram illustrating an embodiment of the drilling risk handling method based on a large language model provided in this specification, applied in a scenario example.
[0060] Figure 7 This is a schematic diagram of the structural composition of a computer device provided in one embodiment of this specification;
[0061] Figure 8 This is a schematic diagram of the structural composition of a drilling risk processing device based on a large language model, provided in one embodiment of this specification.
[0062] Figure 9 This is a schematic diagram illustrating one embodiment of the drilling risk management method based on a large language model provided in this specification, applied in a scenario example. Detailed Implementation
[0063] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0064] It should be noted that the information and data related to users involved in the embodiments of this specification are all information and data authorized by the user or fully authorized by the relevant parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with relevant laws, regulations, and standards, and necessary confidentiality measures have been taken. They do not violate public order and good morals, and corresponding operation entry points are provided for users or relevant parties to choose to authorize or refuse.
[0065] It should also be noted that in the embodiments of this specification, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, they do not mean that the applicant has used or necessarily used the solution.
[0066] See Figure 1 As shown in the embodiments of this specification, a drilling risk handling method based on a large language model is provided. In specific implementation, this method may include the following:
[0067] S101: Obtain the current drilling status parameters of the target well;
[0068] S102: Based on the current drilling status parameters of the target well, determine the matching status perception data for the current drilling status of the target well; wherein, the status perception data includes at least: geological attribute information, wellbore flow information, and change information of key engineering parameters;
[0069] S103: Based on the state-aware data, construct the current drilling status text of the target well;
[0070] S104: Based on the current drilling status text of the target well and a preset external knowledge base, target knowledge data that matches the current drilling status of the target well is determined through similarity matching; wherein, the preset external knowledge base is constructed according to preset construction rules, using drilling domain knowledge information and knowledge information of the target area to which the target well belongs, and is associated with the target well; the preset external knowledge base includes multiple knowledge data;
[0071] S105: Based on the target knowledge data and the current drilling status text of the target well, generate target prompt words for the current drilling status of the target well;
[0072] S106: Using a pre-defined large language model, by processing target prompt words, determine whether there is a drilling risk in the target well, and if it is determined that there is a drilling risk in the target well, determine the corresponding risk cause explanation and / or risk treatment measures.
[0073] Specifically, the target wells mentioned above can be understood as oil and gas wells that are currently being drilled or are awaiting drilling. The drilling risks mentioned include: overflow risk, or lost circulation risk, etc.
[0074] The aforementioned pre-defined external knowledge base can be understood as a knowledge base associated with the target well, constructed in advance according to pre-defined construction rules, combining drilling domain knowledge information with good universality and generalization, and knowledge information of the target area to which the target well belongs with strong adaptability. The specific construction method of the pre-defined external knowledge base will be explained separately later.
[0075] The preset external knowledge base can contain multiple types and contents of drilling knowledge data. Furthermore, the preset external knowledge base can also contain rule-based data, such as reasoning and analysis rules related to drilling risk detection and handling, which are accumulated in advance through machine learning based on the aforementioned drilling knowledge data.
[0076] The aforementioned pre-set large language model can be understood as an algorithm model that is pre-trained based on a large language model and using sample data from drilling construction scenarios, and is capable of performing reasoning analysis related to drilling risks based on input prompt words.
[0077] The aforementioned Large Language Model (LLM) can be understood as a deep learning model trained on a large amount of text data. Based on this model, natural language text can be generated or the meaning of language text can be understood. It can also provide in-depth knowledge about various topics and language production by training on a huge dataset.
[0078] Based on the above embodiments, the current drilling status parameters of the target well can be acquired and used to determine effective status awareness data for drilling risk detection and handling. Then, using this status awareness data and a pre-defined external knowledge base, target knowledge data for drilling risk detection and handling can be determined. This target knowledge data can then be used to generate target prompts that are adaptable to the drilling operation scenario and can effectively guide a pre-defined large language model in reasoning analysis related to drilling risks. The pre-defined large language model can then process these target prompts to perform reasoning analysis and determine whether drilling risks exist in the target well. If drilling risks are determined to exist, corresponding risk cause explanations and / or risk handling measures can be further determined. This allows for the effective use of a large language model, combined with the characteristics of the drilling operation scenario, to accurately and efficiently detect whether drilling risks exist in the target well. Furthermore, it enables the rapid and precise determination of risk cause explanations and / or risk handling measures with high reference value, thereby eliminating corresponding drilling risks in a timely manner and ensuring the safety of drilling operations in the target well.
[0079] In some embodiments, during specific implementation, all current drilling status parameters of the target well can be obtained based on the target well's drilling records, drilling design plan, and seismic detection data. These drilling status parameters include: geological attribute information describing the geological properties of the formation where the target well is located; wellbore flow information describing the well's properties; and key engineering parameters describing the engineering properties during drilling operations for the target well, etc.
[0080] The geological attribute information may include at least: formation lithology, formation pressure, rock mechanics parameters, etc.; the wellbore flow information may include at least: drilling fluid density, viscosity, rheological properties, solid content, etc.; the key engineering parameters may include at least: hook height, drilling pressure, torque, rotational speed, standpipe pressure, pump flush, casing pressure, total tank volume, outlet flow rate, inlet flow rate, outlet density, outlet density, gas measurement value, etc.
[0081] Of course, it should be noted that the geological attribute information, wellbore flow information, and key engineering parameters listed above are only illustrative. In actual implementation, other suitable parameter information may be included depending on the specific circumstances and processing requirements. This manual does not impose any limitations on this.
[0082] In practice, the corresponding geological attribute information can be extracted by acquiring and based on data such as the wellbore structure, wellbore trajectory, target drilling layer, designed well diameter, top and bottom depth of the layer, stratigraphic division, and predicted lithology of the target well.
[0083] After obtaining the current drilling status parameters of the target well, further analysis can be performed based on these parameters to determine the most suitable and effective data for the current drilling risk management of the target well, which can then be used as status awareness data.
[0084] In some embodiments, considering that the types of drilling risks, the causes of drilling risks, and the effects of risk factors that directly or indirectly cause drilling risks may differ across different drilling stages throughout the complete drilling cycle, in order to more accurately detect and manage subsequent drilling risks (including drilling risk detection, risk cause explanation, and risk management measure recommendations), avoid interference from non-primary risk factors, and improve overall processing efficiency, different drilling stages can be distinguished, and targeted and effective drilling status parameters can be obtained for different drilling stages as status perception data.
[0085] The above-mentioned determination of matching status-aware data regarding the current drilling status of the target well based on the current drilling status parameters of the target well may include the following in specific implementation:
[0086] S1: Determine the current drilling stage of the target well based on the current drilling status parameters of the target well; wherein, the drilling stage includes at least one of the following: drilling, circulation, reaming, pump start / stop, casing running, tripping, shut-in, and lost circulation control;
[0087] S2: From the preset rule set, determine the preset processing rule that matches the current drilling stage of the target well, and use it as the target processing rule;
[0088] S3: Based on the target processing rules and the current drilling status parameters of the target well, obtain matching status awareness data about the current drilling status of the target well.
[0089] The preset rule set can store multiple preset processing rules, and each preset processing rule corresponds to at least one drilling stage.
[0090] Before implementation, a large amount of historical data on drilling risk detection and treatment can be collected. Records demonstrating satisfactory detection and treatment results from this historical data can be selected as the first sample. These first sample records are then clustered to obtain multiple data groups. Each data group corresponds to at least one drilling stage and contains drilling parameters with common usage characteristics specific to that stage. Based on these data groups, multiple preset treatment rules corresponding to the drilling stages are generated. Finally, a preset rule set is constructed based on these preset treatment rules.
[0091] In some embodiments, see Figure 2As shown, the above method determines target knowledge data that matches the current drilling status of the target well based on the current drilling status text of the target well and a preset external knowledge base through similarity matching. In specific implementation, this may include the following:
[0092] S1: Convert the current drilling status text of the target well into the corresponding current drilling status vector of the target well;
[0093] S2: Using the current drilling status vector of the target well, perform similarity matching on a preset external knowledge base to determine a preset number of candidate knowledge data with the highest matching degree.
[0094] S3: Determine the target knowledge data based on the candidate knowledge data.
[0095] Specifically, the aforementioned target knowledge data can be understood as text data that can accurately and completely describe the current drilling status of the target well from the perspective of drilling risk detection and management, combined with relevant theories and existing cases.
[0096] Accordingly, the aforementioned target knowledge data can be used to construct effective prompts that guide the pre-defined large language model in reasoning analysis regarding drilling risk detection and handling of the target well.
[0097] In practice, a preset semantic processing model can be used first to perform corresponding semantic vector conversion processing on the current drilling status text of the target well to obtain the required current drilling status vector of the target well.
[0098] Then, similarity matching can be performed using the current drilling state vector of the target well and a preset external knowledge base; wherein, the knowledge data can be stored in the preset external knowledge base in the form of vectors. Specifically, the matching degree between the current drilling state vector of the target well and each knowledge data can be calculated based on vector distance; then, according to the matching degree, each knowledge data is arranged in descending order of matching degree; and a preset number of knowledge data with the highest matching degree (e.g., the top 10 knowledge data with the highest matching degree) are selected from the preset external knowledge base as candidate knowledge data.
[0099] Furthermore, it can be first detected whether the number of candidate knowledge data with a matching degree greater than or equal to a preset first matching degree threshold (e.g., 65%) is greater than or equal to a first reference number. When it is determined that the number is greater than or equal to the first reference number, the candidate knowledge data with a matching degree greater than or equal to the preset first matching degree threshold can be filtered and combined to obtain the corresponding target knowledge data.
[0100] Conversely, when the determined quantity is less than the first reference quantity, candidate knowledge data with a matching degree greater than or equal to a preset second matching degree threshold and less than a preset first matching degree threshold can be selected from the candidate knowledge data as basic knowledge data; then, the basic knowledge data is adjusted and expanded according to the state-aware data to obtain corresponding supplementary knowledge data; finally, the candidate knowledge data with a matching degree greater than or equal to the preset first matching degree threshold is combined with the supplementary knowledge data to obtain the corresponding target knowledge data.
[0101] Specifically, when adjusting and expanding basic knowledge data based on state-aware data, a pre-trained, preset knowledge modification model can be used to process state-aware data and basic knowledge data, and output corresponding supplementary knowledge data. The preset knowledge modification model is a neural network model pre-trained through deep learning, capable of adaptively adjusting and expanding the input basic knowledge data based on the input state-aware data.
[0102] Furthermore, based on the aforementioned target knowledge data, target prompt words can be generated that match the preset large language model and can effectively guide the preset large language model to perform corresponding reasoning analysis.
[0103] In some embodiments, the target prompts may include at least: a first prompt related to drilling risk detection; further, they may include: a second prompt related to the explanation of the causes of drilling risks, and a third prompt related to risk management measures for drilling risks. Different types of prompts correspond to different content inference analyses. By distinguishing and using different types of prompts, the pre-defined large language model can be guided to perform inference analyses of different content in a more refined and accurate manner, resulting in precise and diverse outcomes.
[0104] Furthermore, in order to better guide the pre-set large language model to perform more refined and accurate reasoning analysis of different content, it is also possible to consider distinguishing and using different types of retrieval enhancement rules to process target knowledge data, so as to obtain more targeted and effective prompt words.
[0105] Correspondingly, based on the model feedback data collected during the training of the pre-set large language model, data statistics and machine learning can be used to analyze the model's operational characteristics when it performs inference analysis on different content, such as drilling risk detection, risk cause explanation, and treatment measure recommendation. Then, based on these operational characteristics, the linguistic features of effective prompts guiding the model in inference analysis on different content can be determined. Furthermore, based on these linguistic features, retrieval enhancement rules can be constructed to guide the model in inference analysis on different content. For example, there could be a first retrieval enhancement rule focusing on guiding the model in drilling risk detection, a second retrieval enhancement rule focusing on guiding the model in explaining risk causes, and a third retrieval enhancement rule focusing on guiding the model in recommending treatment measures, and so on.
[0106] In some embodiments, the target cue words include at least: a first cue word related to drilling risk detection;
[0107] Accordingly, see Figure 3 As shown, in specific implementations, the method may also include the following:
[0108] S1: Use a preset semantic processing model to perform semantic vector transformation on the current drilling status text of the target well to obtain the corresponding first intermediate vector;
[0109] S2: Based on the first retrieval enhancement rule, determine the first key feature in the current drilling status text of the target well, as well as the weight coefficient of the first key feature;
[0110] S3: Based on the first key feature and the weight coefficient of the first key feature, adjust the first intermediate vector to obtain the first drilling state vector for drilling risk detection of the current target well.
[0111] S4: Using the first drilling state vector, the first knowledge data is determined by similarity matching with a preset external knowledge base;
[0112] S5: Based on the first knowledge data and the current drilling status text of the target well, construct a first prompt word for the target well; wherein, the first prompt word is used to input into a preset large language model for reasoning to determine whether there is a drilling risk in the target well.
[0113] Specifically, the preset semantic processing model can be a semantic text embedding model that includes at least a bidirectional Transformer encoder. Accordingly, when processing the current drilling status text of the target well based on the preset semantic processing model, the bidirectional Transformer encoder can be used to simultaneously consider the contextual information on both sides of each word in the text, and then the word can be converted into a corresponding word vector; then, the word vectors in the text can be combined and concatenated in sequence to obtain the corresponding first intermediate vector.
[0114] Specifically, the aforementioned preset semantic processing model can also be a model based on the BERT (Bidirectional Encoder Representations from Transformers) structure, or a model based on the SBERT (Sentence-BERT) structure, or a model based on the embedding API structure, etc.
[0115] In practice, according to the first retrieval enhancement rule, keywords reflecting significant changes in the status-aware data (e.g., "mutation," "abnormality," "rise," "fall") can be identified from the current drilling status text of the target well as the first key features. Then, based on the severity of the data changes indicated by these first key features, the weight coefficient of each first key feature is determined. Specifically, according to the first retrieval enhancement rule, a larger weight coefficient (e.g., 2.1) is assigned to the first key feature indicating a greater degree of data change (e.g., "mutation"). Conversely, a smaller weight coefficient (e.g., 1.1) is assigned to the first key feature indicating a less severe degree of data change (e.g., "rise").
[0116] Then, the word vector corresponding to the first key feature can be determined in the first intermediate vector; and the word vector is weighted according to the weight coefficient of the corresponding first key feature to adjust the first intermediate vector and obtain the first drilling state vector suitable for drilling risk detection.
[0117] Then, based on the aforementioned first drilling state vector, similarity matching can be performed on a preset external knowledge base to effectively retrieve first knowledge data that is more suitable for guiding the preset large language model to perform drilling risk detection. This first knowledge data, combined with the current drilling state text of the target well, generates first prompt words suitable for guiding the model to perform drilling risk detection.
[0118] Accordingly, the aforementioned first prompt can be input into a preset large language model, and the preset large language model can be run. Based on the first prompt, and combined with relevant knowledge from a preset external knowledge base—based on generalizable classic cases and historical cases closely related to the target well—the preset large language model can make targeted inferences about whether the target well currently faces drilling risks, and output the final inference result. Based on this inference result, it can be determined whether the target well currently faces drilling risks.
[0119] In some embodiments, after determining that there is a drilling risk in the target well, other prompt words can be generated, and these prompt words can be used to guide a pre-set large model to conduct further reasoning analysis on the explanation of the risk causes and the recommendation of treatment measures.
[0120] Specifically, when it is determined that there is a drilling risk in the target well, the target prompt words may also include: a second prompt word related to the explanation of the cause of the drilling risk, a third prompt word related to the risk management measures for the drilling risk, etc.
[0121] In some embodiments, after obtaining the corresponding first intermediate vector, refer to Figure 4 As shown, in specific implementations, the method may also include the following:
[0122] S1: Based on the second retrieval enhancement rule, determine the second key feature in the current drilling status text of the target well, as well as the weight coefficient of the second key feature;
[0123] S2: Based on the second key feature and the weight coefficient of the second key feature, adjust the second intermediate vector to obtain the second drilling state vector for the risk cause explanation of the current target well;
[0124] S3: Using the second drilling state vector, the second knowledge data is determined by similarity matching with a preset external knowledge base;
[0125] S4: Based on the second knowledge data and the current drilling status text of the target well, construct a second prompt word for the target well;
[0126] S5: Using a pre-defined large language model, the risk causes of drilling risks for the target well are determined by processing the second cue word.
[0127] In practice, based on the second retrieval enhancement rule, keywords that reflect risk factors such as environmental characteristics, geological conditions, and operational behavior can be identified from the current drilling status text of the target well as the second key features. At the same time, based on the interaction between the above risk factors and drilling risks, the weight coefficients of the corresponding second key features are determined.
[0128] The interaction between the aforementioned risk factors and drilling risks can be determined in advance by constructing and using a correlation matrix to conduct correlation analysis based on a large amount of historical drilling risk data.
[0129] Then, the word vector corresponding to the second key feature can be determined in the second intermediate vector; and the word vector is weighted according to the weight coefficient of the corresponding second key feature to adjust the second intermediate vector and obtain the second drilling state vector suitable for drilling risk detection.
[0130] In some embodiments, after determining the risk cause explanation for the drilling risk to the target well, refer to Figure 5 As shown, in specific implementations, the method may also include the following:
[0131] S1: Based on the third retrieval enhancement rule, the current drilling status text of the target well and the explanation of the risk reasons are fused to obtain the corresponding joint text; and the joint text is processed by semantic vector transformation using a preset semantic processing model to obtain the corresponding third intermediate vector.
[0132] S2: Based on the third retrieval enhancement rule, determine the third key feature in the joint text and the weight coefficient of the third key feature;
[0133] S3: Based on the third key feature and its weight coefficient, adjust the third intermediate vector to obtain the third drilling state vector for risk management measures for the current target well.
[0134] S4: Using the third drilling state vector, the third knowledge data is determined by similarity matching with a preset external knowledge base;
[0135] S5: Based on the aforementioned third knowledge data, risk cause explanation, and the current drilling status text of the target well, construct a third prompt word for the target well;
[0136] S6: Using a pre-defined large language model, by processing the third cue word, determine the handling measures for drilling risks of the target well.
[0137] In practice, in order to enable the pre-set large language model to fully and comprehensively understand the current drilling status of the target well, the current drilling status text and the risk cause explanation can be merged to obtain a fused text that includes both the objective status of the target well and the further cause mechanism analysis based on the objective status.
[0138] Then, based on the third retrieval enhancement rule, keywords with a semantic similarity higher than a preset semantic similarity threshold can be identified from the joint text as the third key feature. Specifically, the template keywords can include one or more of the following: handling measures, operational suggestions, successful recovery, etc. Furthermore, the weight coefficient of the third key feature can be determined based on its semantic similarity to the template keywords. Specifically, the higher the semantic similarity, the larger the corresponding weight coefficient; conversely, the lower the semantic similarity, the smaller the corresponding weight coefficient.
[0139] Then, the word vector corresponding to the third key feature can be determined in the third intermediate vector; and the word vector is weighted according to the weight coefficient of the corresponding third key feature to adjust the third intermediate vector and obtain the third drilling state vector suitable for drilling risk detection.
[0140] In some embodiments, see Figure 6 As shown, in specific implementations, the method may also include the following:
[0141] S1: Acquire and construct a basic general knowledge base for the field of drilling engineering based on classic drilling engineering theories, industry standards, and typical case documents;
[0142] S2: Obtain geological background data of the target area where the target well is located, as well as drilling records of adjacent wells already drilled in the target area;
[0143] S3: Based on the geological background data of the target area and the drilling records of adjacent wells of the target well in the target area, extract the geological-engineering data information of the target area;
[0144] S4: Integrate the geological and engineering data information of the target area into the basic general knowledge base to obtain a preset external knowledge base associated with the target well.
[0145] Specifically, the aforementioned classic theories, industry standards, and typical case documents for drilling engineering may include one or more of the following: drilling process flow descriptions, drill bit and drilling fluid usage guidelines, typical formation response descriptions, equipment operation manuals, complex situation handling solutions, construction case reports, and industry standards and specifications, etc.
[0146] In practice, relevant textual data, including classic theories, industry standards, and typical case documents of drilling engineering, based on different types and data structures, can be collected from various sources. This data is then preprocessed, including data cleaning and format standardization, to remove redundant and low-quality content, resulting in preprocessed first textual data. Next, the preprocessed first textual data is segmented according to semantic blocks to obtain multiple text fragments. Feature extraction is then performed on these multiple text fragments to obtain multiple knowledge data. Then, using a pre-defined semantic processing model, semantic vector transformation is performed on these multiple knowledge data to obtain knowledge vectors corresponding to the multiple basic knowledge data. Based on these knowledge vectors, the semantic distance between vectors is calculated and used to identify and merge similar or identical basic knowledge data. Simultaneously, effective information statistics are conducted on the aforementioned basic knowledge data. Based on the statistical results and statistical dimensions, drilling-related knowledge data with good universality and generalization is selected from the knowledge data to construct a basic general knowledge base for the drilling engineering field.
[0147] Specifically, the geological background data of the target area may include one or more of the following: the geological tectonic background, stratigraphic division, fault development, and formation pressure prediction of the target area.
[0148] The drilling records of adjacent drilled wells in the target area may specifically include one or more of the following data: well structure, well inclination, drill string assembly, description of complex situations, logging reports, and construction logs of the adjacent drilled wells.
[0149] In practice, the geological background data of the target area and the drilling records of adjacent wells of the target well in the target area can be preprocessed by data cleaning and format unification to obtain preprocessed second text data; then, the geological-engineering data information of the target area can be extracted by using the preset semantic processing model and the preprocessed second text data.
[0150] Specifically, a pre-defined semantic processing model and pre-processed second text data can be used to obtain multiple related knowledge data associated with the target well, along with corresponding knowledge vectors. Then, based on the knowledge vectors of the related knowledge data, a pre-built basic general knowledge base in the drilling engineering field can be used to learn and acquire knowledge data from the related knowledge data that lacks universality and generalization, but is closely related to the target area where the target well is located, possessing specificity and adaptability to the target area. This knowledge data serves as the geological-engineering data information for the target area.
[0151] Furthermore, based on the basic general knowledge base, a pre-defined external knowledge base associated with the target well can be constructed by integrating the geological and engineering data information of the target area.
[0152] After constructing a pre-defined external knowledge base containing multiple knowledge data related to drilling risk detection and management in the manner described above, further, machine learning can be performed on the knowledge data in the pre-defined external knowledge base to analyze and obtain reasoning and analysis rules related to drilling risk detection and management based on the intrinsic mechanism of action related to drilling risk.
[0153] Among these, the reasoning and analysis rules related to drilling risk detection and management include at least drilling risk detection rules. Specifically, these drilling risk detection rules may include key indicators that need to be considered when detecting and determining the existence of drilling risks, as well as reference thresholds for these key indicators.
[0154] In some embodiments, the aforementioned preset large language model can also be connected to a preset external knowledge base.
[0155] Accordingly, when using a pre-defined large language model to process target prompts and determine whether there is a drilling risk in the target well, the pre-defined large language model can first determine the appropriate reasoning and analysis rules from a pre-defined external knowledge base based on the target prompts. Then, the pre-defined large language model can be used to perform more targeted reasoning and analysis based on the reasoning and analysis rules, using the aforementioned target prompts to accurately determine whether there is a drilling risk in the target well.
[0156] Furthermore, in the process of using the preset large language model and the aforementioned target prompts for more targeted reasoning analysis based on the reasoning analysis rules, the preset large language model can also adaptively adjust the reasoning analysis rules based on the characteristics of the current drilling stage and the data processing experience learned from the drilling risk handling records accumulated for the target well. This makes the reasoning analysis rules adopted by the preset large language model more suitable for the current drilling stage of the target well. Consequently, based on the adaptively adjusted reasoning analysis rules, it is possible to more accurately determine whether there is a drilling risk in the target well. Simultaneously, the adaptively adjusted reasoning analysis rules can also be used to update the reasoning analysis rules stored in the preset external knowledge base.
[0157] In some embodiments, the method may further include the following: obtaining the current cumulative drilling record of the target well at preset time intervals (e.g., every 5 hours); updating a preset external knowledge base based on the current cumulative drilling record of the target well, so that the preset external knowledge base has a stronger correlation with the target well, and then more effective prompts for the target well can be determined based on the preset external knowledge base.
[0158] In some embodiments, the method may further include the following: acquiring and constructing sample data for drilling risk detection and handling based on classic drilling engineering theories, industry standards, and typical case documents; using the sample data to train and learn the initial large language model in multiple rounds; and in each round of training and learning, collecting and adjusting and updating the model based on feedback data through instruction optimization, context learning, etc., to obtain a preset large language model that is adapted to the drilling construction scenario, meets accuracy requirements, and is suitable for drilling risk detection and handling, including drilling risk detection, risk cause explanation, and risk handling measure recommendation.
[0159] As can be seen from the above, the drilling risk handling method based on a large language model provided in this specification can, before implementation, construct a preset external knowledge base associated with the target well by combining drilling domain knowledge information and the knowledge information of the target area to which the target well belongs, according to preset construction rules; and train a preset large language model adapted to the drilling construction scenario. In specific implementation, firstly, based on the current drilling status parameters of the target well, determine the matching state perception data for the current drilling status of the target well; then, based on the state perception data, construct the current drilling status text of the target well; based on the current drilling status text of the target well and the preset external knowledge base, determine the target knowledge data matching the current drilling status of the target well through similarity matching; based on the target knowledge data and the current drilling status text of the target well, generate target prompt words for the current drilling status of the target well; and use the preset large language model to process the target prompt words to determine whether there is a drilling risk in the target well, and if it is determined that there is a drilling risk in the target well, determine the corresponding risk cause explanation and / or risk handling measures. This allows for the effective use of large language models, combined with the characteristics of drilling scenarios, to accurately and efficiently detect whether there are drilling risks in the current target well. Furthermore, it enables the rapid and precise identification of risk causes and / or risk mitigation measures with high reference value, thereby eliminating the corresponding drilling risks in a timely manner and ensuring the safety of drilling operations in the target well.
[0160] This specification provides an embodiment of a computer device, see below. Figure 7 As shown. The computer device includes a network communication port 701, a processor 702, and a memory 703. These structures are connected by internal cables so that they can perform specific data interaction.
[0161] Specifically, the network communication port 701 can be used to obtain the current drilling status parameters of the target well.
[0162] The processor 702 is specifically configured to: determine matching state-aware data about the current drilling status of the target well based on the current drilling status parameters of the target well; wherein the state-aware data includes at least: geological attribute information, wellbore flow information, and change information of key engineering parameters; construct a text representing the current drilling status of the target well based on the state-aware data; determine target knowledge data matching the current drilling status of the target well based on the text representing the current drilling status of the target well and a preset external knowledge base through similarity matching; wherein the preset external knowledge base is constructed according to preset construction rules, using drilling domain knowledge information and knowledge information of the target area to which the target well belongs, and is associated with the target well; the preset external knowledge base includes multiple knowledge data; generate target prompt words for the current drilling status of the target well based on the target knowledge data and the text representing the current drilling status of the target well; and determine whether there is a drilling risk in the target well by processing the target prompt words using a preset large language model, and, if the drilling risk is determined, determine the corresponding risk cause explanation and / or risk treatment measures.
[0163] The memory 703 can be used to store corresponding instruction programs, as well as intermediate data such as status awareness data and current drilling status text.
[0164] Based on the above method, the relevant structural performance of computer equipment can be effectively utilized to improve the data processing speed of electronic equipment and efficiently realize drilling risk data processing based on large language models.
[0165] In this embodiment, the network communication port 701 can be a virtual port bound to different communication protocols, thereby enabling the sending or receiving of different data. For example, the network communication port can be a port responsible for web data communication, a port responsible for FTP data communication, or a port responsible for email data communication. Furthermore, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM or CDMA; it can also be a Wi-Fi chip; or it can be a Bluetooth chip.
[0166] In this embodiment, the processor 702 can be implemented in any suitable manner. For example, the processor can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc. This specification is not limiting.
[0167] In this embodiment, the memory 703 may include multiple layers. In a digital system, anything that can store binary data can be a memory. In an integrated circuit, a circuit with storage function but no physical form is also called a memory, such as RAM, FIFO, etc. In a system, a storage device with a physical form is also called a memory, such as a memory stick, TF card, etc.
[0168] This specification also provides a computer-readable storage medium based on the above-described drilling risk handling method based on a large language model. The computer-readable storage medium stores computer program instructions that, when executed, perform the following: acquire the current drilling status parameters of the target well; determine matching state-aware data regarding the current drilling status of the target well based on the current drilling status parameters; wherein the state-aware data includes at least: geological attribute information, wellbore flow information, and change information of key engineering parameters; construct the current drilling status text of the target well based on the state-aware data; and, based on the current drilling status text of the target well and a preset external knowledge base, through... Similarity matching identifies target knowledge data that matches the current drilling status of the target well. A pre-defined external knowledge base is constructed according to pre-defined rules, combining drilling domain knowledge information and knowledge information from the target region to which the target well belongs, and is associated with the target well. This pre-defined external knowledge base includes multiple knowledge data sets. Based on the target knowledge data and the text indicating the current drilling status of the target well, target prompt words for the current drilling status of the target well are generated. A pre-defined large language model is used to process these target prompt words to determine whether the target well currently faces drilling risks. If drilling risks are determined, corresponding explanations of the risk causes and / or risk mitigation measures are identified.
[0169] In this embodiment, the storage medium includes, but is not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), cache, hard disk drive (HDD), or memory card. The memory can be used to store computer program instructions. The network communication unit can be an interface configured according to standards specified in the communication protocol for network connection communication.
[0170] In this embodiment, the specific functions and effects implemented by the program instructions stored in the computer-readable storage medium can be explained in comparison with other embodiments, and will not be repeated here.
[0171] This specification also provides a computer program product, comprising at least a computer program, which, when executed by a processor, implements the following method steps: acquiring the current drilling status parameters of a target well; determining matching status perception data regarding the current drilling status of the target well based on the current drilling status parameters; wherein the status perception data includes at least: geological attribute information, wellbore flow information, and change information of key engineering parameters; constructing the current drilling status text of the target well based on the status perception data; and determining, through similarity matching, a current drilling status text of the target well that matches the current drilling status of the target well based on the current drilling status text of the target well and a preset external knowledge base. The target knowledge data matches the drilling status; wherein, the preset external knowledge base is constructed according to preset construction rules, using drilling domain knowledge information and knowledge information of the target area to which the target well belongs, and is associated with the target well; the preset external knowledge base includes multiple knowledge data; based on the target knowledge data and the current drilling status text of the target well, target prompt words for the current drilling status of the target well are generated; using a preset large language model to process the target prompt words, it is determined whether there is a drilling risk in the target well, and if it is determined that there is a drilling risk in the target well, the corresponding risk cause explanation and / or risk treatment measures are determined.
[0172] See Figure 8 As shown in the embodiments of this specification, a drilling risk processing device based on a large language model is also provided. This device may specifically include the following structural modules:
[0173] The acquisition module 801 can be used to acquire the current drilling status parameters of the target well.
[0174] The first determining module 802 can be used to determine the matching state perception data about the current drilling state of the target well based on the current drilling state parameters of the target well; wherein, the state perception data includes at least: geological attribute information, wellbore flow information, and change information of key engineering parameters;
[0175] The construction module 803 can be specifically used to construct the current drilling status text of the target well based on the state perception data;
[0176] The second determining module 804 is specifically used to determine target knowledge data that matches the current drilling status of the target well based on the current drilling status text of the target well and a preset external knowledge base through similarity matching; wherein, the preset external knowledge base is a knowledge base constructed according to preset construction rules, using drilling domain knowledge information and knowledge information of the target area to which the target well belongs, and is associated with the target well; the preset external knowledge base includes multiple knowledge data;
[0177] The generation module 805 can be specifically used to generate target prompt words for the current drilling status of the target well based on the target knowledge data and the current drilling status text of the target well.
[0178] The third determination module 806 can be used to determine whether there is a drilling risk in the target well by processing target prompt words using a preset large language model, and to determine the corresponding risk cause explanation and / or risk treatment measures if the target well is determined to have a drilling risk.
[0179] In some embodiments, the geological attribute information may include at least: formation lithology, formation pressure, rock mechanical parameters, etc.; the wellbore flow information may include at least: drilling fluid density, viscosity, rheological properties, solid content, etc.; the key engineering parameters may include at least: hook height, drilling pressure, torque, rotational speed, standpipe pressure, pump flush, casing pressure, total tank volume, outlet flow rate, inlet flow rate, outlet density, outlet density, gas measurement value, etc.
[0180] In some embodiments, when the second determining module 804 is specifically implemented, it can determine the target knowledge data that matches the current drilling status of the target well by similarity matching based on the current drilling status text of the target well and a preset external knowledge base in the following manner: converting the current drilling status text of the target well into a corresponding current drilling status vector of the target well; using the current drilling status vector of the target well to perform similarity matching on the preset external knowledge base to determine a preset number of candidate knowledge data with high matching degree ranking; and determining the target knowledge data based on the candidate knowledge data.
[0181] In some embodiments, the target cue words may include at least: a first cue word related to drilling risk detection;
[0182] Accordingly, in specific implementations, the device can also be used to: perform semantic vector transformation processing on the current drilling status text of the target well using a preset semantic processing model to obtain a corresponding first intermediate vector; determine the first key feature and the weight coefficient of the first key feature in the current drilling status text of the target well according to the first retrieval enhancement rule; adjust the first intermediate vector according to the first key feature and the weight coefficient of the first key feature to obtain a first drilling status vector for drilling risk detection of the current target well; determine the first knowledge data by performing similarity matching on a preset external knowledge base using the first drilling status vector; construct a first prompt word for the target well based on the first knowledge data and the current drilling status text of the target well; wherein, the first prompt word is used to input into a preset large language model for reasoning to determine whether there is a drilling risk in the target well.
[0183] In some embodiments, when it is determined that there is currently a drilling risk in the target well, the target prompt words may further include: a second prompt word related to the explanation of the cause of the drilling risk, a third prompt word related to the risk management measures for the drilling risk, etc.
[0184] In some embodiments, after obtaining the corresponding first intermediate vector, the device may further be used to: determine the second key feature and the weight coefficient of the second key feature in the current drilling status text of the target well according to the second retrieval enhancement rule; adjust the second intermediate vector according to the second key feature and the weight coefficient of the second key feature to obtain a second drilling status vector for the risk cause explanation of the current target well; determine the second knowledge data by performing similarity matching on a preset external knowledge base using the second drilling status vector; construct a second prompt word for the target well based on the second knowledge data and the current drilling status text of the target well; and determine the risk cause explanation of the drilling risk for the target well by processing the second prompt word using a preset large language model.
[0185] In some embodiments, after determining the risk cause explanation for the drilling risk of the target well, the device may further be used to: fuse the current drilling status text of the target well and the risk cause explanation according to a third retrieval enhancement rule to obtain a corresponding joint text; perform semantic vector transformation processing on the joint text using a preset semantic processing model to obtain a corresponding third intermediate vector; determine the third key feature in the joint text and the weight coefficient of the third key feature according to the third retrieval enhancement rule; adjust the third intermediate vector according to the third key feature and the weight coefficient of the third key feature to obtain a third drilling status vector for risk treatment measures for the current target well; determine the third knowledge data by performing similarity matching on a preset external knowledge base using the third drilling status vector; construct a third prompt word for the target well based on the third knowledge data, the risk cause explanation, and the current drilling status text of the target well; and determine the drilling risk treatment measures for the target well by processing the third prompt word using a preset large language model.
[0186] In some embodiments, the device can also be used to: acquire and construct a basic general knowledge base in the field of drilling engineering based on classic drilling engineering theories, industry standards, and typical case documents; acquire geological background data of the target area where the target well is located, as well as drilling records of adjacent wells drilled in the target area; extract geological-engineering data information of the target area based on the geological background data of the target area and the drilling records of adjacent wells drilled in the target area; and integrate the geological-engineering data information of the target area into the basic general knowledge base to obtain a preset external knowledge base associated with the target well.
[0187] It should be noted that the units, devices, or modules described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. For ease of description, the above devices are described by dividing them into various modules according to their functions. Of course, in implementing this specification, the functions of each module can be implemented in one or more software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection between the devices or units shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0188] As can be seen from the above, the drilling risk processing device based on the large language model provided in the embodiments of this specification can effectively utilize the large language model and combine it with the characteristics of the drilling construction scenario to accurately and efficiently detect whether there is a drilling risk in the current target well. It can also quickly and accurately determine the risk cause explanation and / or risk treatment measures with high reference value, so as to eliminate the corresponding drilling risks in a timely manner and ensure the drilling construction safety of the target well.
[0189] In a specific scenario example, the drilling risk handling method based on a large language model provided in this manual can be applied to analyze the causes of drilling risks and recommend appropriate remedial measures. For detailed implementation procedures, please refer to [link / reference needed]. Figure 9 As shown, it includes the following content.
[0190] In this scenario example, traditional risk analysis methods rely primarily on human experience or numerical models based on fixed rules, resulting in bottlenecks such as poor real-time performance, weak cross-domain knowledge integration capabilities, and insufficient interpretability. Because drilling operations require real-time processing of multi-source heterogeneous data, including geological attributes, engineering parameters, and historical risks, traditional methods struggle to achieve rapid correlation and reasoning based on dynamic data. This leads to delayed risk warnings and a lack of causal correlation for decision-making, severely impacting operational safety and efficiency.
[0191] To address the aforementioned issues and their root causes, a further approach is to integrate a large model (large language model) with Retrieval Enhanced Generation (RAG) technology to construct a structured knowledge base (e.g., a pre-defined external knowledge base) encompassing geological, engineering, and adjacent well risk knowledge. This knowledge base, combined with real-time parameter text (e.g., current drilling status text) transformation and similarity matching mechanisms, drives the large model (e.g., a pre-defined large language model) to perform causal reasoning and recommend measures for risks such as spills and wellbore leakage. This not only enhances real-time response capabilities under complex operating conditions but also reduces reliance on human experience through interpretable risk analysis, providing intelligent technical support for the efficient development of deep oil and gas resources.
[0192] In this scenario example, please refer to the following for specific implementation: Figure 9 As shown, the process can include the following steps: First, real-time monitoring of the drilling status, including identifying geological attributes such as formation lithology and calculating formation pressure based on traditional intelligent models, calculating wellbore flow information such as bottom hole pressure, and calculating the changes and rates of change of key engineering parameters (e.g., outlet flow rate, riser pressure, total well volume, inlet flow rate). The calculated indicators are then converted into text descriptions (e.g., the current drilling status text of the target well). Next, an external knowledge base is constructed (e.g., a pre-defined external knowledge base), including knowledge of the oil and gas drilling engineering field and knowledge of neighboring wells within the block (including geological attributes, engineering data, and risk occurrence information). Then, the text descriptions of the real-time drilling status are matched with the knowledge base for similarity. Finally, the reasoning function of a large language model is used to analyze the causes of overflow and lost circulation risks and recommend appropriate remedial measures.
[0193] When constructing a specific external knowledge base, one can integrate classic drilling engineering theories, industry standards, and literature cases to build a standardized and general knowledge base in the field of drilling engineering; or integrate geological and engineering data information of a specific block to form an external knowledge base for that specific block (including geological attributes, engineering data, and risk occurrence information).
[0194] When processing text descriptions, the BERT embedding model can be used to vectorize the real-time drilling status text, converting it into a vector representation. The BM25 algorithm is then used to perform similarity matching between the real-time drilling status text vector and the domain knowledge vector stored in an external knowledge base. By calculating the similarity between the two, the Top K similarity vectors are selected.
[0195] BERT, short for Bidirectional Encoder Representation from Transformers, is a pre-trained language representation model. It employs a novel masked language model (MLM) to generate deep bidirectional language representations. 1) MLM is used to pre-train bidirectional Transformers to generate deep bidirectional language representations. 2) After pre-training, only an additional output layer needs to be added for fine-tuning to achieve state-of-the-art performance on a wide variety of downstream tasks. No task-specific structural modifications to BERT are required during this process.
[0196] BM25 (Best Matching 25) is an algorithm for information retrieval and text mining, widely used in search engines and related fields. BM25 is based on the TF-IDF (Term Frequency-Inverse Document Frequency) concept but improves upon it to consider factors such as document length. Improvements to TF-IDF: BM25 improves TF-IDF calculation by introducing a saturation function and a document length factor for each term in the document. Saturation function: In BM25, a saturation function is introduced to adjust the weight of a term based on its frequency of occurrence (TF). This is to prevent a term from having an excessively high weight due to excessive occurrence in the document. Document length factor: BM25 considers document length by introducing a document length factor, making the impact of document length on weight non-linear. This allows for better adaptation to documents of different lengths.
[0197] When using a large model for reasoning, the constructed prompts can be fed into the large language model for reasoning to obtain drilling risk cause analysis and treatment recommendations based on RAG technology.
[0198] In this scenario example, the construction of an external knowledge base may also include the following: In the stage of building a general knowledge vector base, it is first necessary to collect widely applicable drilling construction-related text data within the industry. This data includes drilling process descriptions, drill bit and drilling fluid usage guidelines, typical formation response descriptions, equipment operation manuals, complex situation handling solutions, construction case reports, and industry standards and specifications. After acquisition, the text is cleaned and formatted, removing redundant and low-quality content, and then segmented or divided by semantic blocks. Next, a pre-trained text embedding model (such as BERT, SBERT, or embedding API) is used to convert it into vectors and store them in a vector database (such as FAISS, Milvus, etc.) to form a general knowledge base available for semantic retrieval. In the target well personalization enhancement stage, it is necessary to collect local data of the target area, such as the geological structure background, stratigraphic division, fault development, and formation pressure prediction of the well, while integrating drilling records from adjacent wells, including ROP, drilling pressure, well inclination, well depth time curves, complex situation descriptions, logging reports, and construction logs. These local textual information are also cleaned, segmented, and embedded, and then merged into the existing vector library as new data. The semantic content related to the target well in the library is continuously enhanced through incremental vector updates. Furthermore, the results can be recalled by segmenting the vector data by "tags" (such as "general," "neighboring wells," and "target wells") and combining them by weight.
[0199] In this scenario example, the external knowledge base can also be updated in real time according to the drilling stage. Specifically, the drilling process has distinct stage characteristics, such as drilling, casing running, tripping, lost circulation control, and stuck pipe handling. Each stage may face different engineering challenges and information needs. Therefore, the entire drilling process can be divided into key nodes or typical events, and targeted data collection can be triggered at each node. For example, when lost circulation first appears, the system automatically collects formation information, circulating loss, drilling fluid properties, and historical lost circulation handling records for the current well section; when stuck pipe is encountered, mechanical parameters (such as drilling pressure and torque), wellbore trajectory, similar events in adjacent wells, and their handling processes are collected.
[0200] In this scenario example, during drilling, when it's necessary to determine the existence of potential risks, a first retrieval strategy can be employed based on real-time collected drilling status text (such as abnormal total pool volume, sudden drop in outlet flow, pump pressure fluctuations, etc.). This strategy emphasizes the weight of features reflecting risk signs (such as keywords like "decline / rise," "abnormal," "mutation," etc.) to construct a first vector. This vector is then compared with text embeddings in an external vector knowledge base to calculate similarity, retrieving historical cases and descriptions of abnormal patterns highly similar to the current state. This constitutes the first knowledge data, used to assist in determining the existence of risks. Once a drilling risk is confirmed, a second retrieval strategy can be employed. Here, higher weights are assigned to semantic dimensions related to "environmental characteristics," "geological conditions," and "operational behavior" in the text, generating a second vector to capture semantics matching the cause of the risk. By retrieving this second vector from an external knowledge base, historical cases, expert explanations, or report fragments that led to similar risks under similar geological structures, equipment parameters, or operating modes can be obtained, forming the second knowledge data to assist engineers in understanding the causes of the risk. When action recommendations are needed, the text information from the first two stages (current drilling status + cause explanation) can be combined and input to construct text with richer semantic relationships. Then, the corresponding feature weight distribution is adjusted according to the third retrieval strategy, with particular emphasis on words such as "treatment measures," "operational recommendations," and "successful recovery," generating a third vector. Using this vector for semantic matching in an external knowledge base, best practices for dealing with the current situation and recommended parameter adjustment schemes can be retrieved, forming third knowledge data to support strategy formulation at the drilling site.
[0201] In this scenario example, the generated prompts may also include the following: During the risk detection phase, real-time collected drilling status text (such as abnormal total pool volume, sudden drop in outlet flow, pump pressure fluctuations, etc.) can be combined with similar historical cases retrieved from an external knowledge base to construct prompts for risk identification. The prompts should focus on the judgment objective of "whether a risk exists," clearly defining the key characteristics of the current state (such as well depth, formation type, parameter variation range), and appropriately introducing historical event summaries or typical patterns to help the large language model form contextual semantic cognition, thereby improving its ability to identify potential risks. For example, the prompts could be organized around "Whether there is a drilling risk currently, please make a judgment based on the following case information," ensuring the prompts are relevant to the current situation and clear in language. During the cause explanation phase, the prompts should be constructed around the goal of "analyzing the causes of the current risk." The input information should include drilling status text and matched cause explanation knowledge data, such as expert analysis or descriptions of typical drilling risk formation mechanisms. In the design of prompts, key parameter change information (such as abnormal torque and well deviation) should be retained, and explanatory fragments of similar historical situations should be embedded to guide the language model to focus on the causal chain most likely related to the current risk. The structure of the prompts can adopt the approach of "combining the current state and the following case analysis content, inferring the cause of the risk and explaining the mechanism," which helps improve the logic and engineering interpretability of the model-generated content. In the measure recommendation stage, the goal of the prompts is to guide the model to propose practical and feasible treatment suggestions based on the current drilling status, risk causes, and historical experience. At this time, the drilling status text, the generated risk explanation content, and the retrieved treatment solution knowledge data should be integrated to generate a comprehensive input. The design of prompts needs to emphasize the semantic intent of "proposing solutions," clearly define the failure point, cause, and the goal to be achieved, and provide a reference path for the model by using historical successful cases or parameter adjustment records. An effective prompt can adopt a format such as "the current situation is..., the known cause is..., please recommend treatment measures based on the following experience fragments," which improves the accuracy and practicality of the model output.
[0202] In this scenario example, training a large-scale model can include the following: First, in terms of instruction tuning, a multi-round instruction corpus centered on drilling tasks should be constructed, including various task-oriented input-output pairs such as "determine if there is a risk of well leakage," "explain possible causes of stuck pipe," and "recommend operating procedures for increased pump pressure." Unlike conventional general tuning, these instructions need to closely match the on-site language style, terminology, and engineering logic, ensuring that the model not only understands the terminology but also provides accurate judgments, analyses, and suggestions in terms of structure. If necessary, expert-annotated question-and-answer pairs should be used as the training set for supervised fine-tuning. Second, in terms of in-context learning, drilling tasks require the model to extract key points from multi-source information fusion (such as state data + knowledge fragments + historical cases). Therefore, context is not only supplementary background but also a source of clues for "controlling thinking." Under the RAG architecture, when providing context through knowledge retrieval, an organized input structure should be designed, such as presenting "state data + similar cases + prior explanations" in blocks, to improve the model's overall grasp of the scenario. Compared to general dialogue models, drilling tasks place greater emphasis on the relationships between "engineering causality" and "parameter evolution" within the context. Regarding the guidance of the chain of thought, to ensure the model not only provides results but also explains the judgment process, a "step-by-step reasoning" approach to answer organization should be introduced during training. For example, in causal analysis, the model should be guided to write: "Given that the lithology of the well section is tight sandstone, the decrease in ROP may be caused by drill bit dulling or cuttings blockage; combined with the increase in pump pressure, it may be due to insufficient bottom hole cleaning," rather than simply outputting "possible bottom hole blockage." Learning this type of reasoning chain requires clearly defining the modeling thought process steps in the training data and explicitly guiding the model to "explain the basis of the judgment step by step" in the prompts.
[0203] In this scenario example, when no matching knowledge data is found, the following steps can be taken. Considering that the current external knowledge base mainly contains conventional typical knowledge and historical drilling data for the target area, when the target well experiences an unprecedented new state, it may be impossible to retrieve matching information from the knowledge base. To address this "knowledge blind spot," several strategies can be adopted: First, the system should have a similarity confidence assessment and recall failure prompt mechanism. When no valid match can be found, the user should be promptly informed that the state may be a new type of anomaly, prompting manual intervention. Second, the common sense reasoning and analogy capabilities of the large language model itself can be utilized. Through carefully designed prompts, the model can be guided to make inferences at the physical or engineering mechanism level based on the current state parameters, providing reasonable analysis even in the absence of direct knowledge support. Furthermore, an expert feedback mechanism can be introduced, allowing manual supplementation of analysis results or handling measures after a new problem is solved, and embedding these contents as new knowledge fragments into the knowledge base to support subsequent retrieval.
[0204] The above scenario examples demonstrate that the drilling risk management method based on a large language model provided in this manual can indeed quickly identify potential risks and provide detailed descriptions of the causes of risks, helping on-site engineers to take timely and targeted measures, thereby improving drilling efficiency and safety.
[0205] While this specification provides the steps of operation for the methods described in the embodiments or flowcharts, more or fewer steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or client product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded. The terms "first," "second," etc., are used to denote names and do not indicate any particular order.
[0206] Those skilled in the art will also know that, besides implementing the controller using purely computer-readable program code, the same functions can be achieved by logically programming the method steps, making the controller function as logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers (PLCs), and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the devices within it used to implement various functions can also be considered structures within that hardware component. Alternatively, the devices used to implement various functions can be considered as both software modules implementing the method and structures within a hardware component.
[0207] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer-readable storage media, including storage devices.
[0208] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this specification can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of this specification can essentially be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments of this specification.
[0209] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. This specification can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.
[0210] Although this specification has been described by way of examples, those skilled in the art will recognize that many variations and modifications are possible without departing from the spirit of this specification, and it is intended that the appended claims cover such variations and modifications without departing from the spirit of this specification.
Claims
1. A drilling risk management method based on a large language model, characterized in that, include: Obtain the current drilling status parameters of the target well; Based on the current drilling status parameters of the target well, determine the matching status perception data for the current drilling status of the target well; wherein, the status perception data includes at least: geological attribute information, wellbore flow information, and change information of key engineering parameters; Based on the state-aware data, the current drilling status text of the target well is constructed; Based on the current drilling status text of the target well and a preset external knowledge base, target knowledge data that matches the current drilling status of the target well is determined through similarity matching. The preset external knowledge base is constructed according to preset construction rules, using drilling domain knowledge information and knowledge information of the target area to which the target well belongs, and is associated with the target well. The preset external knowledge base includes multiple knowledge data. Based on the target knowledge data and the current drilling status text of the target well, generate target prompt words for the current drilling status of the target well; By using a pre-defined large language model and processing target prompts, it can determine whether there is a drilling risk in the target well, and if so, determine the corresponding risk cause explanation and / or risk management measures.
2. The method according to claim 1, characterized in that, The geological attribute information includes at least: formation lithology, formation pressure, and rock mechanics parameters; the wellbore flow information includes at least: drilling fluid density, viscosity, rheological properties, and solid content; the key engineering parameters include at least: hook height, drilling pressure, torque, rotational speed, riser pressure, pump flush, casing pressure, total tank volume, outlet flow rate, inlet flow rate, outlet density, outlet density, and gas measurement value.
3. The method according to claim 1, characterized in that, Based on the current drilling status text of the target well and a pre-set external knowledge base, target knowledge data matching the current drilling status of the target well is determined through similarity matching, including: Convert the current drilling status text of the target well into the corresponding current drilling status vector of the target well; Using the current drilling status vector of the target well, a similarity match is performed on a preset external knowledge base to determine a preset number of candidate knowledge data with the highest matching degree. Based on the candidate knowledge data, the target knowledge data is determined.
4. The method according to claim 3, characterized in that, The target cue words include at least: a first cue word related to drilling risk detection; Accordingly, the method further includes: The current drilling status text of the target well is transformed into a semantic vector using a pre-defined semantic processing model to obtain the corresponding first intermediate vector. Based on the first retrieval enhancement rule, the first key feature in the current drilling status text of the target well and the weight coefficient of the first key feature are determined. Based on the first key feature and its weight coefficient, the first intermediate vector is adjusted to obtain the first drilling state vector for drilling risk detection of the current target well. Using the first drilling state vector, the first knowledge data is determined by similarity matching with a preset external knowledge base; Based on the first knowledge data and the current drilling status text of the target well, a first prompt word is constructed for the target well; wherein, the first prompt word is used to input into a preset large language model for reasoning to determine whether there is a drilling risk in the target well.
5. The method according to claim 4, characterized in that, If it is determined that there is a drilling risk in the target well, the target prompt words also include: a second prompt word related to the explanation of the cause of the drilling risk, and a third prompt word related to the risk management measures for the drilling risk.
6. The method according to claim 5, characterized in that, After obtaining the corresponding first intermediate vector, the method further includes: Based on the second retrieval enhancement rule, the second key feature in the current drilling status text of the target well and the weight coefficient of the second key feature are determined. Based on the second key feature and its weight coefficient, the second intermediate vector is adjusted to obtain the second drilling state vector for explaining the risk causes of the current target well. Using the second drilling state vector, the second knowledge data is determined by similarity matching with a preset external knowledge base; Based on the second knowledge data and the current drilling status text of the target well, construct a second prompt word for the target well; By using a pre-defined large language model and processing the second cue word, the risk causes of drilling risks for the target well are determined.
7. The method according to claim 6, characterized in that, After determining the risk causes for the drilling risk of the target well, the method further includes: Based on the third retrieval enhancement rule, the current drilling status text of the target well and the explanation of the risk reasons are fused to obtain the corresponding joint text; and the joint text is processed by semantic vector transformation using a preset semantic processing model to obtain the corresponding third intermediate vector. Based on the third retrieval enhancement rule, the third key feature in the joint text and the weight coefficient of the third key feature are determined; Based on the third key feature and its weight coefficient, the third intermediate vector is adjusted to obtain the third drilling state vector for risk management measures for the current target well. Using the third drilling state vector, the third knowledge data is determined by similarity matching with a pre-set external knowledge base; Based on the aforementioned third knowledge data, risk cause explanations, and the current drilling status text of the target well, a third prompt word is constructed for the target well; By using a pre-defined large language model and processing third-party prompts, the system determines the appropriate measures to address drilling risks in the target well.
8. The method according to claim 1, characterized in that, The method further includes: We will acquire and construct a basic general knowledge base for the field of drilling engineering based on classic drilling engineering theories, industry standards, and typical case documents. Obtain geological background data of the target area where the target well is located, as well as drilling records of adjacent wells already drilled in the target area; Based on the geological background data of the target area and the drilling records of adjacent wells of the target well in the target area, the geological-engineering data information of the target area is extracted. By integrating geological and engineering data information of the target area into the basic general knowledge base, a preset external knowledge base associated with the target well is obtained.
9. A drilling risk management device based on a large language model, characterized in that, include: The acquisition module is used to obtain the current drilling status parameters of the target well; The first determining module is used to determine matching state perception data about the current drilling state of the target well based on the current drilling state parameters of the target well; wherein, the state perception data includes at least: geological attribute information, wellbore flow information, and change information of key engineering parameters; The construction module is used to construct the current drilling status text of the target well based on the status perception data; The second determining module is used to determine target knowledge data that matches the current drilling status of the target well based on the current drilling status text of the target well and a preset external knowledge base through similarity matching; wherein, the preset external knowledge base is a knowledge base constructed according to preset construction rules, using drilling domain knowledge information and knowledge information of the target area to which the target well belongs, and is associated with the target well; the preset external knowledge base includes multiple knowledge data. The generation module is used to generate target prompt words for the current drilling status of the target well based on the target knowledge data and the current drilling status text of the target well. The third determination module is used to determine whether there is a drilling risk in the target well by processing target prompt words using a preset large language model, and if it is determined that there is a drilling risk in the target well, to determine the corresponding risk cause explanation and / or risk treatment measures.
10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Deep well drilling risk identification method
CN116777221A
Underground coal mine early warning data query method and system based on language large model
CN118733608A
Well drilling difficulty prediction and difficulty solution generation method, device and equipment
CN118885863A
Well engineering intelligent question-answering system and method
CN119088929A
Bid evaluation early warning method and system based on large language model
CN119648128A
Cited By
Drilling scheme recommendation method and device, computer program product and electronic equipment
CN121119783A
Rescue drilling machine operation parameter prediction method and system based on large model
CN121168675A
A large model-based rescue drilling rig operation parameter prediction method and system
CN121168675B
Rock-soil layer intelligent analysis method and device based on drilling data
CN121637196A