A large model-based rescue drilling rig operation parameter prediction method and system
By fusing sensor data with natural language descriptions using a large model-based approach, high-precision predictions of drilling pressure, rotational speed, and torque are generated. This solves the problem of existing rescue drilling rigs relying on human experience, enabling intelligent decision-making and efficient rescue.
Patent Information
- Application Number
- CN202511713682.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-11-21
AI Technical Summary
Current rescue drilling operations rely on manual experience and lack intelligent and systematic parameter decision-making mechanisms, resulting in slow start-up, delayed decision-making, and unclear operation plans, which reduces rescue efficiency and success rate, and makes it difficult to handle complex and dynamic geological conditions and sudden working conditions.
By employing a large model-based approach that integrates drilling rig sensor data with natural language descriptions of operating conditions, and through semantic understanding and causal reasoning of the large model, high-precision, real-time predictions of parameters such as drilling pressure, rotational speed, and torque are generated, providing interpretable operational suggestions and constructing a closed-loop system for real-time data acquisition, semantic representation of operating conditions, predictive reasoning, and feedback optimization.
It achieves high-precision parameter prediction, reduces reliance on human experience, adapts to complex working conditions, improves rescue efficiency and safety, outputs interpretable risk assessments and operational recommendations, and enhances the success rate of rescue operations.
Smart Images

Figure CN121168675B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for predicting the operational parameters of a rescue drilling rig based on a large model, belonging to the field of intelligent technology for rescue equipment. Background Technology
[0002] Mining environments are complex, with frequent occurrences of emergencies such as gas explosions, water inrushes, collapses, and blowouts. These accidents are often characterized by their suddenness, unpredictability, and high difficulty in rescue. Once an accident occurs, a lifeline must be established using rescue drilling rigs as quickly as possible. However, current rescue drilling rigs rely primarily on manual experience and on-site judgment, lacking intelligent and systematic parameter-based decision-making mechanisms. This often leads to slow startup, delayed decision-making, and unclear operational plans in actual rescue operations, severely reducing rescue efficiency and success rates.
[0003] Current research on improving the efficiency and intelligence of rescue drilling rigs mainly focuses on two aspects: first, improving the performance of mechanical equipment, such as drilling rig power, drill bit materials, and hydraulic systems; and second, decision support based on geological modeling and expert experience bases. However, intelligent decision-making methods based on expert knowledge bases still have the following significant shortcomings: 1. Insufficient real-time performance: Traditional methods rely on historical experience bases or static geological modeling, making it impossible to quickly analyze and predict large-scale time-series data collected in real time. 2. Insufficient intelligence: Existing decision-making systems are mostly rule-driven, making it difficult to handle complex working condition semantics and nonlinear relationships between multidimensional parameters. 3. Insufficient adaptability: Faced with complex and dynamic geological conditions and sudden working conditions, traditional methods lack flexible reasoning and prediction capabilities, making it difficult to generate personalized and scenario-based drilling rig operation suggestions.
[0004] With the development of artificial intelligence and large language models (LLM), it has become possible to introduce them into the prediction of operational parameters for rescue drilling rigs. Large models can not only integrate multi-source heterogeneous data, but also combine historical cases and operational semantics to achieve contextual understanding, causal reasoning, and real-time prediction. This provides a new technological path for intelligent decision-making and autonomous control of rescue drilling rigs. Summary of the Invention
[0005] The technical problem this invention aims to solve is to overcome the shortcomings of existing technologies and provide a method and system for predicting operational parameters of rescue drilling rigs based on a large model. Its main objectives include: 1. Improving prediction accuracy and real-time performance: By fusing raw sensor data from drilling rig sensors with natural language descriptions of operating conditions, high-precision, real-time prediction of key parameters such as drilling pressure, rotational speed, and torque of the rescue drilling rig is achieved. 2. Enabling intelligent decision-making and assisted control: Utilizing the semantic understanding and reasoning capabilities of the large model, interpretable structured prediction results and operational suggestions are generated to assist operators in making rapid and scientific decisions. 3. Improving emergency rescue efficiency and safety: Through intelligent prediction and operational suggestions, the speed of opening rescue drilling rig access routes is accelerated, safety risks during the rescue process are reduced, and the lives of miners are protected. In summary, this invention not only introduces the contextual understanding and reasoning capabilities of a large model at the algorithm level but also constructs a complete closed loop at the system level, encompassing real-time data acquisition, semantic representation of operating conditions, predictive reasoning, historical calibration, and feedback optimization, thereby significantly improving the intelligence level and rescue efficiency of rescue drilling rigs.
[0006] Prior to this invention, a method for predicting operational parameters of a rescue drilling rig based on a large model is provided, comprising: preprocessing raw sensor data to obtain standardized data; identifying the working condition of the standardized data based on preset rule criteria to obtain working condition category labels; semanticizing the instantaneous state of the standardized data to obtain JSON objects; semanticizing the process of the standardized data to obtain natural language fragments; generating intelligent drilling operation prompts based on the standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions; inputting the intelligent drilling operation prompts into a large model, using the large model to establish relationships between the standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions, and performing causal reasoning to output structured prediction results and alarm information including working condition category labels and predicted values of key parameters within a preset future time period.
[0007] Prioritizing the output of structured prediction results and alarm information, including working condition category labels and predicted values of key parameters within a preset future timeframe, a query vector is generated based on pre-acquired contextual descriptions, pre-acquired historical cases, predicted values of key parameters within the preset future timeframe, corresponding operational suggestions, and a risk assessment of the current work status. Based on the query vector, similarity matching is performed in a vector database to retrieve several historical cases from the vector database that exceed a preset relevance threshold. Similarity matching is then performed between the query vector and the historical cases from the vector database, and weights are assigned to the historical cases based on the similarity matching. The predicted values of key parameters within the preset future timeframe are then weighted and fused to obtain the calibrated final prediction value. Based on the calibrated final prediction value, corresponding operational suggestions, and a risk assessment of the current work status, structured operational suggestions and natural language explanations are generated.
[0008] First, instantaneous state semantics are performed on the standardized data. Sensor data at the end of the current time window is selected and combined with the working condition category label, operation status, equipment operating status and working environment status to obtain a JSON object.
[0009] Prior to, after performing instantaneous state semanticization on the standardized data, process semanticization is performed on the standardized data to obtain natural language fragments, including: extracting trend features of the standardized data based on the least squares method; extracting volatility features of the standardized data based on the coefficient of variation; comparing the standardized data with threshold intervals in the expert knowledge base to obtain threshold determination features; and obtaining natural language fragments based on trend features, volatility features, and threshold determination features.
[0010] Prioritizing the generation of intelligent drilling operation prompts based on standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions, the process includes: constructing an intelligent drilling operation prompt template using pre-defined predictive analysis prompts, diagnostic and attribution prompts, optimization suggestion prompts, and emergency response prompts; extracting working condition information, key sensor numerical sequences from standardized data, contextual information, analogical information from historical cases, and pre-acquired task instructions from standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions; and inputting the working condition information, key sensor numerical sequences from standardized data, contextual information, analogical information from historical cases, and pre-acquired task instructions into a large model based on the intelligent drilling operation prompt template to generate intelligent drilling operation prompts.
[0011] Prior to this, the predicted values of key parameters within the preset time frame include the mechanical drilling rate ROP_next_60s and the rescue rig torque Torque_next_60s; alarm information includes alarm type, severity, triggering evidence, and corresponding handling suggestions.
[0012] Prioritized, both raw sensor data and sensor data include rescue rig drilling pressure, rescue rig rotation speed, rescue rig lifting force, rescue rig suspended weight, rescue rig power head torque, rescue rig lifting pressure, rescue rig pressing pressure, rescue rig rotation pressure, and well depth; operating condition categories include tripping in / out, normal drilling, rescue drilling, and emergency stoppage; JSON objects include timestamps, operating condition types, operating status, and instantaneous values of key parameters; contextual descriptions include operator logs and geological environment information.
[0013] Firstly, the raw sensor data is preprocessed to obtain standardized data, including preprocessing the raw sensor data such as cleaning, smoothing, normalization and Z-score standardization to obtain standardized data.
[0014] Prior to this, a rescue drilling rig operation parameter prediction system based on a large model, based on any one of the large model-based rescue drilling rig operation parameter prediction methods, includes: a data acquisition layer module for preprocessing raw sensor data to obtain standardized data; a working condition identification and semanticization layer module for performing working condition identification on the standardized data based on preset rule criteria to obtain working condition category labels; performing instantaneous state semanticization on the standardized data to obtain JSON objects; performing process semanticization on the standardized data to obtain natural language fragments; and a large model inference layer module for generating intelligent drilling operation prompts based on standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions; inputting the intelligent drilling operation prompts into a large model, and using the large model to establish a system based on standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, and pre-acquired task instructions. The system analyzes the relationship between historical cases and pre-acquired task instructions, performs causal reasoning, and outputs structured prediction results including work condition category labels, predicted values of key parameters within a preset future timeframe, and alarm information. The result calibration layer module generates a query vector based on pre-acquired contextual descriptions, pre-acquired historical cases, predicted values of key parameters within a preset future timeframe, corresponding operational suggestions, and a risk assessment of the current work status. Based on the query vector, it performs similarity matching in a vector database to retrieve several historical cases from the vector database that exceed a preset relevance threshold. It then performs similarity matching between the query vector and the historical cases in the vector database, assigns weights to the historical cases based on the similarity matching, and performs weighted fusion of the predicted values of key parameters within a preset future timeframe to obtain the final calibrated prediction value. The result feedback layer module generates structured operational suggestions and natural language explanations based on the final calibrated prediction value, corresponding operational suggestions, and a risk assessment of the current work status.
[0015] Preferably, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the methods described herein.
[0016] Preferably, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described herein.
[0017] The beneficial effects achieved by this invention are as follows:
[0018] 1. This invention offers high prediction accuracy and interpretable results. By integrating drilling rig time-series data with operational semantic information, and combining historical case comparisons with RAG calibration, this invention enables high-precision prediction of key parameters such as rescue rig load on bit (WOB), rescue rig rotational speed (RPM), and rescue rig torque. Furthermore, the output of this invention includes the reasoning process and risk assessment, ensuring the interpretability of the structured prediction results and facilitating operator understanding and adoption.
[0019] 2. This invention is highly adaptable and suitable for complex working conditions. The system can automatically adjust the prediction logic and parameter weights according to different working conditions, exhibiting high adaptability. Even in situations with complex geological formations, variable environments, or unstable equipment conditions, it can provide stable and reliable predictions and suggestions.
[0020] 3. This invention reduces reliance on human experience. Traditional rescue drilling rig operations heavily depend on experienced operators, while the intelligent prediction and decision support provided by this invention can effectively reduce reliance on human experience, providing actionable operating plans for inexperienced operators, thereby improving the overall level of rescue operations.
[0021] 4. This invention improves safety and rescue success rate. In addition to providing predictive values, this invention also outputs risk assessments and response strategies, helping operators to identify potential risks in a timely manner and avoid dangerous operations in advance, thereby significantly improving the safety and success rate of rescue operations.
[0022] 5. This invention has broad application prospects. In addition to coal mine rescue drilling rigs, the technical framework of this invention can also be extended to fields such as oil drilling, geothermal development, and tunnel excavation that require complex parameter prediction and working condition understanding, thus possessing wide application value. Attached Figure Description
[0023] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating the present invention.
[0025] Figure 2 This is a schematic diagram of the overall architecture of the rescue drilling rig operation parameter prediction system in this invention. Detailed Implementation
[0026] See Figure 1This invention provides a method for predicting operational parameters of rescue drilling rigs based on a large model. The core advantage of this method lies in its ability to effectively integrate massive amounts of real-time sensor data generated by the rescue drilling rig with natural language text information describing the operational conditions. By leveraging the powerful contextual understanding and generation capabilities of the large model, it can accurately predict key parameters such as drilling pressure and rotational speed of the rescue drilling rig over a future period and provide interpretable prediction basis.
[0027] Step 1: Data Acquisition and Preprocessing. Step 1.1: Data Acquisition and Channel Definition. This invention utilizes a programmable logic controller (PLC) system for the rescue drilling rig, combined with an industrial gateway and sensor interface, to acquire raw sensor data from the rescue drilling rig during its operational cycle at a fixed sampling frequency, forming a time-series data set.
[0028] Raw sensor data includes, but is not limited to: wobble weight (WOB), rotational speed (RPM), lifting force, rotor weight, torque, lifting pressure, downforce pressure, rotation pressure, and bit depth. Bit depth is a core parameter in drilling engineering, referring to the vertical distance from the wellhead to the bottom of the well, usually with the rotary table centering surface as the reference point.
[0029] To ensure efficient data transmission and synchronization, this invention uses the Modbus-TCP protocol as the interface protocol for raw sensor data, and achieves rapid transmission of raw sensor data by converting CAN bus to Ethernet.
[0030] Raw sensor data is aggregated and centrally processed through an industrial gateway to form a continuous, time-consistent data stream, which serves as input for subsequent processing steps.
[0031] Step 1.2, Data Cleaning. The acquired raw sensor data is preprocessed to enhance the stability and generalization ability of the large model. Forward imputation and moving average methods are used to clean the raw sensor data. Standardization and normalization are then applied to eliminate scale effects between different physical quantities in the raw sensor data, making it easier for the large model to process.
[0032] The preprocess_data method ensures that all key feature columns in the original sensor data are present in the original sensor data, and arranges the original sensor data in a preset order, removing irrelevant or missing features to ensure the consistency and reproducibility of subsequent processing.
[0033] Furthermore, a forward-filling method is used to correct missing values in the original sensor data. That is, for any missing original sensor data at a given moment, it is directly filled in using valid original sensor data from the previous moment. This method has the following advantages: 1. It facilitates efficient and rapid processing of real-time data streams; 2. It is applicable to the actual situation of short-term sensor packet loss under drilling conditions, and can restore the true working conditions to the greatest extent; 3. It ensures the continuity and temporal consistency of the original sensor data, avoiding fluctuations without physical meaning introduced by interpolation.
[0034] For a very small number of consecutive missing points in the original sensor data that cannot be covered by forward padding, backward padding is used to supplement the missing original sensor data to ensure the integrity of the original sensor data at all times.
[0035] For outliers detected in the raw sensor data, such as extreme values exceeding the physically reasonable range, they are directly removed and corrected using forward padding to prevent the abnormal raw sensor data from interfering with the training and prediction of large models. Specifically, the `fillna(method='ffill')` and `fillna(method='bfill')` functions from the Pandas library are used to perform forward padding first, followed by backward padding, ensuring that all missing values are effectively filled. The `fillna` function is a method in the pandas library used to handle missing values, obtaining the padded raw sensor data.
[0036] After the missing value imputation step is completed, the imputed raw sensor data needs to be smoothed. For each sensor channel, a simple moving average (SMA) method is used to smooth the imputed raw sensor data to obtain smoothed raw sensor data. Assume the sampling frequency is... 1Hz, let the original time series be: Let the size of the sliding window be... Then at time moving average for: Prediction standard deviation for: This method effectively extracts medium- and long-term trend changes in drilling operations while filtering out short-term high-frequency noise, enhancing the large model's ability to discriminate the overall working condition. A sliding window smoothing mechanism is implemented using Pandas' `rolling().mean()` function. The window size is set to 48, which effectively smooths short-term fluctuations, highlights the long-term trend of drilling conditions, and balances real-time performance with noise suppression capabilities, making it suitable for the actual needs of rescue drilling rigs operating at a 1Hz sampling frequency.
[0037] Step 1.3, Normalization and Standardization. After smoothing, in order to eliminate the scale effect between different physical quantities and improve the convergence speed and stability of large model training, this invention performs normalization and standardization on all smoothed original sensor data.
[0038] First, min-max normalization is used to scale the smoothed raw sensor data to the [0,1] interval to obtain normalized time series data: Secondly, the distribution-sensitive model module uses the Z-score normalization method to further convert the normalized raw sensor data into normalized time series data.
[0039] Distribution-sensitive model modules refer to temporal modeling structures whose performance depends on the distribution of input features during the training and inference of large models. Such temporal modeling structures typically include, but are not limited to: recurrent neural networks (such as RNN, LSTM, GRU), convolutional neural networks (CNN), Transformer-type temporal models based on attention mechanisms, and other deep learning models that require the input to satisfy a zero-mean, unit-variance distribution.
[0040] For the aforementioned distribution-sensitive model module, to avoid instability in large model training due to dimensional differences and mean / variance differences between different sensor channels, this invention employs the Z-score normalization method for its input time series data to obtain standardized data. The calculation formula is as follows: ,in, σ represents the mean of all normalized data points in the standardized data of this sensor channel within the current sample window, and σ represents the standard deviation of all normalized data points in the current sample window for this sensor channel. This standardization process ensures that the data for each type of sensor channel follows a zero-mean, unit-variance distribution, making it more suitable for time-series deep learning model structures that are sensitive to feature distribution.
[0041] Step 2: Operating Condition Identification and Semantic Labeling. Effective processing and interpretation of massive amounts of standardized data during rescue drilling operations is a primary prerequisite for intelligent decision-making. The core task of this step is to transform continuous, purely numerical standardized data into structured information with clear semantics that can be understood by both large models and human experts. This process includes two key steps: First, based on a predefined operating condition classification system, the standardized data is discriminated to identify the operating condition category corresponding to the rescue drilling rig within the current time window, obtaining operating condition category labels. Second, the identified operating condition categories and their judgment criteria are structured and semantically processed to output standardized operating condition category labels and semantic descriptions, providing high-quality input for the prediction and inference of large models.
[0042] Step 2.1, Operating Condition Categories and Definitions. An operating condition refers to a specific working state formed by the interaction between the drill string and the formation during the drilling process of a rescue drilling rig. This invention includes four operating conditions for rescue drilling rig operations: tripping in / out, normal drilling, rescue drilling, and emergency shutdown. The purpose of distinguishing these operating conditions is to provide clear and standardized contextual semantics for subsequent large-scale models, enabling them to activate corresponding prediction and inference mechanisms under different operating conditions.
[0043] Step 2.2, Working Condition Identification Logic. Since the large model does not directly process the raw sensor data, but rather performs inference based on these "labeled" working conditions, this system must first complete the working condition identification.
[0044] This invention utilizes preset key parameter thresholds and feature combinations as rule criteria to achieve real-time identification of standardized data and working condition categories. The rule criteria are shown in Table 1.
[0045] Table 1. Rule Criteria
[0046]
[0047] Based on the above-mentioned rules and criteria, this system can identify the macroscopic working condition category corresponding to the current window from standardized data, and use it as input for subsequent semantic annotation and structured representation.
[0048] Step 2.3, Semantic Labeling and Vectorization. After identifying the macroscopic operating condition categories, it is necessary to further transform them into structured semantic information that can be processed by large models. This invention adopts two complementary structured methods: instantaneous state semanticization and process semanticization.
[0049] Step 2.3.1: During the semantic markup process, the output format of the standardized data is JSON (JavaScript Object Notation), a lightweight data exchange format characterized by its simplicity, readability, and ease of parsing. JSON is widely used for data transmission and storage, facilitating subsequent standardized data processing, storage, and transmission.
[0050] The system extracts the value of the last time point t from the standardized data and combines it with other non-time-series labels to form a JSON object describing the current time. For standardized data at a single time point, the structured JSON object has the following structure: {"timestamp":"2025-06-26T12:30:00","working_condition_type":"Drilling start / stop condition","operation_state":"Drilling start","pressure":5,"equipment_status":"Normal","environment":"Stable",}.
[0051] The JSON object includes timestamp, working_condition_type, pressure, equipment_status, environment, and operation_state. Specifically: timestamp represents the data timestamp; working_condition_type represents the working condition category; operation_state represents the current specific operational behavior; pressure represents the current drilling pressure of the rescue drilling rig, in kN; equipment_status represents the equipment operating status; and environment represents the status of the working environment.
[0052] Step 2.3.2: To describe the dynamic evolution over a period of time, this invention generates a process-level JSON structure based on standardized data within a sliding window. First, all standardized data within the current sliding window is obtained. Based on the standardized data, structured operation condition data is obtained: {"timestamp_series":["12:29:56","12:29:57",...,"12:30:00"],"operation_state":"Drilling","pressure_series":[48,49,49.5,50,50.3],"rpm_series":[120,120,120,120,120],"equipment_status":"Normal","environment":"Stable"}.
[0053] The structured chemical condition data includes a continuous time stamp sequence (timestamp_series), operation status (operation_state), time value sequence of drilling pressure of the rescue drilling rig (pressure_series), time value sequence of rotation speed of the rescue drilling rig (rpm_series), equipment status identifier of the rescue drilling rig (equipment_status), and environment status of the work site (environment).
[0054] Based on this, the system further extracts trend characteristics, volatility characteristics, and threshold judgment characteristics from the structured chemical condition data through a built-in discriminant model, identifies information such as drilling status and equipment status, and converts this information into natural language descriptions: 1. Trend Analysis Module: Based on least squares method fitting of straight lines, calculates the slope. The system categorizes the trend characteristics of the structured chemical condition data as "stable," "slowly rising," "rapidly rising," "slowly falling," or "rapidly falling" based on preset thresholds. 2. Statistical Analysis Module: Calculates the coefficient of variation. ,in The mean within the window. The standard deviation is used. The volatility characteristics of the structured chemical condition data are judged as "very stable," "volatile," or "volatile" based on the threshold. 3. Threshold Comparison Module: This module compares the key parameters in the structured chemical condition data with the threshold ranges in the expert knowledge base, outputting the threshold judgment characteristics as "normal," "alarm," or "severely exceeded."
[0055] Finally, the above tags are integrated into natural language statements. For example, through the combination of the three specific algorithm modules mentioned above, the system can deterministically extract trend features, volatility features, and interval state features from a piece of structured condition data, and transform these features into standardized tags that can be directly used by the subsequent natural language generation module, thereby completing the entire semanticization process.
[0056] After integrating these analysis results, the following natural language fragment was formed: "Currently in drilling state, the drilling pressure of the rescue drilling rig continues to rise, the speed of the rescue drilling rig is stable at 120 rpm / min, the pressure is 50 KN, the equipment is operating normally, and the environment is stable." This can provide accurate and rich information for subsequent large model inference.
[0057] The built-in discriminant model in this invention is a set of built-in algorithm modules for analyzing time series characteristics, mainly consisting of three modules: 1. Trend analysis module: a trend determination algorithm based on least squares linear regression. Input: standardized data from step 1.3. Solution: Fit a straight line using the Ordinary Least Squares method. The core of this step is calculating the slope of the line. .
[0058] ,in: It is the average of time indices; It is the average of standardized data.
[0059] Define slope threshold =0.05 or =0.5. Based on the calculated slope The value, based on the slope threshold, is categorized into a preset trend category, and a trend description label is output, such as "greater than the slope threshold". It means "rapid rise".
[0060] 2. Statistical Analysis Module: Volatility Quantification Algorithm Based on Coefficient of Variation. The input to the statistical analysis module is the standardized data from step 1.3. This module is used to evaluate the dispersion and volatility of standardized data within a time window. It calculates the arithmetic mean of the standardized data. Calculate the standard deviation σ of the standardized data to measure the absolute dispersion of the standardized data: Similar to the trend analysis module, this invention also defines a minimum volatility level threshold for the coefficient of variation. =0.02 and the highest volatility level threshold =0.5, by The Cvvolatile variable is calculated.
[0061] Based on the mean within the window Standard deviation And a preset threshold, determine the fluctuation level: if If the value is ≤0.02, the volatility level is "Very Stable"; if 0.02 ≤ If ≤0.5, then the volatility level = "Volatility exists"; if ≥0.5 indicates a volatility level of "Volatility". Output of the statistical analysis module: When... At that time, the volatility level is "violent volatility"; when When the fluctuation level is "very stable"; when In CV stable and CV volatility_max When it falls between these two, the volatility level is "Volatility Exists".
[0062] In addition, this module can also output the average value of the corresponding sensor channel within the time window. , which is used to visually describe the stable level of this channel during this period, such as "the drilling pressure of the rescue drill is stable at about 12 kN" or "the torque of the rescue drill is stable at about 3.0 kN·m". Among them, is the arithmetic mean of all sampling points of this sensor channel within the time window.
[0063] 3. Threshold comparison module: A state determination algorithm based on a multi-level early warning interval. The input of the threshold comparison module is the standardized data , and the preset threshold intervals in the expert knowledge base. This module is used to determine whether the standardized data in step 1.3 is in the normal, warning or dangerous working interval.
[0064] First, for each key parameter in the standardized data, predefined thresholds are set in the expert knowledge base, so as to divide into four thresholds including Critical_Low (L2): severely too low, Warning_Low (L1): warning too low, Warning_High (H1): warning too high, Critical_High (H2): severely too high. Then, traverse the standardized data within the time window for each data point , and determine its belonging interval. The belonging intervals include the interval (severely low): <L2, interval 2 (warning low): L2 ≤ <L1, interval 3 (normal): L1 ≤ ≤ H1, interval 4 (warning high): H1 < ≤ H2, interval 5 (severely high): > H2.
[0065] The final output of this module is determined by the worst state that appears within the time window. For example, even if 99% of the standardized data within the time window is in the "normal" interval, as long as any one point of the standardized data falls into the "severely high" interval, the state of the entire window is determined to be "severely overlimit". The final output of this module: a clear device or parameter status label, such as "normal", "high load warning", "severely overlimit". <000019^0>
[0066] Step 3: Construct an intelligent drilling operation prompt word template. The core task of this step is to act as the "senior translator" and "task scheduler" between the front-end data processing system and the back-end large model. It combines the structured working condition data and semantic working condition information generated in step 2 with real-time data, background knowledge and task objectives, and constructs a standardized input format that the large model can deeply understand and perform complex reasoning based on it - an intelligent drilling operation prompt word (Prompt).
[0067] Step 3.1, Categorization of Prompt Templates. Based on the different decision-making needs in drilling operations, this invention pre-defines four core prompt templates. Each prompt template targets a specific task, guiding the large model to complete different intelligent analyses by adjusting the structure and content of the prompts. The four core prompt templates are:
[0068] 1. Predictive Analysis Prompt. Purpose: To predict the changing trends of one or more key drilling parameters over a future period. This is the most commonly used template, designed to provide forward-looking guidance to operators.
[0069] 2. Diagnostic & Causal Analysis Prompt. Purpose: To investigate the root cause of phenomena when abnormal parameters or sudden changes in operating conditions are detected. For example, asking, "What could be the cause of the sudden and drastic fluctuation in the torque of the rescue drilling rig?"
[0070] 3. Optimization & Recommendation Prompt. Purpose: To request a set of optimal drilling parameter combinations from the large model for specific objectives (such as increasing drilling speed or protecting the drill bit).
[0071] 4. Emergency Response Prompt. Purpose: To request the safest and fastest emergency response plan and operating procedures from the large model when high-risk situations such as emergency drilling stop are detected.
[0072] Step 3.2: Unified Prompt Message Main Structure. In this invention, the condition category label and semantic description identified in Step 2 will serve as the core contextual information for intelligent drilling operation prompts. Furthermore, a unified prompt message main structure is constructed by combining real-time standardized data, operation logs, historical analogy cases, and task instructions. The prompt message main structure also adopts JSON format and contains five main parts, as shown below: 1. A condition information module is used to represent the condition information of the rescue drilling rig at the current moment. This module includes a condition category label (condition_label) and a natural language semantic description (semantic_description); 2. A real-time sequence data module is used to represent the key sensor value sequence in the standardized data within the most recent time window. This module includes a sequence description, a normalized rescue drilling rig drill pressure sequence (normalized_wob), a normalized rescue drilling rig speed sequence (normalized_rpm), and a normalized rescue drilling rig torque sequence (normalized_torque); 3. A contextual information module is used to represent contextual information during the operation process. This module includes operator... The system includes: 1. Operator log, 2. Geological information, and 3. Task goal; 4. A historical analogy module for representing analogy information with historical cases, including the case number (retrieved_case_id), similarity score (similarity_score), and case summary (case_summary); 5. A task instruction module for representing the instruction requirements of the current task, including the task type (query_type), task object (query_target), and output requirements (output_requirements), including the output format (format), whether to include reasoning (include_reasoning), and whether to include risk assessment (include_risk_assessment).
[0073] Step 4: Large-Scale Model Inference Analysis. Based on the intelligent drilling operation prompt word template, the large-scale model is input with operating condition information, key sensor value sequences from standardized data, contextual information, analogical information from historical cases, and pre-acquired task instructions to generate intelligent drilling operation prompt words. The large-scale model then uses contextual understanding and causal reasoning mechanisms to generate risk assessments, operational suggestions, and optimization plans regarding the current drilling status.
[0074] First, the intelligent drilling operation prompts are fed into the pre-trained BERT model. The BERT model, based on its built-in contextual understanding capabilities, parses and interprets the input standardized data and natural language descriptions. During this process, the BERT model generates corresponding semantic embedding vectors, transforming the standardized data into vector form in a high-dimensional space, facilitating subsequent reasoning and analysis.
[0075] Next, the large model performs inference analysis based on the input prompts. First, the large model decodes the standardized input data using its "contextual understanding" capabilities, identifying temporal features and operational background. For example, the large model can automatically identify the relationship between rescue rig bit weight (WOB) and rescue rig rotation speed (RPM), understanding their impact on drilling operations.
[0076] Subsequently, the large model uses a "causal inference" mechanism to analyze the correlation between standardized data and historical cases. This inference process is one of the model's core capabilities; it infers potential problems in the current operation based on pattern recognition in the standardized data. For example, the large model discovers that the current drilling pressure of the rescue rig is close to the threshold where problems have occurred historically, thus predicting the potential vibration risk in the output.
[0077] Based on the inference results of the large model, predictions are generated regarding the future WOB (Whole Bore Weight) and RPM (Rotation Speed) of the rescue rig, along with corresponding operational recommendations. The large model not only predicts future values but also provides specific operational suggestions and risk assessments. These outputs help operators understand operational risks and support their decision-making.
[0078] The output of large models is not limited to text suggestions; it can also generate structured prediction results. This facilitates quick understanding and application by operators. Among them, the rescue drilling rig torque Torque refers to the rotational force applied to the drill bit; the working condition semantics is a description of the current drilling operation conditions, such as rock formation type and environmental conditions. This is the error term, representing random variations or uncertainties that the large model failed to fully capture. This formula indicates that a future point in time can be predicted using current and past rescue rig load factor (WOB), rescue rig rotation speed (RPM), rescue rig torque (Torque), and operating condition descriptions. Error term This indicates the uncertainty that may exist in the forecast.
[0079] Finally, the system evaluates the confidence level of the structured prediction results and analyzes their reliability. High-confidence structured predictions provide operators with more confident operational suggestions, while low-confidence predictions encourage more cautious operation. The prediction function... Automatically generated from the contextual relationships encoded by the large model; For model fitting residuals, large models can automatically output confidence intervals: This formula quantifies the uncertainty of structured prediction results.
[0080] Step 5: Retrieval Enhanced Inference and Calibration RAG-CA. After completing the large model inference analysis in Step 4, the system enters the retrieval enhanced inference and calibration phase. By introducing external knowledge bases and historical sensor data, the structured prediction results of the large model are further improved, making the final operational recommendations more reliable and targeted. The key objective of this step is to combine real-time sensor data with historical experience to calibrate the output of the large model through retrieval and weighted calculation, providing operators with accurate decision support.
[0081] First, the system transforms the standardized data, operating condition descriptions, historical cases, and task objectives generated in step 4 into query vectors. These query vectors are generated based on the BERT model and a time-series model, and are combined with historical sensor data, current operating condition semantics, and task objectives to form the initial query input for retrieval. These query vectors include: text-based operating condition description vectors, time-series query vectors based on standardized data such as the rescue drilling rig's WOB and RPM, and historical case vectors extracted from an external knowledge base. These query vectors are represented as follows: ,Will Input into the vector retrieval system.
[0082] During the retrieval process, the vector retrieval system performs similarity matching between historical sensor data and current raw sensor data using the FAISS vector database. The FAISS vector database can quickly process large-scale data and supports similarity searches in high-dimensional space. Using the index stored in the FAISS vector database, the most similar working condition vector can be found based on the input query, enabling efficient retrieval and similarity matching in step 4. Specifically, the vector retrieval system retrieves the most relevant historical cases by calculating the cosine similarity between the query vector and historical sensor data vectors. (Cosine similarity) Here, q represents the query vector, d represents the historical data vector, and a higher cosine similarity value indicates a higher similarity. Through this process, the vector retrieval system can obtain historical cases most similar to the current working environment, operating conditions, and time-series data, and extract relevant information from them as evidence. .
[0083] Next, as Figure 2 As shown, the vector retrieval system performs contextual fusion on the retrieved evidence packets through a weighted mechanism. A weight is assigned based on the similarity between each evidence packet and the current query vector. The weights are calculated based on the retrieved similarity values. By using weighted summation and considering the correlation between various evidence packages, the vector retrieval system calibrates historical sensor data and structured prediction results to obtain predicted values for key parameters. This can be expressed by the following formula: ,in, These are predicted values of key parameters obtained from historical cases or search results. These are weighting coefficients. These are the predicted values of key parameters after final fusion. Based on this, the vector retrieval system further generates a risk assessment and operational recommendations regarding the current drilling status. The risk assessment is based on the confidence interval of the predicted values. The calculation basis for the confidence interval is as follows: ,in, It is the critical value of the normal distribution. It is the standard deviation of the structured prediction results. This represents the confidence interval of the structured prediction result. In this way, the vector retrieval system can quantify the reliability of the structured prediction result and generate a risk level accordingly.
[0084] When the confidence level is high, the vector retrieval system provides the operator with clear operational suggestions; when the confidence level is low, the system issues a warning and suggests careful adjustments. The formula for generating operational suggestions is: ,in, It is a function that generates operation suggestions. These are the predicted values for key parameters. To determine the optimal operating strategy based on environmental factors (such as rock type and equipment condition).
[0085] Finally, the vector retrieval system will determine the results based on the actual operation. With key parameter prediction values The system provides feedback and optimization to address discrepancies between the retrieved data and the structured prediction results. If significant deviations occur in practice, the vector retrieval system incrementally updates the large model to improve future prediction accuracy. This process involves adjusting the weights of the retrieved historical sensor data and evidence packages. The formula for weight updates is: ,in, This is the learning rate.
[0086] Through this feedback mechanism, the vector retrieval system can continuously optimize its retrieval algorithm and large model, thereby better adapting to the changing drilling environment. Each feedback and optimization process makes the vector retrieval system more intelligent and accurate when facing new drilling tasks.
[0087] Next, input the problem and the recalled knowledge into the large model together to generate task response content that conforms to the domain context, has accurate expression, and clear structure.
[0088] The output of the large model is in the comprehensive output format of a JSON object and a natural language explanation. An example of the output is as follows: {"condition":"normal_drilling", "prediction":{"ROP_next_60s":2.1, "Torque_next_60s":3.0}, "recommendations":[{"parameter":"RPM", "action":"increase", "target_range":[1500,1600], "confidence":0.85, "rationale":"ROP below expected, torquestable"}], "alerts":[{"type":"low_efficiency", "severity":"medium", "evidence":["ROP<threshold", "WOB stable"], "suggestion":"increase RPM gradually"}], "explanation":"Currently in normal drilling conditions, the predicted ROP in the next 1 minute is slightly low. It is recommended to increase the RPM of the rescue rig to 1500–1600 rpm to improve drilling efficiency"}.
[0089] The prediction and decision output results include the following parts: 1. Working condition category label "condition": indicating the current working condition category of the rescue drill rig, for example, "normal_drilling" represents the normal drilling working condition; 2. Structured prediction result "prediction": used to represent the predicted values of key parameters in the future time period, including the rate of penetration ROP_next_60s in the next 60 seconds and the torque of the rescue drill rig Torque_next_60s in the next 60 seconds; 3. Operation recommendations "recommendations": used to represent the parameter adjustment recommendations generated by the vector retrieval system based on the structured prediction results, including parameter name "parameter" (such as the RPM of the rescue drill rig), operation action "action" (such as increasing the propulsion speed "increase"), target range "target_range" (such as 1500–1600 rpm), recommended confidence "confidence", and the reason "rationale" for generating this recommendation; 4. Alarm information "alerts": used to represent the risk alarms identified by the vector retrieval system, including alarm type "type" (such as low efficiency "low_efficiency"), severity "severity" (such as medium), triggering evidence "evidence" (such as "ROP < threshold", "WOB stable"), and corresponding handling suggestions "suggestion" (such as "increase RPM gradually"); 5. Explanation "explanation": used to represent a comprehensive explanation in natural language form, for example, "Currently in the normal drilling working condition, it is predicted that the rate of penetration in the next 1 minute is slightly low. It is recommended to increase the RPM of the rescue drill rig to 1500–1600 rpm to improve the drilling efficiency."
[0090] In summary, by integrating historical data, real-time monitoring, and large model inference analysis, the vector retrieval system can provide more accurate and reliable structured prediction results and operation recommendations for operators, helping them make scientific and rational decisions in complex drilling operations and ensuring the efficient and safe execution of drilling operations.
[0091] In an embodiment of the present application, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in any one of the above are implemented.
[0092] In an embodiment of the present application, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0093] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0094] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention described herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not invented herein. The specification and embodiments are to be considered exemplary only.
[0095] The above specific embodiments further illustrate the purpose, technical solution and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this application should be included within the scope of protection of this application.
Claims
1. A method for predicting operational parameters of rescue drilling rigs based on a large model, characterized in that, include: The raw sensor data is preprocessed to obtain standardized data; Based on preset rule criteria, standardized data is used to identify working conditions and obtain working condition category labels; Perform instantaneous state semanticization on standardized data to obtain a JSON object; Semantizing standardized data to obtain natural language fragments; Based on standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions, intelligent drilling operation prompts are generated. The intelligent drilling operation prompts are input into a large model. The large model is used to establish the relationship between standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions. Causal reasoning is then performed to output structured prediction results and alarm information, including working condition category labels and predicted values of key parameters within a preset future time. After outputting structured prediction results and alarm information, including working condition category labels and predicted values of key parameters within a preset future time, a query vector is generated based on pre-acquired contextual descriptions, pre-acquired historical cases, predicted values of key parameters within a preset future time, corresponding operational suggestions, and risk assessment of the current work status. Based on the query vector, similarity matching is performed in the vector database to retrieve several historical cases from the vector database that are higher than the preset relevance threshold. The query vector and several historical cases in the vector database are matched for similarity. Based on the similarity matching, weights are assigned to the historical cases in the vector database. The predicted values of key parameters within a preset time period are weighted and fused to obtain the calibrated final predicted value. Based on the final calibrated predictions, corresponding operational recommendations, and risk assessments of the current operational status, structured operational recommendations and natural language explanations are generated.
2. The method for predicting rescue drilling rig operation parameters based on a large model according to claim 1, characterized in that, The standardized data is semantically represented in real time. Sensor data at the end of the current time window is selected and combined with the working condition category label, operation status, equipment operating status and working environment status to obtain a JSON object.
3. The method for predicting rescue drilling rig operation parameters based on a large model according to claim 2, characterized in that, After performing instantaneous state semantics on the standardized data, process semantics is then performed on the standardized data to obtain natural language fragments, including: Based on the least squares method, trend features of standardized data are extracted from standardized data. Based on the coefficient of variation, the volatility characteristics of the standardized data are extracted; Standardized data is compared with threshold ranges in an expert knowledge base to obtain threshold determination features; Natural language segments are obtained based on trend features, volatility features, and threshold determination features.
4. The method for predicting rescue drilling rig operation parameters based on a large model according to claim 3, characterized in that, Based on standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions, intelligent drilling operation prompts are generated, including: Intelligent drilling operation prompt word templates are constructed by using preset predictive analysis prompt words, diagnostic and attribution prompt words, optimization suggestion prompt words, and emergency response prompt words; Based on standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions, we extract working condition information, key sensor numerical sequences in standardized data, contextualized information, analogical information from historical cases, and pre-acquired task instructions. Based on the intelligent drilling operation prompt word template, the working condition information, key sensor value sequences in standardized data, contextual information, analogy information from historical cases, and pre-acquired task instructions are input into the large model to generate intelligent drilling operation prompt words.
5. The method for predicting rescue drilling rig operation parameters based on a large model according to claim 4, characterized in that, The predicted values of key parameters within the preset time frame include the mechanical drilling rate ROP_next_60s and the rescue rig torque Torque_next_60s; alarm information includes alarm type, severity, triggering evidence, and corresponding handling suggestions.
6. The method for predicting rescue drilling rig operation parameters based on a large model according to claim 2, characterized in that, Both raw sensor data and sensor data include rescue drilling rig drilling pressure, rescue drilling rig rotation speed, rescue drilling rig lifting force, rescue drilling rig suspended weight, rescue drilling rig power head torque, rescue drilling rig lifting pressure, rescue drilling rig pressing pressure, rescue drilling rig rotation pressure, and well depth; Operating conditions include tripping in and out of the hole, normal drilling, rescue drilling, and emergency stoppage. The JSON object includes a timestamp, work condition type, operation status, and instantaneous values of key parameters; The contextual description includes operator logs and geological environment information.
7. The method for predicting rescue drilling rig operation parameters based on a large model according to claim 6, characterized in that, The raw sensor data is preprocessed to obtain standardized data, including: The raw sensor data undergoes preprocessing including cleaning, smoothing, normalization, and Z-score standardization to obtain standardized data.
8. A rescue drilling rig operation parameter prediction system based on a large model, characterized in that, A method for predicting operational parameters of a rescue drilling rig based on a large model, as described in claim 1, includes: The data acquisition layer module is used to preprocess the raw sensor data to obtain standardized data; The working condition recognition and semantic layer module is used to identify working conditions in standardized data based on preset rule criteria to obtain working condition category labels; to perform instantaneous state semanticization on standardized data to obtain JSON objects; and to perform process semanticization on standardized data to obtain natural language fragments. The large model inference layer module is used to generate intelligent drilling operation prompts based on standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions. The intelligent drilling operation prompts are input into the large model, which uses the large model to establish relationships between standardized data, working condition category labels, JSON objects, natural language fragments, pre-acquired contextual descriptions, pre-acquired historical cases, and pre-acquired task instructions, and performs causal inference. The output includes structured prediction results including working condition category labels, predicted values of key parameters within a preset future time, and alarm information. The result calibration layer module is used to generate query vectors based on pre-acquired contextual descriptions, pre-acquired historical cases, predicted values of key parameters within a preset future time, corresponding operational suggestions, and risk assessments of the current work status. Based on the query vector, similarity matching is performed in the vector database to retrieve several historical cases in the vector database that are higher than a preset relevance threshold; similarity matching is performed between the query vector and several historical cases in the vector database, and weights are assigned to the historical cases in the vector database based on the similarity matching. The predicted values of key parameters within a preset time period are weighted and fused to obtain the calibrated final predicted value. The results feedback layer module is used to generate structured operational suggestions and natural language explanations based on the final calibrated predictions, corresponding operational suggestions, and risk assessments of the current operational status.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Drilling data processing method, device and equipment and readable storage medium
CN116307122A
Well drilling risk processing method and device based on large language model
CN120931065A