Fracturing wellhead device failure analysis method and system based on large model and knowledge graph
By building an intelligent question-answering system based on large models and knowledge graphs, the problems of manual inspection and reliance on expert experience in failure analysis of fracturing wellhead equipment have been solved, achieving fast and accurate fault diagnosis and knowledge acquisition, and improving the intelligence level of the oil and gas industry.
Patent Information
- Application Number
- CN202510865729.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-12
AI Technical Summary
In existing technologies, failure analysis of fracturing wellhead devices relies on manual inspection, on-site monitoring, and expert judgment. This results in long inspection cycles, isolated monitoring data, expert decision-making that relies on experience and is difficult to standardize, limited accuracy and timeliness of fault diagnosis, and a lack of systematic knowledge acquisition.
Using a method based on large models and knowledge graphs, by constructing a knowledge graph, introducing a self-evolution mechanism, LoRA technology and RAG module, combining the Transformer architecture and the Qwen2.5-7B-instruct model, an intelligent question-answering system is built to achieve rapid and accurate failure analysis of fracturing wellhead devices.
It significantly improves the intelligence level of fracturing wellhead device failure analysis, improves the timeliness and accuracy of diagnosis, optimizes knowledge acquisition efficiency, forms a professional data set, and solves the limitations and inaccuracies of traditional methods.
Smart Images

Figure CN120633800A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent diagnosis of oil and gas field development, and in particular to a failure analysis method and system for a fracturing wellhead device based on a large model and a knowledge graph. Background Art
[0002] Fracturing wellheads are critical surface equipment used in oil and gas field production. They are primarily designed to withstand high-pressure operating environments, control wellhead fluid flow, and ensure the safety and stability of downhole fracturing operations. However, they often experience component failures such as wear, erosion, and leakage. This is especially true for cross-connects, gate valves, bolts, and manifolds, which are critical components of the wellhead. These components can suffer irreversible damage due to long-term stress, pressure, and vibration.
[0003] Currently, failure analysis of fracturing wellhead equipment primarily relies on manual inspection, on-site monitoring, and expert judgment. These methods have limitations, such as long inspection cycles, isolated monitoring data, and expert decision-making that relies on experience and is difficult to standardize. This limits the accuracy and timeliness of fault diagnosis. Furthermore, relevant knowledge is scattered across technical manuals, standards, equipment operation and maintenance reports, and patent literature. Traditional knowledge acquisition methods (such as manual review and document retrieval) are not only time-consuming and labor-intensive, but also difficult to form a systematic knowledge system, making them unable to meet the needs of rapid, intelligent decision-making.
[0004] To improve knowledge acquisition efficiency, oil and gas companies urgently need question-answering systems that can quickly and accurately answer professional technical questions. The rapid development of artificial intelligence (AI) technology has provided new solutions for failure analysis of fracturing wellhead equipment. Large models and knowledge graphs, with their powerful reasoning capabilities and knowledge representation, have become important research areas for intelligent oil and gas fields.
[0005] There is an urgent need for a new type of failure analysis method for fracturing wellhead devices that can solve the above problems. Summary of the Invention
[0006] The present invention proposes a failure analysis method and system for fracturing wellhead devices based on a large model and knowledge graph, which solves the problems in the existing technology that failure analysis of fracturing wellhead devices mainly relies on manual inspection, on-site monitoring and expert experience judgment, resulting in long inspection cycles, isolated monitoring data, expert decision-making relying on experience and difficult to standardize, resulting in limited accuracy and timeliness of fault diagnosis.
[0007] The technical solution of the present invention is as follows: A method for analyzing failure of a fracturing wellhead device based on a large model and a knowledge graph, comprising the following steps: Step 1: Build a knowledge graph: (1) Data collection and processing: Collect relevant information in the field of fracturing wellhead through industry standards, failure analysis reports, expert experience documents, corporate cases and scientific research literature to ensure the professionalism and accuracy of the question-answering system; and use NLP technology to perform word segmentation and part-of-speech tagging on the processed fracturing wellhead text data. Basic language analysis; identifying entities such as equipment parts and failure modes from text, such as equipment names, parts, and failure modes; performing knowledge extraction and structuring; standardizing extracted terms and eliminating homonyms and synonyms through entity standardization and disambiguation; identifying relationships between entities, such as causal relationships and dependency relationships, through relationship extraction; and extracting relevant features of entities through attribute extraction. (2) Knowledge fusion: After completing the initial entity and relationship extraction, information from structured data and unstructured data is integrated, and redundant or inconsistent triples in the data sources are unified and integrated into a set of knowledge quadruples with clear structure and unified semantics: ; Where s represents the subject; p represents the relationship; o represents the object; t represents time; E represents the entity set; R represents the relationship set; (3) Knowledge graph: Construct a knowledge graph with fracturing wellhead entities as nodes and fracturing wellhead entity relationships as edges; Step 2: Establish a dynamic update mechanism for the knowledge graph based on the fault evolution process. By introducing information such as timestamps, frequency statistics, and impact path weights, the knowledge graph is given the ability to self-evolve, thereby improving the timeliness and pertinence of diagnosis. Step 3: Build an open-source model: Use the self-attention mechanism in the Transformer architecture to calculate the correlation between samples to determine the importance of each sample in a sequence and assign weights to these samples that represent their importance.
[0008] Among them, Q (query), K (key), and V (value) are three sets of vectors in the attention mechanism; softmax is used to normalize the score, which represents the attention weight of the current token to other tokens; KL divergence loss is used to optimize the generation quality of the model by calculating the difference between the model's predicted distribution and the true answer distribution, making it better fit the target answer and reducing unreasonable distribution deviations. It further measures the deviation between the probability distribution of tokens generated by the model at each time step and the true token:
[0009] in, is the probability distribution of the true answer, is the model’s predicted distribution; Training on the complexity and accuracy of fracturing wellhead device failures ensures the model has strong reasoning capabilities; Step 4: Introduce a closed-loop feedback mechanism to verify the model output and provide feedback results in actual applications. The feedback results are used to optimize the channels of the open source large model through incremental learning, without retraining the entire model. Only when the knowledge changes significantly, small-scale fine-tuning of LoRA is performed to ensure that the model can quickly adapt to new information; and it can continuously optimize model performance and the accuracy of the knowledge base.
[0010] As a further technical solution, step 3 also includes constructing a prompt word project to guide the large model to extract question-answer pairs from the document by designing specific prompt words.
[0011] In a further technical solution, step 3 is fine-tuned using LoRA technology and the Qwen2.5-7B-instruct model. By freezing the original weights of the pre-trained model and training only the newly added low-rank matrix, the number of parameters and computational overhead are greatly reduced. ; in, is the pre-trained weight matrix, It is the adjustment matrix obtained by low-rank decomposition, and only these parameters are updated during training.
[0012] By introducing a low-rank update matrix into the model's weight matrix, domain adaptation or task optimization is achieved, ensuring that the model can quickly adapt to new information. This addresses the high computational cost and storage requirements of traditional fine-tuning methods, which require adjusting all model parameters. The Qwen2.5-7B-instruct model excels in Chinese language understanding and generation, boasting strong bilingual capabilities in both Chinese and English, an interactive style that closely mimics human instructions, and overall performance that outperforms similar models in multiple authoritative evaluations. It can provide a more natural and accurate question-answering and interactive experience for specific domains, such as fracturing wellheads. A further technical solution includes retrieval enhancement after step 3: retrieval enhancement generation is performed through the RAG module. After obtaining the knowledge graph and the fine-tuned large model, the RAG technology is used to enhance the ability of the large model in reasoning, thereby achieving more accurate failure analysis and suggestion generation; in actual applications, the RAG module is channel-optimized through incremental learning based on the model output verification and feedback results, thereby achieving dual-channel optimization of the RAG module and the open source large model; the RAG module dynamically adjusts the retrieval strategy and knowledge base content based on the feedback data, while the large model is only fine-tuned on a small scale through LoRA when the knowledge changes significantly, thereby achieving efficient model adaptation and optimization.
[0013] In a further technical solution, the RAG module is specifically: (a) In the data preprocessing stage, RAG technology uses an embedding model. The bge-large-zh model is optimized for Chinese and has high-precision semantic retrieval capabilities. This model is suitable for RAG tasks and can efficiently match relevant knowledge blocks and improve the accuracy of the question-answering system. (b) In the retrieval phase, FAISS vector retrieval technology is used to perform fast nearest neighbor searches within the vector database. The HNSW indexing strategy is used to improve retrieval efficiency and identify documents most semantically similar to the query. Subsequently, the FAISS search results are preliminarily screened based on the BM25 keyword matching method. Finally, the screened results are re-ranked by calculating vector similarity to ensure that the most relevant knowledge blocks are returned, providing accurate contextual support for the generation phase. The vector index is constructed using the FAISS vector database. (c) In the generation phase, the prompt word engineering constructed in the large model fine-tuning phase is directly applied to the RAG system to guide the model to generate answers; (d) Use BLEU and ROUGE to evaluate the quality of answers generated by large models. The BLEU metric measures the match between the generated answer and the standard answer, focusing on n-gram exact matches. The ROUGE metric measures the recall rate of the generated answer, ensuring that the answer covers key knowledge points and improves the completeness of the answer.
[0014]
[0015] Among them, BP is the penalty factor for shorter sentences; Indicates n-gram matching degree; Indicates the weight of n-gram, usually weight equal( ); Represents the count of matched n-grams, that is, the overlap between the generated text and the reference answer; Represents the count of all n-grams in the reference answer, indicating the complete range of key information.
[0016] The present invention discloses a failure analysis system for a fracturing wellhead device based on a large model and a knowledge graph, comprising an intelligent agent hub, a decision-making and planning module, a data storage module, an execution module, and a tool module; the intelligent agent hub uses fine-tuning and prompting contexts to associate relevant domain knowledge, enabling it to have control, perception, and action capabilities; combining long-term and short-term memory, fine-tuning and prompting contexts are used to inject relevant domain knowledge into the intelligent agent; the decision-making and planning module assists in reasoning, analysis, and decision-making through RAG technology, knowledge graphs, external data, and tool modules; and the execution module implements support for the execution of business tasks. The capabilities of each intelligent agent are applied to specific businesses in the form of intelligent question-answering; intelligent business application services based on intelligent question-answering are constructed, with intelligent question-answering oriented to the intelligent agent as the back-end service form, and the large model analysis results are provided to the business front-end, generating knowledge retrieval and other services in the field of fracturing wellhead failure for users.
[0017] Further technical solutions also include a multi-strategy scheduling module and a tool chain linkage module to further improve the task processing efficiency and accuracy of the intelligent agent center; after task analysis, the intelligent agent center can automatically generate an execution graph according to task requirements, and the decision planning module selects appropriate decision strategies based on historical feedback and tool response quality to ensure personalized and efficient decision-making in different task scenarios; the intelligent agent center supports the end-to-end process of language-structure-execution through the tool chain linkage module; through the structured decomposition of natural language input, the intelligent agent can call knowledge graphs, case libraries, external APIs and computing tools to achieve closed-loop support from reasoning to execution, thereby improving response speed and decision-making accuracy in the event of fracturing wellhead failure.
[0018] Further technical solutions also include user interaction pages, such as web terminals, mobile terminals or API interfaces; Semantic Parsing and Intent Recognition Module: This module understands natural language, identifies user intent, extracts key entities, and determines the query type. It is responsible for parsing user questions and invoking the agent hub to provide answers. If a user asks about the cause of a device failure, the system will invoke the Failure Analysis Module; if the user needs to retrieve a specific standard, the knowledge base will be used for the query. The present invention discloses a method and system for analyzing failure of a fracturing wellhead device based on a large model and a knowledge graph, which has the following beneficial effects: 1. This invention applies large models, knowledge graphs, and RAG technology to fracturing wellhead failure analysis, significantly improving the intelligence level of the oil and gas industry; 2. This invention forms a set of professional data sets. By integrating and fine-tuning domain knowledge, it optimizes the problem that large models have strong generalization but lack professionalism, and improves the accuracy of generated answers; 3. This invention embeds knowledge graphs and RAG technology into the large model, significantly optimizing the large model's expertise in fracturing wellhead failure analysis, avoiding the hallucination problem that may exist in general large models, and ensuring the accuracy of answers; 4. This invention expands the application of large models, knowledge graphs and RAGs in vertical field question-answering systems. Aiming at the complexity of fracturing wellhead failure analysis, it constructs a professional knowledge-enhanced reasoning framework and verifies the application of multiple technologies in industry through actual cases. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 , a schematic diagram of the combined framework of the present invention; Figure 2 , a schematic diagram of constructing the knowledge graph provided by the present invention; Figure 3 , Application of this invention in vertical field question answering system. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] A knowledge graph is a structured knowledge representation method that clearly displays complex relationships between entities. In the field of fracturing wellhead failure analysis, a knowledge graph can structure and organize dispersed knowledge about equipment failures, establishing connections between equipment components, failure modes, cause analysis, and remediation measures, providing explainable reasoning support for intelligent question answering.
[0023] Big model technology refers to large-scale neural network models trained using deep learning algorithms, which possess powerful feature extraction and inference capabilities. In the field of fracturing wellhead failure analysis, big models can efficiently process massive amounts of failure cases, equipment operation and maintenance documentation, and engineering standards, automatically extract key information, and perform intelligent inference based on specific operating conditions, providing efficient and accurate failure diagnosis and decision support. While big models perform well in open-domain question answering, their accuracy and professional application in specialized fields such as fracturing wellheads face challenges.
[0024] Single knowledge graphs or large model technologies still have limitations in fracturing wellhead device failure analysis. Although knowledge graphs have structured expression capabilities, they lack intelligent reasoning capabilities; while large models have strong reasoning and language interaction capabilities, they lack a deep understanding of the professional knowledge of fracturing wellhead devices. In addition, to compensate for the deviations or hallucinations that occur in the reasoning process of large models, which affect the accuracy and relevance of professional knowledge answers, RAG retrieval-enhanced generation technology has emerged. RAG technology combines generative models with efficient information retrieval mechanisms, using the retrieval system to extract relevant document fragments from external knowledge bases and generate more accurate and context-relevant answers through large model processing. Therefore, combining large models, knowledge graphs, and RAG technology to construct an intelligent question-answering system has important research significance and application value.
[0025] Figure 2 The present invention provides a schematic diagram for constructing a knowledge graph and generates a professional data set driven by the knowledge graph. The graph optimizes the problem of large models having strong generalization but insufficient professionalism by integrating and structuring the domain knowledge of fracturing wellhead devices, thereby improving the accuracy of generated answers.
[0026] Figure 3 , the application of this invention in vertical field question-answering systems; in view of the complexity of fracturing wellhead failure analysis, a professional knowledge-enhanced reasoning framework was constructed, and the application of multiple technologies in industry was verified through actual cases.
[0027] The application of big models, knowledge graphs and RAG technology to fracturing wellhead failure analysis has significantly improved the intelligence level of the oil and gas industry.
[0028] like Figure 1 As shown in the combined framework diagram of the present invention, a knowledge base is constructed to process and store documents such as the acquired knowledge on the fracturing wellhead field, including: parsing and cleaning data related to the fracturing wellhead; storing the processed fracturing wellhead data in a database; organizing and summarizing the fracturing wellhead knowledge to form a fracturing wellhead failure analysis database; the data types required to be stored include structured data and unstructured data. Structured data such as production data, equipment parameters, and industry standards are stored in the database; industry standard texts, failure analysis reports, expert experience documents, corporate cases, and scientific research literature related to fracturing wellheads are vectorized and factor extracted, and then placed into knowledge graphs and vector databases according to business needs to form a fracturing wellhead failure knowledge base.
[0029] A prompt engineering project was built to design specific prompts to guide the large model in extracting question-answer pairs from documents. This included: configuring the key and URL for the Alibaba Cloud Bailian API to lay the foundation for subsequent document extraction and ensure the large model could efficiently access and process fracturing wellhead data; segmenting collected long text data documents into small paragraphs to improve the performance and accuracy of the large model or RAG system; cleaning the text content to remove extra spaces and non-ASCII characters; building a prompt engineering project and calling the Alibaba Cloud Bailian API to generate professional question-answer pairs and directly return JSON data that conforms to the format for subsequent large model fine-tuning. The extracted content must focus on fracturing wellhead failure analysis and criteria for determining fracturing wellhead failure. The format of the question-answer pair is: "instruction", which describes the specific requirements of the task; "input", which provides the input content for the model; and "output", which contains the generated content of the model.
[0030] Construct a knowledge graph and extract key information from documents related to fracturing wellheads through natural language processing technology, including: using natural language processing technology to segment and tag the processed fracturing wellhead text data. Break down long texts into individual words, and part-of-speech tagging will mark each word with its grammatical role, such as noun, verb, etc. Knowledge extraction and structuring. Use named entity recognition to identify entities such as equipment parts and failure modes from the text, such as equipment names (such as wellhead, casing head, fracturing valve), components (such as sealing rings, connecting bolts), and failure modes (such as corrosion, fatigue, and wear); standardize and disambiguate entities, standardize the extracted terms (such as "high pressure" and "high pressure working conditions" are normalized as the same entity), and eliminate homonyms or synonyms; identify the relationships between entities through relationship extraction, such as causal relationships ("high pressure leads to sealing failure"), dependency relationships ("wellhead relies on casing head structural support"), etc.; attribute extraction is used to extract relevant features of entities (occurrence frequency, operating parameters, etc.); To improve the timeliness and evolutionary capabilities of the knowledge graph, this paper further proposes a dynamic knowledge graph update mechanism tailored to the fault evolution process. This mechanism first extracts events and performs time series analysis on failure documents related to fracturing wellhead equipment. Based on the key operational nodes, equipment status changes, and fault occurrence times described in the documents, a fault evolution time series is generated. Each edge, or fault relationship, can be assigned a timestamp or time interval, indicating the moment or time range in which the fault relationship occurred.
[0031] Through knowledge fusion, scattered multi-source data information is transformed into structured quadruple data: ; Where s represents the subject; p represents the relationship; o represents the object; t represents time; E represents the entity set; R represents the relationship set; A framework for event evolution modeling was proposed, and a knowledge graph of fracturing wellhead failures was constructed using Neo4j. Each node represents a specific failure event, while edges represent temporal sequences or causal logic. Based on this, the system introduced various semantic tags, such as "node timestamp" (reflecting the time of event occurrence), "event frequency weight" (measuring the probability of the event occurring in historical cases), and "path evolution credibility" (path occurrence probability learned based on historical data), giving the graph a richer evolutionary dimension.
[0032] Finally, the reasoning process is based on the entities and relationships stored in the knowledge graph, using graph database queries and rule engines for intelligent inference. The inference engine traverses the nodes (entities) and edges (relationships) in the graph, analyzing causal and dependency relationships between entities to discover underlying knowledge.
[0033] Furthermore, the system supports incremental, dynamic updates to the graph. When new cases or feedback data arrives, the system automatically revises the graph structure if it identifies new event chains or significant changes in the frequency of existing paths. This update mechanism includes adding new nodes, updating edge weights, or replacing paths. Through this incremental update, the system continuously optimizes fault evolution paths, providing more accurate intelligent question-and-answer support.
[0034] Fine-tuning was performed using LoRA technology and the Qwen2.5-7B-instruct large model, including: The Qwen2.5-7B-instruct model is an open source large model that has achieved significant improvements in instruction execution, generating long texts, understanding structured data, and generating structured outputs, especially JSON. The core architecture of the large model is the transformer architecture. The transformer is mainly composed of an encoder and a decoder. The Qwen2.5-7B-instruct is a decoder model. At each step, the decoder predicts the next token based on the input prompt and the previously generated content. Each step of the model relies on the results generated previously, thereby gradually reasoning and generating text; the core components of the Transformer architecture include self-attention mechanism (Self-Attention), positional encoding (Positional Encoding) and multi-head attention mechanism (Multi-Head Attention); the self-attention mechanism determines the importance of each sample to a sequence by calculating the correlation between samples, and assigns these samples weights that represent their importance:
[0035] Among them, Q (query), K (key), and V (value) are three sets of vectors in the attention mechanism; softmax is used to normalize the score, which represents the attention weight of the current token to other tokens; Positional encoding is mainly used to capture the positional relationship between words so that the model can understand the order and structure between words. The positional encoding in Transformer mainly represents the position of each word through a set of fixed vectors generated by sine and cosine functions:
[0036]
[0037] Where pos represents the position of the input sequence; d represents the total dimension of the position encoding vector; The introduction of the multi-head attention mechanism allows the model to calculate multiple attention heads in parallel. Each head can learn and understand the input data from a different perspective, thereby enhancing the model's ability to capture different dimensions of information in the text. For each attention head h: , , ; in , , is the learning parameter matrix for head h, X is the input vector; Calculate the attention weights through the self-attention mechanism described above; Finally, the outputs of multiple heads are merged, and each attention head obtains a weighted output vector by calculation: ; in, represents the output of the h-th head, Represents the weight matrix of the output layer, which is used to merge the outputs of all heads and map them to the final representation space.
[0038] LoRA is an efficient parameter fine-tuning technology for large pre-trained models. It achieves domain adaptation or task optimization by introducing a low-rank update matrix into the model's weight matrix. LoRA freezes the original weights of the pre-trained model and only trains the newly added low-rank matrix, significantly reducing the number of parameters and computational overhead.
[0039]
[0040] in, is the pre-trained weight matrix, is the adjusted matrix obtained through low-rank decomposition, and only this part of the parameters will be updated during training.
[0041] Specifically, for the weight matrix W (with dimensions d×k) in the pre-trained model, LoRA introduces a low-rank decomposition form ∆W = A·B, where A (with dimensions d×r) and B (with dimensions r×k) are two smaller trainable matrices, and r is the rank (r << min(d,k)), thus restricting parameter updates to a low-dimensional subspace. During training, only A and B are optimized, while W remains unchanged. This method not only reduces the VRAM usage and training time but also facilitates switching between multiple tasks, as only the A and B matrices for different tasks need to be stored.
[0042] The present invention selects Lumma Factory to perform visual fine-tuning on Qwen2.5-7B-instruct. LummaFactory is a large model fine-tuning tool that provides a visual low-code interface, which can simplify the model training process and support efficient fine-tuning techniques such as LoRA.
[0043] Load the processed JSON data related to the failure of the fracturing wellhead and the base model into Lumma Factory, and enable LoRA for efficient fine-tuning. Lumma Factory will automatically load the LoRA adaptation layer and freeze the original weights of Qwen2.5-7B-instruct, only optimizing the low-rank matrix, thereby reducing VRAM usage and accelerating training.
[0044] Visualize and monitor the fine-tuning process, observe the decline of the loss function during training, and dynamically adjust parameters such as the learning rate to optimize the training results.
[0045] Combined with the RAG technology for retrieval-augmented generation to enhance the capabilities of the large model during inference and achieve more accurate failure analysis and recommendation generation, including: Text vector representation. Use the bge-large-zh text embedding model to convert the fracturing wellhead text data into a high-dimensional vector representation: ; Among them, is the text data input by the user, is the high-dimensional vector representation of the text; For each piece of text in the knowledge base, pre-compute their vector representations: ; In this way, all the documents are converted into vectors and stored in the FAISS vector database; Knowledge search. When the user inputs a query the system finds the most relevant document in the knowledge base; The FAISS vector database + HNSW (Hierarchical Navigable Small World) index is used for efficient nearest neighbor search.
[0046] Calculate query text vector With knowledge base vector The cosine similarity between: ; Select the top-K documents with the highest similarity as candidate knowledge: ; Re-ranking to ensure the most relevant content comes first. The BM25 algorithm is used to calculate the keyword matching degree of the text. The higher the BM25 calculation score, the higher the text matching degree: ; ; in, Used to measure the importance of words; Expressive words In the documentation Frequency of occurrence in Represents a document length; represents the average length of all documents in the knowledge base; and are free parameters that adjust the effects of word frequency and document length, respectively; Combining the similarity of vector retrieval and the keyword matching of BM25, the final score is weighted and calculated: ; Select the highest-ranked document as contextual support information for the large model to generate answers; Generate the final answer: Based on the retrieved knowledge, use Qwen2.5-7B-instruct to generate the answer: ; BLEU and ROUGE are used to evaluate the quality of answers generated by large models. The BLEU metric measures the match between the generated answer and the ground truth answer, focusing on exact n-gram matches. The ROUGE metric measures the recall of the generated answer, ensuring that the answer covers key knowledge points and improves the completeness of the answer.
[0047]
[0048]
[0049] Among them, BP is the penalty factor for shorter sentences; Indicates n-gram matching degree; Indicates the weight of n-gram, usually weight equal( ); Represents the count of matched n-grams, that is, the overlap between the generated text and the reference answer; Represents the count of all n-grams in the reference answer, indicating the complete range of key information.
[0050] A closed-loop feedback mechanism is introduced. Through a dual-channel incremental mechanism, the RAG module dynamically adjusts the retrieval strategy and knowledge base content based on feedback data, while the large model only performs small-scale incremental fine-tuning through LoRA when knowledge changes significantly, thereby achieving efficient model adaptation and optimization. In the closed-loop feedback mechanism, the RAG module is first optimized through incremental learning. As new data is introduced, the system dynamically adjusts the knowledge base content and retrieval strategy to ensure it can adapt to new knowledge or failure modes. If new knowledge is added, such as new failure cases or industry reports, embeddings are calculated for the new data, and bge-large-zh is used to calculate the vector representation of the new data:
[0051] in is the new text data, is the generated vector.
[0052] Then the new vector is directly added to the FAISS vector library to ensure that FAISS can still perform the nearest neighbor search (HNSW) efficiently, avoiding the need to delete the library and rebuild it, thus saving computational effort.
[0053] When FAISS incrementally updates new knowledge, RAG automatically uses the new knowledge during the retrieval phase without the need for manual adjustment.
[0054] When the feedback data contains significant new knowledge, the system will initiate a fine-tuning process for the large model. Using LoRA (Low Rank Adaptation) technology, the system only adjusts some parameters of the large model without retraining the entire model.
[0055] Fine-tuning the large model works in tandem with the incremental adaptation mechanism of the RAG module to improve the system's responsiveness and accuracy. The RAG module is responsible for rapidly capturing and retrieving new knowledge, while the large model enhances its reasoning capabilities through fine-tuning. Together, these two ensure the system can continuously optimize in a dynamically changing environment, providing high-quality intelligent decision support.
[0056] Construct an intelligent agent framework for fracturing wellhead failure analysis, including an agent hub, decision-making and planning modules, data storage modules, execution modules, and tool modules. The agent framework enables the invocation of corresponding modules, such as analyzing failure reports, retrieving failure cases, and proposing preventive measures. The fine-tuned large model is selected to build the intelligent agent hub, giving it control, perception, and action capabilities. After receiving natural language tasks, the intelligent agent hub can identify the intent and analyze the structure of the task, and dynamically plan the call process through the decision module.
[0057] The system incorporates a mechanism for fusing short-term and long-term memory: long-term memory includes a knowledge graph, a fault case library, and a predefined rule set, while short-term memory stores context and task variables generated during the current interaction. A prompt injection mechanism is used to introduce domain knowledge into the model inference process.
[0058] RAG technology, external data, and external tools are used to assist in reasoning, analysis, and decision-making, while also providing support for the execution of business tasks on the back end.
[0059] A multi-strategy scheduling mechanism is built to dynamically match different reasoning paths and tool call sequences based on task type (such as retrieval, reasoning, and generation). After parsing structured tasks, the agent automatically generates an execution graph and selects a strategy based on task content, historical feedback, and tool response quality, enabling personalized and efficient decision-making.
[0060] In terms of the tool chain linkage mechanism, the system supports an end-to-end linkage process of language-structure-execution: the model can complete structured decomposition (such as extracting fault events, time series, parameter conditions) based on the user's natural language input, and call the knowledge graph engine, case library retrieval module, external API interface or professional computing tools accordingly to complete the closed-loop execution process from reasoning to action.
[0061] We further introduce policy evolution capabilities, namely, continuously optimizing the scheduling mechanism through a policy meta-learning framework. Based on user interaction feedback data, the system can perform reinforcement learning-style updates on task scheduling paths, gradually improving task matching accuracy and tool call efficiency, demonstrating the agent's adaptive evolutionary capabilities across multiple rounds of business tasks.
[0062] Build intelligent business application services based on intelligent question-answering, use intelligent question-answering for intelligent agents as the back-end service, provide large model analysis results to the business front-end, and generate knowledge retrieval and other services in the field of fracturing wellhead failure for users.
[0063] The knowledge retrieval process in the field of fracturing wellhead failure includes: User input issues: You need to design a user interaction page, such as a web terminal, mobile terminal, or API interface; Semantic parsing and intent recognition: This primarily involves building an intelligent question-and-answer backend, responsible for parsing user questions and invoking intelligent agents to provide responses. The backend must possess natural language understanding capabilities, be able to identify user intent, extract key entities, and determine the query type. For example, if a user inquires about the cause of a device failure, the system will invoke the failure analysis module; if the user needs to retrieve a specific standard, the knowledge base will be used for the query. Knowledge retrieval: including structured knowledge retrieval generated by prompt word engineering, knowledge graph retrieval and RAG retrieval.
[0064] Of course, without departing from the spirit and essence of the present invention, technicians familiar with the field should be able to make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.
Claims
1. A fracturing wellhead device failure analysis method based on a large model and knowledge graph, characterized by: The following steps are involved: Step 1: Build a knowledge graph: (1) Data collection and processing: Collect relevant information in the fracturing wellhead field through industry standards, failure analysis reports, expert experience documents, corporate cases and scientific research literature; and use NLP technology to perform basic language analysis such as word segmentation and part-of-speech tagging on the processed fracturing wellhead text data; Identify equipment components, failure mode entities, components, and failure modes from text; perform knowledge extraction and structuring; Through entity standardization and disambiguation processing, the extracted terms are standardized and the homonymy or synonymy problems are eliminated; Identify the relationships between entities through relationship extraction; Attribute extraction is used to extract relevant features of entities; (2) Knowledge fusion: After completing the initial entity and relationship extraction, information from structured data and unstructured data is integrated, and redundant or inconsistent triples in the data sources are unified and integrated into a set of knowledge quadruples with clear structure and unified semantics: ; Where s represents the subject; p represents the relationship; o represents the object; t represents time; E represents the entity set; R represents the relationship set; (3) Knowledge graph: Construct a knowledge graph with fracturing wellhead entities as nodes and fracturing wellhead entity relationships as edges; Step 2: Establish a dynamic update mechanism for the knowledge graph based on the fault evolution process. By introducing information such as timestamps, frequency statistics, and impact path weights, the knowledge graph is given the ability to self-evolve. Step 3: Build an open source large model: Calculate the correlation between samples through the self-attention mechanism in the Transformer architecture: Among them, Q (query), K (key), and V (value) are three sets of vectors in the attention mechanism; softmax is used to normalize the score, which represents the attention weight of the current token to other tokens; Using KL divergence loss, by calculating the difference between the model's predicted distribution and the true answer distribution, the model's generation quality is optimized to better fit the target answer and reduce unreasonable distribution deviations: in, is the probability distribution of the true answer, is the model’s predicted distribution; Step 4: Introduce a closed-loop feedback mechanism to verify the model output and provide feedback results in actual applications. The feedback results are used to optimize the channel of the open source large model through incremental learning.
2. The method for analyzing failure of a fracturing wellhead device based on a large model and a knowledge graph according to claim 1, characterized in that: The step 3 also includes constructing a prompt word project, by designing specific prompt words to guide the large model to extract question-answer pairs from the document.
3. The method for analyzing failure of a fracturing wellhead device based on a large model and a knowledge graph according to claim 2, characterized in that: Step 3 is fine-tuned using LoRA technology and the Qwen2.5-7B-instruct model. By freezing the original weights of the pre-trained model, only the newly added low-rank matrix is trained, which greatly reduces the number of parameters and computational overhead; ; in, is the pre-trained weight matrix, It is the adjustment matrix obtained by low-rank decomposition, and only these parameters are updated during training.
4. The method for analyzing failure of a fracturing wellhead device based on a large model and a knowledge graph according to claim 3, characterized in that: Step 3 also includes retrieval enhancement: retrieval-augmented generation is performed through the RAG module (Retrieval-Augmented Generation). After obtaining the knowledge graph and the fine-tuned large model, the RAG technology is used to enhance the large model's ability in reasoning, thereby achieving more accurate failure analysis and suggestion generation. In actual applications, the RAG module is optimized through incremental learning based on the model output verification and feedback results.
5. The method for analyzing failure of a fracturing wellhead device based on a large model and a knowledge graph according to claim 4, characterized in that: The RAG module is specifically: (a) In the data preprocessing stage, RAG technology uses the Embedding model, and the Embedding model used is the bge-large-zh model; (b) During the retrieval phase, the FAISS vector retrieval technique is used to perform a fast nearest neighbor search within the vector database. The HNSW indexing strategy is used to improve retrieval efficiency and identify documents most similar to the query semantics. Subsequently, the FAISS search results are preliminarily screened based on the BM25 keyword matching method. Finally, the screened results are re-ranked by calculating vector similarity. The vector index is constructed using the FAISS vector database. (c) In the generation phase, the prompt word engineering constructed in the large model fine-tuning phase is directly applied to the RAG system to guide the model to generate answers; (d) Use BLEU and ROUGE to evaluate the quality of the answers generated by the large model; the BLEU metric is used to measure the match between the generated answer and the standard answer; The ROUGE metric is used to measure the recall of the generated answers; Among them, BP is the penalty factor for shorter sentences; Indicates n-gram matching degree; Indicates the weight of n-gram, usually weight equal( ); Represents the count of matched n-grams, that is, the overlap between the generated text and the reference answer; Represents the count of all n-grams in the reference answer, indicating the complete range of key information.
6. A fracturing wellhead device failure analysis system based on a large model and knowledge graph, characterized by: It includes intelligent body center, decision-making and planning module, data storage module, execution module and tool module; The agent hub associates relevant domain knowledge using fine-tuning and prompting contexts; The decision-making planning module assists in reasoning, analysis and decision-making through RAG technology, knowledge graph, external data and tool modules; and implements execution support for business tasks through the execution module.
7. The fracturing wellhead device failure analysis system based on a large model and a knowledge graph according to claim 6, characterized in that: It also includes a multi-strategy scheduling module and a tool chain linkage module. After task analysis, the intelligent agent hub can automatically generate an execution graph based on task requirements. The decision planning module selects appropriate decision strategies based on historical feedback and tool response quality to ensure personalized and efficient decision-making in different task scenarios. The intelligent agent hub supports the end-to-end process of language-structure-execution through the tool chain linkage module. By structured decomposition of natural language input, the intelligent agent can call upon knowledge graphs, case libraries, external APIs and computing tools.
8. The fracturing wellhead device failure analysis system based on a large model and a knowledge graph according to claim 6 or 7, characterized in that: It also includes user interaction pages; Semantic parsing and intent recognition module: understands natural language, identifies user intent, extracts key entities, and determines the query type. It is responsible for parsing user questions and calling the intelligent agent hub to answer them.
Citation Information
Cited By
Model cluster-based complex system intelligent design method and system
CN121189188A
A model cluster-based intelligent design method and system for complex systems
CN121189188B
Drilling operation assistance and fault analysis method and system based on multi-agent cooperation
CN121561744A