Petroleum knowledge question-answering method and system based on large language model
By preprocessing petroleum data and fine-tuning a large language model, the bias problem in existing petroleum knowledge question answering technologies has been solved, achieving more efficient and accurate petroleum knowledge question answering.
Patent Information
- Application Number
- CN202511378527.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing model-based knowledge question answering technologies fail to reach their full potential when faced with specific tasks, resulting in significant biases in petroleum knowledge question answering.
The platform collects raw petroleum-related data from various data sources, performs usability preprocessing and verification, constructs a set of petroleum-related question-and-answer pairs, fine-tunes the large language model, evaluates the effectiveness of question-and-answer output, and iteratively optimizes the question-and-answer process.
It improves the accuracy and efficiency of petroleum knowledge Q&A, ensures the reliability and quality of data, enhances the ability of large language models to understand petroleum terminology and complex processes, and accurately matches user needs.
Smart Images

Figure CN120893584A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computing models, in particular to a petroleum knowledge question and answer method and system based on a large language model. BACKGROUND
[0002] A large language model (LLM) is a language model composed of a neural network with many parameters (usually billions of weights or more), which is trained on a large amount of unlabelled text using self-supervised learning or semi-supervised learning. With the continuous innovation of technology, such as deep sea exploration technology, shale oil exploitation technology, and the emergence of emerging fields, a large amount of new knowledge continues to flow in, while traditional petroleum knowledge has been accumulated for many years and has formed a deep deposit, which makes petroleum practitioners face great challenges in knowledge acquisition. Whether it is for new employees to quickly get familiar with the business, or for experienced experts to track the latest developments and solve cross-disciplinary problems, efficient and accurate knowledge support is needed.
[0003] For example, the invention patent with publication number CN111881279B discloses a question and answer method, a question and answer device and a storage device based on a Transformer model, the question and answer method comprising: obtaining a question text input by a user, processing the question text to obtain a question sequence; decoding the question sequence to obtain a plurality of candidate answers related to the question sequence; concatenating the question sequence with each candidate answer; scoring each concatenation result to select the candidate answer corresponding to the highest score as the optimal answer for the question sequence.
[0004] For example, the invention patent with publication number CN114416953B discloses a question and answer processing method, a training method, device, equipment, medium and product of a question and answer model, relating to the technical field of artificial intelligence, specifically to the technical field of natural language processing, deep learning and knowledge graph, the question and answer processing method comprising: obtaining processing data, wherein the processing data includes question data and candidate answers; performing general semantic understanding on the processing data to obtain general data features; selecting a target question and answer processing mode from candidate question and answer processing modes based on the general data features; processing the general data features using the target question and answer processing mode to obtain a target answer in the candidate answers for the question data.
[0005] In combination with the above technical solutions, it is found that the existing model-based knowledge question and answer technical solutions generally have general applicability, so that they may not be able to perform well when facing specific tasks, resulting in large deviations in knowledge question and answer, and ultimately affecting the use of petroleum knowledge question and answer. SUMMARY
[0006] In view of the deficiencies of the prior art, the present application provides a petroleum knowledge question and answer method and system based on a large language model, which can effectively solve the problems involved in the above background art.
[0007] To achieve the above object, the present application is implemented by the following technical solutions: The present application provides a petroleum knowledge question and answer method based on a large language model, comprising: data availability preprocessing: a petroleum knowledge question and answer platform collects original petroleum domain data from each data source, denoted as an original petroleum domain data set, a data processing component performs availability preprocessing operation on the original petroleum domain data set, determines the availability evaluation value of the original petroleum domain data set, checks the availability evaluation threshold value predefined, determines the availability state of the original petroleum domain data set, and records the original petroleum domain data set after the availability preprocessing operation as a petroleum domain available data set; question and answer pair application quality determination: the petroleum knowledge question and answer platform constructs a petroleum domain question and answer pair set according to a predefined construction method, and performs calibration operation, determines the application quality index of each petroleum domain question and answer pair, compares with the predefined application quality reference index, and screens to obtain a petroleum domain question and answer pair data set; model output effect determination: the petroleum knowledge question and answer platform stores the petroleum domain question and answer pair data set in a large language model to which a petroleum question and answer component belongs, the petroleum knowledge question and answer platform issues a fine-tuning instruction, and the petroleum question and answer component performs fine-tuning operation on the large language model according to a predefined fine-tuning method, evaluates the question and answer output effect index of the large language model, checks the predefined question and answer output effect index, and the petroleum knowledge question and answer platform determines whether to perform repeated fine-tuning operation on the large language model, thereby completing the petroleum knowledge question and answer based on the large language model.
[0008] The second aspect of the present application provides a large language model-based oil knowledge Q&A system, comprising: a data availability preprocessing module for an oil knowledge Q&A platform to collect oil field raw data of each data source, denoted as an oil field raw data set, a data processing component to perform availability preprocessing operation on the oil field raw data set, to determine the availability evaluation value of the oil field raw data set, to check with a predefined availability evaluation threshold, to determine the availability state of the oil field raw data set, and to denote the oil field raw data set after availability preprocessing operation as an oil field available data set; a Q&A pair application quality module for the oil knowledge Q&A platform to construct an oil field Q&A pair set from the oil field available data set according to a predefined construction method, and to perform calibration operation to determine the application quality index of each oil field Q&A pair, and to compare with a predefined application quality reference index to obtain an oil field Q&A pair data set; a model output effect module for the oil knowledge Q&A platform to store the oil field Q&A pair data set in a large language model to which the oil Q&A component belongs, the oil knowledge Q&A platform to issue a fine-tuning instruction, and the oil Q&A component to perform fine-tuning operation on the large language model according to a predefined fine-tuning method, to evaluate the Q&A output effect index of the large language model, to check with a predefined Q&A output effect index, and to determine whether to perform repeated fine-tuning operation on the large language model, so as to complete the large language model-based oil knowledge Q&A.
[0009] Compared with the prior art, the embodiments of the present application have at least the following advantages or beneficial effects: (1) The present application provides a large language model-based oil knowledge Q&A method and system. The oil knowledge Q&A platform first collects oil field raw data from each data source, and obtains an available data set after availability preprocessing and checking by a data processing component; then constructs a Q&A pair set according to a predefined method based on the data set, and filters out a Q&A pair data set after calibration and comparison, to provide more accurate information for the oil Q&A component and improve the efficiency and accuracy of the Q&A component; then stores it in the large language model of the oil Q&A component, the platform issues a fine-tuning instruction, the oil Q&A component fine-tunes the model according to the predefined method, and evaluates the Q&A output effect index of the model and checks it with the predefined index, to determine whether to repeat the fine-tuning, so as to realize the large language model-based oil knowledge Q&A.
[0010] (2) The present application performs availability preprocessing operation on the oil field raw data set by a data processing component, determines the data preprocessing performance index of the data processing component, can understand the resource requirements of the data processing component at different stages, optimizes the data processing process to improve the overall data processing speed, and determines the stability of the data processing component according to the data preprocessing performance index, which helps to ensure the data processing quality.
[0011] (3) The present application evaluates the availability of the original data set in the oil field by determining the availability evaluation value, and corrects the errors and missing problems of the data in the oil field caused by various factors through availability preprocessing, so that the data is more accurate and reliable, and can provide higher quality data basis for subsequent knowledge question and answer work.
[0012] (4) The present application fine-tunes the large language model according to the pre-defined fine-tuning mode through the oil question and answer component, and through the fine-tuning operation, the oil question and answer component can better understand the oil professional terms, concepts and complex process flow, the fine-tuned component can more accurately combine the oil geology and mining knowledge to answer, reduce ambiguous or incorrect expressions, so as to accurately match the user's oil knowledge demand, evaluate the question and answer output effect index of the large language model, optimize the scientificity of the answer, and enhance the relevance and pertinence of the component. BRIEF DESCRIPTION OF DRAWINGS
[0013] The present application is further illustrated by the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present application. For ordinary skilled in the art, other drawings can be obtained without creative labor on the basis of the following drawings.
[0014] Figure 1 The present application is further illustrated by the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present application. For ordinary skilled in the art, other drawings can be obtained without creative labor on the basis of the following drawings.
[0015] Figure 2 The present application is further illustrated by the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present application. For ordinary skilled in the art, other drawings can be obtained without creative labor on the basis of the following drawings.
[0016] Figure 3 The present application is further illustrated by the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present application. For ordinary skilled in the art, other drawings can be obtained without creative labor on the basis of the following drawings. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary skilled in the art without creative labor are within the scope of protection of the present application.
[0018] Referring to Figure 1 The first aspect of the present application provides a large language model-based oil knowledge question and answer method, comprising: an oil knowledge question and answer platform collects original data in the oil field from various data sources, denoted as an original data set in the oil field, a data processing component performs availability preprocessing operation on the original data set in the oil field, determines the availability evaluation value of the original data set in the oil field, checks with the pre-defined availability evaluation threshold, determines the availability state of the original data set in the oil field, and the original data set in the oil field after availability preprocessing operation is denoted as an available data set in the oil field.
[0019] The above-mentioned oil knowledge Q&A platform is a digital platform focusing on knowledge exchange and answer in the field of oil, aiming to provide professional, accurate and convenient knowledge service for oil industry practitioners, researchers, students and people interested in oil knowledge. It relies on advanced information technology and rich oil field resources to realize the function of user asking and obtaining answers. Users can ask various oil-related questions on the platform, such as oil exploration technology, mining technology, refining process, market dynamics, industry policy, etc. The oil knowledge Q&A platform uses natural language processing technology to understand the user's problem intention, analyzes, semantically understands and classifies the natural language input by the user, and converts it into a form that the computer can understand, so as to accurately match the relevant knowledge and answers, and with the help of the powerful language understanding and generation ability of the large language model, the oil field knowledge is learned and inferred to generate high-quality answers. The data processing component is a key part of the oil knowledge Q&A platform, which is a tool for data processing in the oil knowledge Q&A platform, used to process, convert and optimize the collected oil field raw data to improve the usability and quality of the data.
[0020] The oil field raw data of each data source can be oil enterprise internal data such as production data, exploration data, management data, etc.; industry organization published data such as industry association data, etc.; academic research data such as academic journal paper data, research reports of colleges and research institutions, etc.; market data such as oil transaction data, market research agency data, etc.
[0021] The above-mentioned availability preprocessing operation specifically includes missing value processing: methods such as mean filling and median filling can be used for data perfection; error value correction: the data processing component can compare with historical data, associate with data of other related equipment, or make judgments and corrections according to physical principles, for example, in oil production, if the daily oil production data of an oil well suddenly fluctuates significantly and is inconsistent with the stable production trend in the past few months or years, further judgment is needed, for example, the daily oil production in a period of time is 100, 105, 101, 150, 106, 98, and the stable production interval in the past few months is [95, 110], then 150 indicates that the daily oil production data has a significant fluctuation, which may need to be corrected. Specifically, the slope, fluctuation range and other characteristics of the historical production curve can be compared to determine the anomaly; repeated value deletion: repeated records are identified and deleted by comparing key data fields such as oil product transaction time; data format unification: for example, date formats may have "year-month-day", "day / month / year" and other forms, which need to be unified into a standard format during preprocessing.
[0022] Specifically, the data processing component performs availability preprocessing operations on the original data set in the oil field, obtains preprocessing performance data of the data processing component, and obtains availability preprocessing operation data of the original data set in the oil field.
[0023] The preprocessing performance data of the data processing component specifically includes CPU occupancy of the data processing component at each preprocessing time point, memory occupancy of the data processing component at each preprocessing time point, preprocessing duration of the data processing component in a preprocessing period, and data processing amount of the data processing component in the preprocessing period.
[0024] The preprocessing time point is specifically a plurality of preprocessing time points obtained by dividing the preprocessing period by time points, and the division method can be 30 seconds. The preprocessing period is specifically a duration of the data processing component performing availability preprocessing operations on the original data set in the oil field. The preprocessing period is specifically a period of time from when the data processing component receives an availability preprocessing operation instruction to when the data processing component completes all original data operations and feeds back to the oil knowledge Q&A platform.
[0025] The preprocessing performance data can be extracted from the running report of the data processing component.
[0026] The CPU occupancy of the data processing component at each preprocessing time point is processed by averaging to obtain the CPU average occupancy of the data processing component in the preprocessing period.
[0027] The data processing amount of the data processing component in the preprocessing period is processed by ratio with the preprocessing duration of the data processing component in the preprocessing period to obtain the data processing throughput of the data processing component in the preprocessing period.
[0028] The CPU average occupancy limit value and the memory maximum occupancy reference value are extracted from the oil knowledge Q&A information base.
[0029] The CPU average occupancy of the data processing component in the preprocessing period, the memory occupancy of the data processing component at each preprocessing time point, the preprocessing duration of the data processing component in the preprocessing period, and the data processing throughput of the data processing component in the preprocessing period are comprehensively processed to obtain the data preprocessing performance index of the data processing component. The specific analysis method is as follows: ; In the formula, is the data preprocessing performance index of the data processing component, is the CPU average occupancy of the data processing component in the preprocessing period, is the CPU average occupancy limit value, is the memory occupancy rate of the data processing component at the i-th preprocessing time point, i is the number of each preprocessing time point, M is the total amount of preprocessing time points, max represents the maximum value, is the maximum memory occupancy reference value, is the preprocessing duration of the data processing component in the preprocessing period, is the data processing throughput of the data processing component in the preprocessing period, is the preprocessing performance influence parameter corresponding to the predefined preprocessing duration in the oil knowledge Q&A information base, is the preprocessing performance influence parameter corresponding to the data processing throughput in the oil knowledge Q&A information base, and e is a natural constant.
[0030] The data preprocessing performance index of the data processing component in the embodiment is used to measure the efficiency and effect of the data processing component in the data preprocessing stage, such as data cleaning and transformation, and is also used to evaluate the ability and stability of the data preprocessing component in processing data.
[0031] The preprocessing performance influence parameter corresponding to the preprocessing duration and the preprocessing performance influence parameter corresponding to the data processing throughput are both extracted from the oil knowledge Q&A information base. For example, the preprocessing duration and the data processing throughput form a mapping set with the preprocessing performance influence parameter corresponding to the preprocessing duration and the preprocessing performance influence parameter corresponding to the data processing throughput which are preset in the oil knowledge Q&A information base, and the real-time preprocessing duration and the data processing throughput are brought into the mapping set to obtain the preprocessing performance influence parameter corresponding to the preprocessing duration and the preprocessing performance influence parameter corresponding to the data processing throughput.
[0032] In the embodiment, if the CPU average occupancy rate is small within a suitable range, the preprocessing duration will be shortened, and the data preprocessing performance of the data processing component will be increased. When the CPU average occupancy rate is too high and greater than the CPU defined occupancy rate, the system performance may be reduced, and the phenomenon of lag may occur, which will prolong the preprocessing duration and have a negative impact on the data preprocessing performance of the data processing component. Similarly, appropriate memory occupancy can improve the preprocessing speed and shorten the preprocessing duration. However, if the maximum memory occupancy is too high and greater than the maximum memory occupancy reference value, the system will frequently perform memory exchange, which will greatly reduce the data processing speed, prolong the preprocessing duration, and reduce the processing performance of the data processing component. The data processing throughput and the preprocessing duration are inversely proportional. The higher the data processing throughput, the more data is processed in a unit of time, and the shorter the time required for preprocessing the same size of data, so as to increase the processing performance of the data processing component.
[0033] The availability preprocessing operation data of the original data set in the oil field specifically includes the number of numerical value missing of each data group in the original data set in the oil field, the number of data garbled codes of the original data set in the oil field, the number of key elements of the original data set in the oil field, the total data amount of the original data set in the oil field, and the number of data formats of the original data set in the oil field. The above availability preprocessing operation data can be extracted from the preprocessing execution report of the data processing component.
[0034] Further, the availability evaluation value of the original data set in the oil field is determined, and the specific determination process is as follows: The number of key elements of the original data set in the oil field is subjected to ratio processing with the total data amount of the original data set in the oil field to obtain the key element proportion of the original data set in the oil field.
[0035] The number of numerical value missing of each data group in the original data set in the oil field is added to obtain the total number of numerical value missing of the original data set in the oil field. The data format reference number is extracted from the oil knowledge Q&A information base.
[0036] The total number of numerical value missing of the original data set in the oil field, the number of data garbled codes of the original data set in the oil field, the key element proportion of the original data set in the oil field, the number of data formats of the original data set in the oil field, and the data preprocessing performance index of the data processing component are comprehensively analyzed to obtain the availability evaluation value of the original data set in the oil field, and the specific analysis method is as follows: ; In the formula, is the availability evaluation value of the original data set in the oil field, is the total number of numerical value missing of the original data set in the oil field, is the number of data garbled codes of the original data set in the oil field, is the key element proportion of the original data set in the oil field, and is the number of data formats of the original data set in the oil field, is the data format reference number, is the data preprocessing performance index of the data processing component, is the availability influence parameter corresponding to the total number of numerical value missing predefined in the oil knowledge Q&A information base, is the availability influence parameter corresponding to the number of data garbled codes predefined in the oil knowledge Q&A information base, is the availability influence parameter corresponding to the key element proportion predefined in the oil knowledge Q&A information base, is the availability influence parameter corresponding to the data preprocessing performance index predefined in the oil knowledge Q&A information base, and e is a natural constant.
[0037] The availability evaluation value of the original data set in the oil field in this embodiment is a quantitative evaluation of the ability of the original data set in the oil field to meet the use demand in subsequent oil field analysis, learning, etc. The higher the availability evaluation value, the more the initial data set can meet the demand of subsequent question and answer.
[0038] It should be explained that the total number of numerical value missing refers to the total number of missing specific numerical values in each independent data group in the original data in the oil field; the number of data garbled refers to the number of data units that do not conform to the data format specification and cannot be normally recognized and analyzed in the original data in the oil field; the proportion of key elements refers to the proportion of the number of data elements that are of key importance to oil knowledge question and answer, analysis, etc. in the original data set in the oil field, to the total number of elements in the data set, wherein the key elements can be oil exploration key elements such as formation lithology data, seismic wave reflection characteristic data, oil reservoir porosity and permeability data, oil production key elements such as oil well production data, working parameters of production equipment (such as stroke and number of strokes of a pumping unit), wellhead pressure data, etc. The non-key elements can be environmental data, management data, etc. in the original data set in the oil field, which are not directly used for oil exploration, development or production.
[0039] The availability influence parameter corresponding to the total number of numerical value missing, the availability influence parameter corresponding to the number of data garbled, the availability influence parameter corresponding to the proportion of key elements, and the availability influence parameter corresponding to the data preprocessing performance index are all extracted in the oil knowledge question and answer information library, for example, the total number of numerical value missing, the number of data garbled, the proportion of key elements, and the data preprocessing performance index are respectively mapped with the preset availability influence parameter corresponding to the total number of numerical value missing, the availability influence parameter corresponding to the number of data garbled, the availability influence parameter corresponding to the proportion of key elements, and the availability influence parameter corresponding to the data preprocessing performance index in the oil knowledge question and answer information library to form a mapping set. The real-time total number of numerical value missing, the number of data garbled, the proportion of key elements, and the data preprocessing performance index are brought into the mapping set to obtain the availability influence parameter corresponding to the total number of numerical value missing, the availability influence parameter corresponding to the number of data garbled, the availability influence parameter corresponding to the proportion of key elements, and the availability influence parameter corresponding to the data preprocessing performance index.
[0040] In this embodiment, if a large number of missing values exist in the key elements, it will directly affect the effectiveness of the key element proportion. When there are a large number of missing values in the key elements, the availability of these key elements for analysis or question answering is reduced, resulting in a decrease in the availability evaluation value of the original data set in the oil field. The more the number of data formats, the higher the risk of data garbled code. Different data formats have different encoding methods, storage rules and parsing requirements. When the data processing component processes data in multiple formats, it may not be familiar with some special formats or make mistakes during format conversion, resulting in data garbled code. Similarly, too many data formats may increase the total number of missing values, because different data formats may have different integrity requirements and data recording specifications. When processing data in multiple formats, some data may not be recorded correctly, thereby increasing the total number of missing values. Therefore, the number of data formats should be within an appropriate range to reduce the total number of missing values and data garbled codes, thereby improving the availability of the original data set in the oil field.
[0041] In this embodiment, the larger the data preprocessing performance index of the data processing component, the more effectively the missing values in the data set can be processed. The availability evaluation value increases due to the improvement of data integrity. At the same time, a high-performance data processing component performs well in error value recognition rate and correction accuracy, can accurately identify and correct error data, reduce the interference of error data on subsequent analysis, thereby improving the accuracy of the data. In addition, a good performance data processing component performs well in data standardization effect evaluation and normalization range compliance rate, and can compare and analyze different orders of magnitude of data (such as oil composition content, equipment operating parameters, etc.) in the original data set in the oil field on the same level, thereby improving the availability evaluation value of the data.
[0042] Specifically, the method for determining the availability state of the original data set in the oil field comprises: performing availability state verification on the availability evaluation value of the original data set in the oil field and a pre-defined availability evaluation threshold value to obtain an availability state verification result, so as to determine the availability state of the original data set in the oil field.
[0043] The availability state verification result is a first availability state verification result or a second availability state verification result. The first availability state verification result is that the availability evaluation value of the original data set in the oil field is greater than or equal to the pre-defined availability evaluation threshold value in the oil knowledge question and answer information base. The second availability state verification result is that the availability evaluation value of the original data set in the oil field is less than the availability evaluation threshold value.
[0044] If the available state verification result is the first available state verification result, it is determined that the oil field original data set is in an available state, and if the available state verification result is the second available state verification result, it is determined that the oil field original data set is in an unavailable state, and data optimization is performed on the oil field original data set.
[0045] The above data optimization of the oil field original data set can be specifically that when the oil field original data set is in an unavailable state, the distribution of the data is observed, if there are a large number of blanks or obvious inconsistencies with the expected distribution on some variables, there may be missing values, then a multiple imputation method can be used, which is not simply filled with mean or median, but generates reasonable imputation values based on the distribution of the data and the relationship between variables; similar data records are compared one by one, for example, in the oil equipment procurement data, records with the same procurement time, equipment model, supplier and procurement amount are compared, which may be repeated, then a variety of duplicate value identification methods are used to comprehensively check and remove duplicate values; check if the values of key data fields have changed or been lost, such as after converting the oil equipment maintenance record format, if the values of key fields such as equipment model, maintenance time and maintenance content are incorrect or missing, it indicates that there is a format conversion problem, then a format conversion verification mechanism is used to check the integrity and accuracy of the data after each format conversion, to ensure that there is no data loss or deformation.
[0046] In this embodiment, it needs to be explained that the data in the oil field often has high complexity, a single parameter often cannot fully reflect the true situation of the data element, and there is usually a relationship between the variables in the oil data, if only looking at a single parameter of production, some internal relationships may be ignored, by considering multiple parameters comprehensively, the rationality of the data element can be more accurately grasped, the comprehensive formula can combine multiple related parameters to more comprehensively evaluate the data element, so as to more accurately judge whether the current production data is reasonable, avoid misjudgment caused by single parameter judgment, and improve the accuracy of the data.
[0047] The oil knowledge Q&A platform constructs an oil field Q&A pair set according to the available oil field data set and a predefined construction method, performs calibration operation, determines the application quality index of each oil field Q&A pair, compares with the predefined application quality reference index, and screens to obtain an oil field Q&A pair data set.
[0048] The above constructing an oil field Q&A pair set according to a predefined construction method is specifically that the construction process is: first, constructing an oil field knowledge graph according to the available oil field data set, and second, constructing an oil field Q&A pair set according to the oil field knowledge graph.
[0049] The oil field knowledge graph is constructed, specifically, the key elements in the oil field such as oil resources, oil exploration and development, oil products, etc. are extracted from the available data set in the oil field as nodes; the elements representing the relationship between entities such as exploration, production, supply, etc. are extracted from the available data set in the oil field as the edges connecting the nodes; the key elements that may be attributes are extracted from the available data set in the oil field, and according to a plurality of named entities and a plurality of entity relationships, a plurality of knowledge triples are constructed, and according to the plurality of knowledge triples, a corresponding oil field knowledge graph is constructed, wherein the entity refers to the node in the knowledge graph, which represents the key element in the oil field, and the relationship refers to the connection and interaction between entities; extraction can be carried out by natural language processing technology for text mining, for example, using a keyword search algorithm, searching for keywords related to "oil resources" (such as "reserves", "distribution", "unconventional oil", etc.) in the available data set in the oil field, extracting data containing these keywords, and organizing and extracting them.
[0050] The specific construction of the oil field knowledge graph is shown in Figure 3 , Figure 3 is a structural diagram of the knowledge graph, and Figure 3 It can be seen that the oil field knowledge graph is specifically composed of entities, attributes and relationships, through this structure, the oil knowledge question and answer platform can clearly define the data objects that need to be stored in the available data set in the oil field, the characteristics of these objects and their mutual relationship, thereby providing a data basis for the construction of the oil field question and answer set.
[0051] According to the oil field knowledge graph, the oil field question and answer set is constructed, first, the potential question and answer pair generation object is determined according to the entity and relationship in the knowledge graph, the question type is set, for example, the question: what are the main methods of oil exploration? The answer: the main methods of oil exploration include seismic exploration, electromagnetic exploration, drilling sampling, etc. Seismic exploration analyzes the reflection of underground rock layers to infer the location of oil and gas reservoirs, while drilling confirms the existence of oil and gas reservoirs through actual drilling; using the set question as a template, automatically generating and optimizing the question through GPT-4, and using Rag to obtain the question and answer pair.
[0052] Rag (Retrieval Augmentation Technology) is a technology that combines information retrieval with generation models, the core idea of Rag is to retrieve relevant information in a large text library, and then combine the retrieved information with a generation model to generate more accurate and contextually relevant answers.
[0053] According to the existing GPT-4 generated question, the answer is given by Rag to get the question and answer pair, and the basic workflow of Rag is as follows: 1. Retrieval system - this system can retrieve relevant document fragments from the document library of the oil knowledge Q&A platform, using the BM25 retrieval algorithm to retrieve relevant document fragments according to query matching; 2. Define query and retrieval - for a given question (query), the retrieval system will find the most relevant documents or paragraphs in the knowledge base, which will be used as input for the RAG model to generate more accurate answers; 3. Use Rag to generate answers - RAG model will retrieve documents and questions as input, pass into a GPT-4, and generate answers according to RAG-Sequence or RAG-Token.
[0054] Example 1: 1) User asks: What is the chemical formula of water? 2) Generate query and perform information retrieval: The model reformulates "What is the chemical formula of water?" as a retrieval query, such as "water chemical formula". Retrieve the paragraph from the chemistry basics entry: "Water is composed of hydrogen and oxygen elements in a ratio of 2:1 in terms of the number of atoms present in the molecule, and its chemical formula is H2O." 3) Generate answer: Combine the retrieval content with the question to generate a complete answer, for example: "The chemical formula of water is H2O." Example 2: 1) User asks: What is the commonly used approximation of pi? 2) Generate query and perform information retrieval: The model reformulates the question as "pi approximation", and retrieves the paragraph: "Pi (π) is the ratio of a circle's circumference to its diameter, and its commonly used approximation is 3.14159." 3) Generate answer: For example: "The commonly used approximation of pi (π) is 3.14159." Example 3: 1) User asks: What is the name of Earth's natural satellite? 2) Generate query and perform information retrieval: Reformulate as "Earth natural satellite name", retrieve the paragraph: "Earth has a natural satellite, commonly known as 'the Moon'." 3) Generate answer: For example: "The natural satellite of Earth is called the Moon."
[0055] The above calibration operation, in this embodiment, is first calibrated by the oil knowledge Q&A platform. For example, use semantic detection tools to detect the semantics of oil field Q&A pairs, and use artificial evaluation methods to further calibrate Q&A pairs. The evaluation criteria usually involve the accuracy, relevance, fluency, and completeness of the Q&A pair, and delete the problematic Q&A pair to obtain the oil field Q&A pair dataset.
[0056] Furthermore, the specific process for determining the application quality index of each question-and-answer pair in the petroleum field is as follows: Obtain calibration operation data from the petroleum knowledge Q&A platform, specifically including the calibration time for each petroleum-related Q&A pair and the number of semantically erroneous sentences in each pair. This calibration operation data can be extracted from the calibration execution report of the petroleum knowledge Q&A platform.
[0057] The application data for question-and-answer pairs in the petroleum field was obtained, specifically including the number of literature citations for each pair, the answer codes for each pair, and the knowledge update interval for each pair. This application data can be extracted from the search reports of the petroleum knowledge question-and-answer platform.
[0058] The number of literature citations and the answer reference codes for each petroleum-related question-and-answer pair were extracted from the petroleum knowledge question-and-answer database. The answer codes and their answer reference codes for each petroleum-related question-and-answer pair were then processed to obtain and denoted as the answer code similarity for each petroleum-related question-and-answer pair.
[0059] The above data processing can specifically involve calculating the similarity of answer codes using a hash algorithm. In a specific embodiment, suppose the answer code for a question-and-answer pair in the petroleum field is... Convert this encoded sequence to MD5 (hash function) for Similarly, for the answer reference code This encoded sequence is converted into an MD5 hash function, which is used to... The two hash values are compared bit by bit. For a 128-bit MD5 hash value, the comparison starts from the first bit. and Let's count the number of identical positions. Assuming the number of identical positions is K, then the similarity can be expressed as... , where H represents the hash value and 128 represents the fixed number of bits in MD5.
[0060] The application quality index of each petroleum-related question-and-answer pair is obtained by comprehensively analyzing the answer code similarity, calibration time, number of semantic errors, number of literature citations, knowledge update interval, and usability evaluation value of the original petroleum dataset. The specific analysis method is as follows: ; In the formula, Let g be the application quality index of the g-th question-and-answer pair in the petroleum field, where g is the number of each question-and-answer pair in the petroleum field. G is the total number of oil field question and answer pairs, is the answer encoding similarity of the gth oil field question and answer pair, is the calibration duration of the gth oil field question and answer pair, is the number of semantic error sentences of the gth oil field question and answer pair, is the number of literature references of the gth oil field question and answer pair, is the number of literature reference adaptations of the gth oil field question and answer pair, is the knowledge update interval duration of the gth oil field question and answer pair, is the availability evaluation value of the oil field original data set, is the application quality impact parameter corresponding to the pre-defined answer encoding similarity in the oil knowledge question and answer information base, is the application quality impact parameter corresponding to the pre-defined calibration duration in the oil knowledge question and answer information base, is the application quality impact parameter corresponding to the pre-defined number of semantic error sentences in the oil knowledge question and answer information base, is the application quality impact parameter corresponding to the pre-defined knowledge update interval duration in the oil knowledge question and answer information base, is the application quality impact parameter corresponding to the pre-defined availability evaluation value in the oil knowledge question and answer information base, and e is a natural constant.
[0061] In this embodiment, the application quality index of each oil field question and answer pair is used to measure the quality level of each group of question and answer pairs in the oil field in actual application. The higher the application quality index, the higher the answer quality and the stronger the accuracy of the question and answer pair in actual application.
[0062] It should be explained that the above-mentioned answer encoding similarity refers to the similarity between the encoding obtained after encoding the answer of the oil field question and answer pair and the reference encoding of the answer of the question and answer pair. The number of semantic error sentences refers to the number of sentences that are incorrect due to inaccurate semantic expression, inconsistency with logic, violation of oil professional knowledge, or ambiguity, etc. in the question and answer sentences in the oil field. For example, the question is "What should be paid attention to in the maintenance of oil pipeline?", and the answer is "Pay attention to that thing to avoid problems". Here, "that thing" is semantically ambiguous and unclear about what it specifically refers to, which may be pipeline corrosion, pressure change or other factors. Such a sentence is a semantic ambiguity error sentence.
[0063] The application quality influence parameter corresponding to the answer coding similarity, the application quality influence parameter corresponding to the calibration time length, the application quality influence parameter corresponding to the semantic error sentence number, the application quality influence parameter corresponding to the knowledge update interval time length and the application quality influence parameter corresponding to the usability evaluation value are extracted from the oil knowledge Q&A information library. For example, the answer coding similarity, the calibration time length, the semantic error sentence number, the knowledge update interval time length and the usability evaluation value are respectively mapped to the application quality influence parameter corresponding to the preset answer coding similarity, the application quality influence parameter corresponding to the calibration time length, the application quality influence parameter corresponding to the semantic error sentence number, the application quality influence parameter corresponding to the knowledge update interval time length and the application quality influence parameter corresponding to the usability evaluation value in the oil knowledge Q&A information library. The real-time answer coding similarity, the calibration time length, the semantic error sentence number, the knowledge update interval time length and the usability evaluation value are brought into the mapping set to obtain the application quality influence parameter corresponding to the answer coding similarity, the application quality influence parameter corresponding to the calibration time length, the application quality influence parameter corresponding to the semantic error sentence number, the application quality influence parameter corresponding to the knowledge update interval time length and the application quality influence parameter corresponding to the usability evaluation value.
[0064] In the embodiment, the high answer coding similarity usually means that the answer of the Q&A pair is consistent with the reference standard in content structure and expression of key knowledge points, and the semantic error sentence number is relatively small, so as to improve the application quality index of the Q&A pair in each oil field. On the contrary, the large semantic error sentence number will cause a large difference between the answer coding and the reference coding, thereby reducing the answer coding similarity, and bringing negative effects to the application quality of the Q&A pair in each oil field. If the semantic error sentence number increases, the calibration time length will increase, because more time is needed to identify and correct the errors, and the answer needs to be re-audited to ensure the accuracy. When the knowledge update interval time length is too long, a large amount of knowledge needs to be updated and errors need to be corrected, thereby increasing the calibration time length and reducing the application quality of the Q&A pair in each oil field. Reasonable literature reference can reduce the semantic error sentence number, because the cited literature usually undergoes strict audit and verification, and can provide accurate knowledge support for the answer. However, when the number of literature references is too large, the content readability of the Q&A pair will be reduced, the focus will not be prominent, and the application quality of the Q&A pair will be reduced. Too few literature references will reduce the credibility and timeliness of the Q&A pair, and also reduce the application quality of the Q&A pair.
[0065] The greater the availability evaluation value of the original data set in the oil field in this embodiment, the higher the accuracy of the data, and the more accurate data support can be provided for question and answer when constructing question and answer pairs in each oil field, thereby improving the application quality index; the data set with a large availability evaluation value has better integrity, which enables more comprehensive information to be provided when answering questions in the oil field, better meets the user's demand for comprehensive knowledge acquisition, and further improves the application quality index. Meanwhile, the high-availability original data set can ensure the timeliness of the data and provide the user with the latest knowledge to improve the application quality index.
[0066] Specifically, the screening obtains the oil field question and answer pair data set, and the specific analysis process is as follows: If the application quality index of each oil field question and answer pair is greater than or equal to the application quality reference index predefined in the oil knowledge question and answer information library, each oil field question and answer pair is recorded as an oil field question and answer pair data set. If the application quality index of a certain oil field question and answer pair is less than the application quality reference index, the oil field question and answer pair is deleted, and the remaining oil field question and answer pairs are recorded as an oil field question and answer pair data set.
[0067] The oil knowledge question and answer platform stores the oil field question and answer pair data set in the large language model to which the oil question component belongs. The oil knowledge question and answer platform issues a fine-tuning instruction, and the oil question component fine-tunes the large language model according to the predefined fine-tuning manner. The oil knowledge question and answer platform checks the question and answer output effect index of the large language model against the predefined question and answer output effect index, and determines whether to repeat the fine-tuning operation on the large language model, thereby completing the oil knowledge question and answer based on the large language model.
[0068] The above oil question component is a key part of the oil knowledge question and answer platform, which is a module specially used for processing oil field question and answer related tasks. This component is mainly based on natural language processing technology and related oil field knowledge, and is responsible for receiving oil field questions raised by users, searching and matching in the stored question and answer pair data set, and then generating and returning appropriate answers. The oil question component is optimized and customized for the oil field based on the large language model, so that the question and answer service is more professional and targeted, and can better utilize the oil field question and answer pair data set to answer user questions.
[0069] The above oil question component fine-tunes the large language model according to the predefined fine-tuning manner. Specifically, the large language model to which the oil question component belongs rewrites the user's input according to the constructed knowledge graph through LoRA fine-tuning and converts it into natural language output using NLG capability.
[0070] Rewriting the question input by the user is to optimize the question through the structural model of the knowledge graph, and to make the answer of the large language model more professional and comprehensive through reasoning of the knowledge graph. In a specific embodiment, when the user inputs "when was Apple Inc. established?", based on the "Apple Inc." entity and its "establishment time" attribute in the knowledge graph, the model can rewrite the question as "in what year was Apple Inc. established?" or provide related supplementary questions such as "who is the founder of Apple Inc.?" and the like.
[0071] The rewriting of the question by the large language model is actually to structure and semanticize the information in natural language, and then use the knowledge of entities, relationships, attributes, etc. in the graph to understand and reconstruct the question. In simple terms, question rewriting is the process of the large model understanding the user's input question, and finding the optimal answer in a vast knowledge base and returning it to the user through NLP technology, which usually includes the following steps: 1. Analyze the meaning of the question; 2. Introduce the knowledge graph and map the parsed semantics to structured knowledge; 3. Rewrite the question according to the reasoning of the graph to generate a more accurate question; 4. Generate an answer and return the accurate result; 5. Convert it to natural language output with NLP technology.
[0072] Further, the evaluation of the question and answer output effect index of the large language model includes the following steps: Obtain the question and answer output data of the large language model, including the number of knowledge points covered by the large language model in each question and answer output, the amount of additional information provided by the large language model in each question and answer output, the answer output time of the large language model in each question and answer output, and the number of user praise times of the large language model in the question and answer output period.
[0073] Each question and answer output above refers to a number of question and answer output processes in the question and answer output period. The determination of the question and answer output period is based on the state of the oil question and answer component, the actual output demand, and the specific knowledge question, etc. The above question and answer output data can be extracted from the output report of the oil question and answer component.
[0074] Obtain the total amount of answer information of the large language model in each question and answer output, which can be extracted from the output report of the oil question and answer component. The amount of additional information provided by the large language model in each question and answer output is compared with the total amount of answer information of the large language model in each question and answer output to obtain the proportion of additional information of the large language model in each question and answer output.
[0075] The number of knowledge points covered by the large language model in each question and answer output is compared with the number of knowledge points covered by the pre-defined question and answer output in the oil knowledge question and answer information library to obtain the knowledge point coverage degree value of the large language model in each question and answer output. The additional information reference proportion under each question and answer output, the answer output reference time length under each question and answer output, and the user praise definition times are extracted from the oil knowledge question and answer information library.
[0076] The knowledge point coverage degree value of the large language model in each question and answer output, the additional information proportion of the large language model in each question and answer output, the answer output time length of the large language model in each question and answer output, the user praise times of the large language model in the question and answer output period, the availability evaluation value of the oil field original data set, and the application quality index of each oil field question and answer pair are comprehensively analyzed to obtain the question and answer output effect index of the large language model, and the specific analysis method is as follows: ; In the formula, is the question and answer output effect index of the large language model, is the knowledge point coverage degree value of the large language model in the dth question and answer output, d is the number of each question and answer output, , D is the total amount of question and answer output, is the additional information proportion of the large language model in the dth question and answer output, is the additional information reference proportion under the dth question and answer output, is the answer output time length of the large language model in the dth question and answer output, is the answer output reference time length under the dth question and answer output, is the user praise times of the large language model in the question and answer output period, is the user praise definition times, is the application quality index of the gth oil field question and answer pair, g is the number of each oil field question and answer pair, , G is the total amount of oil field question and answer pairs, is the availability evaluation value of the oil field original data set, is the output effect influence parameter corresponding to the pre-defined knowledge point coverage degree value in the oil knowledge question and answer information library, is the output effect influence parameter corresponding to the pre-defined application quality index in the oil knowledge question and answer information library, is the output effect influence parameter corresponding to the pre-defined availability evaluation value in the oil knowledge question and answer information library, and e is a natural constant.
[0077] The question and answer output effect index of the large language model in this embodiment is used to measure the performance and quality of the large language model in answering questions, helping to evaluate whether the model can accurately, completely, clearly and reasonably answer various questions raised by users, especially in the application scenario of this specific field of oil knowledge question and answer, reflecting the practicality and reliability of the model output answer.
[0078] It should be explained that the knowledge point coverage value mentioned above refers to the proportional relationship between the oil field knowledge points involved in the answers output by the large language model in each question and answer output and the number of pre-defined relevant knowledge points involved in the answers. For example, the principle of seismic exploration technology includes the generation, propagation and reflection law of seismic waves, and how to infer the underground geological structure by receiving reflected waves. The principle of logging technology, such as electrical logging and acoustic logging, is how different logging methods obtain formation information. When answering questions such as "how does seismic exploration find oil", these technical principles are knowledge points. For example, the principle, applicable conditions and advantages and disadvantages of different production methods such as self-flowing production, mechanical production (rod pump production, rodless pump production), etc. When answering the question "under what circumstances should self-flowing production be selected", the relevant knowledge of production methods is the knowledge point.
[0079] The proportion of additional information refers to the proportion of other information in the output answer of the large language model in each question and answer output, in addition to the core information required to directly answer the question. These additional information may include relevant background knowledge, supplementary explanation, example or other related but unnecessary content. The reference proportion of additional information refers to the corresponding adaptation value of the pre-defined proportion of additional information in each question and answer output. The answer output time length refers to the time span from when the large language model receives the user's question to when the answer is completely output. This time span covers the entire process of the model's understanding of the question, searching for relevant information in the knowledge base, generating answers, and outputting the answers.
[0080] The output effect influence parameter corresponding to the knowledge point coverage value, the output effect influence parameter corresponding to the usability evaluation value and the output effect influence parameter corresponding to the application quality index are extracted from the oil knowledge question and answer information base. For example, the knowledge point coverage value, the usability evaluation value and the application quality index form a mapping set with the output effect influence parameter corresponding to the pre-set knowledge point coverage value in the oil knowledge question and answer information base, the output effect influence parameter corresponding to the usability evaluation value and the output effect influence parameter corresponding to the application quality index. The knowledge point coverage value, the usability evaluation value and the application quality index are brought into the mapping set to obtain the output effect influence parameter corresponding to the knowledge point coverage value, the output effect influence parameter corresponding to the usability evaluation value and the output effect influence parameter corresponding to the application quality index.
[0081] In this embodiment, when the knowledge point coverage value of the large language model under each question and answer output is high within a reasonable range, the comprehensiveness of the user's question can be met, thereby improving the number of user good comments, and thus improving the question and answer output effect index of the large language model; appropriate additional information can enrich the answer content and help the user better understand the core knowledge, which can also improve the number of user good comments, but when the proportion of excessive information is too high and greater than the reference proportion of additional information, it will lead to negative problems such as unclear answer focus and information redundancy, and reduce the question and answer output effect of the large language model; at the same time, too high a reference proportion of additional information will generally also lead to a longer answer output time, and if the answer output time is too long and deviates from the answer output reference time, it will lead to a decrease in user good comments, thereby greatly reducing the question and answer output effect of the large language model.
[0082] In this embodiment, the greater the availability evaluation value of the original data set in the oil field and the greater the application quality index of each oil field question and answer pair, the higher the accuracy of the original data in the oil field, reducing the interference of error data, making the model output answer more in line with the actual situation, thereby improving the accuracy in the question and answer output effect index; a good data set is also excellent in data consistency, and unified data can ensure that the model output answer is clear and accurate, and will not mislead the user due to the confusion of units or definitions, which helps to improve the clarity in the question and answer output effect index, in addition, a high application quality index means that the accuracy of the question and answer pair itself is higher, and the large language model can more accurately answer the user's question when referring to these high-quality question and answer pairs to generate answers, which also improves the accuracy in the question and answer output effect index.
[0083] Specifically, the oil knowledge question and answer platform determines whether to perform repeated fine-tuning operation on the large language model, specifically, the oil question and answer component performs output effect verification on the question and answer output effect index of the large language model and the predefined question and answer output effect index, obtains an output effect verification result and uploads it to the oil knowledge question and answer platform, so that the oil knowledge question and answer platform determines whether to perform repeated fine-tuning operation on the large language model.
[0084] The output effect verification result is a first output effect verification result or a second output effect verification result. The first output effect verification result is that the question and answer output effect index of the large language model is greater than or equal to the pre-defined question and answer output effect index in the oil knowledge question and answer information base. The second output effect verification result is that the question and answer output effect index of the large language model is less than the question and answer output effect index. If the output effect verification result is the first output effect verification result, the oil knowledge question and answer platform determines that the large language model does not need to be repeatedly fine-tuned. If the output effect verification result is the second output effect verification result, the oil knowledge question and answer platform determines that the large language model needs to be repeatedly fine-tuned, and issues a repeated fine-tuning instruction. The oil question and answer component repeatedly fine-tunes the large language model according to the pre-defined repeated fine-tuning manner.
[0085] Further, the pre-defined repeated fine-tuning manner is that the oil question and answer component uses the built-in LoRA method to repeatedly fine-tune the large language model in the oil knowledge field.
[0086] The LoRA method is a technology that introduces an adaptive matrix into the large language model. The repeated fine-tuning operation is to keep the weights of the original model in the oil question and answer component unchanged, add a low-rank matrix to adjust the output of the original model, and make the original model become a large language model in the oil knowledge field, so as to complete the repeated fine-tuning operation of the large language model.
[0087] In a simple way, it is to embed the above-mentioned oil field data set into the original model of the oil question and answer component, change the parameter weights related to the oil field in the original model, and make the original large language model more proficient in the knowledge of the oil related field. The essence of model modification is the modification of model parameters - parameters (in natural language processing, it can be understood that the user's question is converted into a set of numbers according to a specific method, and is usually in the form of a matrix), the essence of fine-tuning is to change the original parameters of the model from one state to another, so that the model is more proficient in a certain field, and LoRA fine-tuning has lower storage requirements than full fine-tuning, and limited computing resources can also be adapted.
[0088] Based on the embodiments of the present application, the oil knowledge question and answer system can also improve the operation instruction analysis accuracy of the oil pumping equipment, dynamic risk assessment and real-time adaptive control. The specific examples are as follows: The oil knowledge question and answer system adopts a full-flow design of multi-modal input-layered processing-dynamic execution-feedback iteration, including: (1) Multi-modal command input and preprocessing module Input collection: Receive instructions input by operators through voice (dialect supported) or keyboard (e.g., "Adjust the pumping unit stroke to 5 times per minute, it's that well with high downhole pressure").
[0089] Transcription and de-duplication: Convert voice to text through the Coze audio-to-text plugin, call the fine-tuned LLM (the oil knowledge Q&A system in this embodiment) to filter redundant information (e.g., "just that" "a little"), and retain core actions ("adjust stroke"), objects ("pumping unit"), and parameters ("5 times per minute" "high downhole pressure").
[0090] Feature extraction: Generate a "action-object-parameter-condition" quadruple (e.g., "adjust-pumping unit-stroke 5 times per minute-high downhole pressure"), providing structured data for subsequent processing.
[0091] (2) Professional command generation module in the oil field Professional instruction conversion: Input the quadruple into the oil knowledge Q&A system in this embodiment, and convert it into standardized control language (e.g., "pumping unit stroke 3 meters, stroke 5 times per minute, based on downhole pressure 1.2 MPa dynamic adjustment").
[0092] Instruction splitting and verification: Split the steps by logic (e.g., "1. Check the current stroke; 2. Adjust to 5 times per minute; 3. Monitor pressure changes in real time"), support operators to modify instructions through voice or keyboard until meet the requirements.
[0093] (3) Code generation and multi-layer risk assessment module Code generation: Call the Coze code generation plugin to convert the split instructions into executable code for pumping equipment embedded systems (e.g., C++ script).
[0094] Expert model evaluation Expert model 1 (safety rule verification): Based on the oil engineering rule library (e.g., "sand content > 3% stroke ≤ 4 times per minute") and real-time parameters of the equipment (e.g., current sand content 3.5%), evaluate the risk of code execution, if the stroke setting violates the rules, trigger an early warning and correct it.
[0095] Expert model 2 (parameter rationality verification): Check the parameter matching in the code (e.g., "whether the stroke 3 meters and the stroke 5 times per minute are suitable for the pumping rod strength"), output optimization suggestions (e.g., "adjust the stroke to 4 times per minute to match the stroke").
[0096] Expert model 3 (syntax and logic verification): Verify the correctness of the code syntax and the coherence of the execution logic to ensure that there are no program errors.
[0097] (4) Control optimization and execution module Dynamic parameter optimization: Combined with real-time data from downhole sensors (temperature, pressure, sand content), the adaptive control parameters (such as "when the pressure rises to 1.5 MPa, the stroke frequency automatically decreases by 10%") are generated by the oil knowledge Q&A system of the embodiment.
[0098] Execution and emergency stop mechanism: The optimized code is transmitted to the embedded end for execution, and the emergency stop triggering mechanism is started - when the sensor detects a sudden abnormality (such as sudden increase in sucker rod load, wellhead leakage), the execution is immediately terminated and an alarm is triggered.
[0099] (5) Data feedback and model iteration module Data acquisition and processing: The embedded end collects real-time equipment operation data (stroke, stroke frequency, energy consumption) and environmental data (downhole pressure, temperature) and uploads them to the cloud database after compression by the Coze data packaging plug-in.
[0100] Oil knowledge fusion feedback: By comparing the actual execution effect with the expected target (such as whether adjusting the stroke frequency reduces the risk of pump sticking), and combining the oil field knowledge graph (such as "the relationship between sucker rod wear and stroke frequency"), the model iteration direction is generated, continuously optimizing the instruction analysis and decision-making ability of the oil knowledge Q&A system of the embodiment.
[0101] It needs to be explained that in a specific embodiment, the oil field LLM fine-tuning method of the present application based on the above embodiment can be as follows: (1) Knowledge graph construction Entity and relationship extraction: Extract entities (such as "pumping unit", "stroke", "sand content") and relationships (such as "stroke frequency affects pumping efficiency", "high sand content leads to pump sticking") from oil engineering literature and pumping equipment manuals to construct knowledge triples.
[0102] Graph application: Embed the knowledge graph into the LLM to make the model understand professional associations such as "reducing stroke frequency can reduce sucker rod wear" and improve the professionalism of instruction analysis.
[0103] (2) Question and answer pair generation and calibration Graph-based question and answer pair design: generate professional oil field question and answer pairs based on the knowledge graph, such as: Question: "When the sand content exceeds 3%, how should the stroke frequency of the pumping unit be adjusted?" Answer: "When sand content exceeds 3%, the pumping frequency should be reduced to less than 4 times per minute to reduce the wear of the sucker rod and oil pipe and the risk of pump sticking." RAG enhancement: search for oil industry standards, failure cases, etc. documents using BM25 algorithm, generate high-precision question and answer pairs with GPT-4o, and form fine-tuning data set after manual calibration (accuracy, professional verification).
[0104] (3) LLM fine-tuning implementation Base model selection: based on ChatGLM3-6B model, using LoRA (low rank adaptation) method for parameter efficient fine-tuning, only adjusting the model parameters related to the oil field, reducing the demand for computing resources.
[0105] Fine-tuning target: enable the model to analyze oil professional instructions, correlate equipment parameters and downhole working conditions, and generate reasonable control logic.
[0106] 3. Question rewriting and intelligent decision-making Question rewriting mechanism: when the user inputs a vague instruction (such as "this well is not pumping"), the model analyzes the semantics and rewrites it into a precise question (such as "is the pumping machine stuck due to high sand content? Do you need to adjust the pumping frequency?").
[0107] NLG output: through natural language generation technology, the device state and decision suggestion are converted into colloquial language (such as "current sand content is 4%, it is recommended to reduce the pumping frequency from 5 times per minute to 3 times per minute to avoid pump sticking"), which is easy for operators to understand.
[0108] Reference Figure 2 The second aspect of the present application provides a large language model-based oil knowledge question and answer system, comprising: a data availability preprocessing module, a question and answer pair application quality determination module, and a model output effect determination module.
[0109] The second aspect of the present application provides a large language model-based oil knowledge question and answer system, further comprising: an oil knowledge question and answer information base for storing CPU average occupancy rate adaptation value, memory maximum occupancy rate reference value, data format reference number, literature reference adaptation number of each oil field question and answer pair, answer reference encoding of each oil field question and answer pair, additional information reference proportion under each question and answer output, answer output reference time under each question and answer output, user praise definition times, and pre-set values of various factors.
[0110] The data availability preprocessing module is connected to the question and answer pair application quality determination module, the question and answer pair application quality determination module is connected to the model output effect determination module, and the data availability preprocessing module, the question and answer pair application quality determination module and the model output effect determination module are all connected to the oil knowledge question and answer information base.
[0111] The data availability preprocessing module is used for the oil knowledge Q&A platform to collect original data of the oil field from various data sources, denoted as an original data set of the oil field, and a data processing component performs an availability preprocessing operation on the original data set of the oil field, determines an availability evaluation value of the original data set of the oil field, checks the availability evaluation value with a predefined availability evaluation threshold, determines the availability state of the original data set of the oil field, and records the original data set of the oil field after the availability preprocessing operation as an available data set of the oil field.
[0112] The Q&A pair application quality module is used for the oil knowledge Q&A platform to construct an oil field Q&A pair set from the available data set of the oil field according to a predefined construction manner, and perform a calibration operation to determine an application quality index of each oil field Q&A pair, compare the application quality index with a predefined application quality reference index, and screen an oil field Q&A pair data set.
[0113] The model output effect module is used for the oil knowledge Q&A platform to store the oil field Q&A pair data set in a large language model to which the oil Q&A component belongs, and the oil knowledge Q&A platform issues a fine-tuning instruction, so that the oil Q&A component performs a fine-tuning operation on the large language model according to a predefined fine-tuning manner, evaluates a Q&A output effect index of the large language model, checks the Q&A output effect index with a predefined Q&A output effect index, and determines whether to perform a repeated fine-tuning operation on the large language model, so as to complete the oil knowledge Q&A based on the large language model.
[0114] The above content is merely an example and description of the structure of the present application, and those skilled in the art can make various modifications or supplements or use similar ways to replace the described specific embodiments, as long as they do not deviate from the structure of the present application or exceed the scope defined by the present application, and should belong to the protection scope of the present application.
Claims
1. A petroleum knowledge question-answering method based on a large language model, characterized in that, include: Data availability preprocessing: The petroleum knowledge Q&A platform collects raw petroleum data from various data sources, which is referred to as the raw petroleum dataset. The data processing component performs availability preprocessing on the raw petroleum dataset, determines the availability assessment value of the raw petroleum dataset, verifies it against the predefined availability assessment threshold, determines the availability status of the raw petroleum dataset, and records the raw petroleum dataset after availability preprocessing as the available petroleum dataset. Question-answer pair application quality assessment: The petroleum knowledge question-answering platform constructs a petroleum field question-answer pair set from available datasets in the petroleum field according to a predefined construction method, performs calibration operations, determines the application quality index of each petroleum field question-answer pair, compares it with a predefined application quality reference index, and selects the petroleum field question-answer pair dataset. Model output performance evaluation: The petroleum knowledge question-and-answer platform stores the petroleum-related question-and-answer pair dataset in the large language model to which the petroleum question-and-answer component belongs. The petroleum knowledge question-and-answer platform issues fine-tuning instructions, and the petroleum question-and-answer component performs fine-tuning operations on the large language model according to the predefined fine-tuning method. The question-and-answer output performance indicators of the large language model are evaluated and verified against the predefined question-and-answer output performance indicators. The petroleum knowledge question-and-answer platform determines whether to repeat the fine-tuning operation on the large language model, thereby completing the petroleum knowledge question-and-answer based on the large language model.
2. The petroleum knowledge question-answering method based on a large language model according to claim 1, characterized in that: The data processing component performs usability preprocessing on the raw dataset in the petroleum field, thereby obtaining the preprocessing performance data of the data processing component and the usability preprocessing operation data of the raw dataset in the petroleum field. The preprocessing performance data of the data processing component specifically includes the CPU utilization rate of the data processing component at each preprocessing time point, the memory utilization rate of the data processing component at each preprocessing time point, the preprocessing duration of the data processing component within the preprocessing cycle, and the data processing volume of the data processing component within the preprocessing cycle. The average CPU utilization of the data processing component at each preprocessing time point is averaged to obtain the average CPU utilization of the data processing component during the preprocessing cycle. The data processing throughput of the data processing component in the preprocessing period is obtained by comparing the amount of data processed by the data processing component in the preprocessing period with the preprocessing time of the data processing component in the preprocessing period. The average CPU utilization threshold and the maximum memory utilization reference value were extracted from the petroleum knowledge Q&A database. The data processing component's average CPU utilization during the preprocessing cycle, memory utilization at each preprocessing time point, preprocessing duration, and data processing throughput during the preprocessing cycle are combined to obtain the data preprocessing performance index of the data processing component. The availability preprocessing operation data of the original dataset in the petroleum field specifically includes the number of missing values in each data group of the original dataset in the petroleum field, the number of garbled data in the original dataset in the petroleum field, the number of key elements in the original dataset in the petroleum field, the total data volume of the original dataset in the petroleum field, and the number of data formats in the original dataset in the petroleum field.
3. The petroleum knowledge question-answering method based on a large language model according to claim 2, characterized in that: The specific process for determining the usability assessment value of the original dataset in the petroleum field is as follows: The ratio of the number of key elements in the original dataset of the petroleum field to the total amount of data in the original dataset of the petroleum field is processed to obtain the proportion of key elements in the original dataset of the petroleum field. The total number of missing values in the original dataset of the petroleum field is obtained by summing the number of missing values in each data group. The number of data format references was extracted from the petroleum knowledge Q&A database; By comprehensively analyzing the total number of missing values, the number of garbled characters, the proportion of key elements, the number of data formats, and the data preprocessing performance indicators of the data processing components in the original dataset of the petroleum field, a usability assessment value for the original dataset of the petroleum field is obtained.
4. The petroleum knowledge question-answering method based on a large language model according to claim 3, characterized in that: The determination of the availability status of the original dataset in the petroleum field specifically involves verifying the availability assessment value of the original dataset in the petroleum field against a predefined availability assessment threshold to obtain the availability status verification result, thereby determining the availability status of the original dataset in the petroleum field. The availability status verification result is either the first availability status verification result or the second availability status verification result. The first availability status verification result is specifically the availability assessment value of the original dataset in the petroleum field being greater than or equal to the availability assessment threshold. The second availability status verification result is specifically that the availability assessment value of the original dataset in the petroleum field is less than the availability assessment threshold; If the availability status verification result shows the first availability status verification result, the original dataset in the petroleum field is determined to be in an available state. If the availability status verification result shows the second availability status verification result, the original dataset in the petroleum field is determined to be in an unavailable state, and data optimization is performed on the original dataset in the petroleum field.
5. The petroleum knowledge question-answering method based on a large language model according to claim 1, characterized in that: The specific process for determining the application quality index of each question-and-answer pair in the petroleum field is as follows: Obtain calibration operation data from the petroleum knowledge Q&A platform, specifically including the calibration time of each petroleum field Q&A pair and the number of semantically incorrect sentences in each petroleum field Q&A pair; Obtain application data for question-and-answer pairs in the petroleum field, specifically including the number of literature citations for each question-and-answer pair, the answer codes for each question-and-answer pair, and the knowledge update interval for each question-and-answer pair. The number of literature citations for each question-and-answer pair in the petroleum field and the answer reference codes for each question-and-answer pair in the petroleum field were extracted from the petroleum knowledge question-and-answer database. The answer codes of each petroleum field question-and-answer pair are processed with the answer reference codes of each petroleum field question-and-answer pair to obtain and record the answer code similarity of each petroleum field question-and-answer pair. The application quality index of each petroleum field question-answer pair is obtained by comprehensively analyzing the answer encoding similarity, calibration time, number of semantic errors, number of literature citations, knowledge update interval, and availability assessment value of the original petroleum field dataset.
6. The petroleum knowledge question-answering method based on a large language model according to claim 1, characterized in that: The filtering process yielded a question-and-answer pair dataset in the petroleum field, and the specific analysis process is as follows: If the application quality index of each petroleum field question-and-answer pair is greater than or equal to the application quality reference index, then the data corresponding to each petroleum field question-and-answer pair will be recorded as the petroleum field question-and-answer pair dataset. If the application quality index of a certain question-and-answer pair in the petroleum field is less than the application quality reference index, then the question-and-answer pair in the petroleum field will be deleted. Several question-and-answer pairs in the petroleum field whose application quality index is greater than or equal to the application quality reference index will be counted, and the data corresponding to several question-and-answer pairs in the petroleum field will be recorded as the petroleum field question-and-answer pair dataset.
7. The petroleum knowledge question-answering method based on a large language model according to claim 1, characterized in that: The evaluation metrics for assessing the question-answering output performance of the large language model are evaluated as follows: Obtain the question-and-answer output data of the large language model, specifically including the number of knowledge points covered by the large language model in each question-and-answer output, the amount of additional information provided by the large language model in each question-and-answer output, the answer output duration of the large language model in each question-and-answer output, and the number of positive user reviews of the large language model within the question-and-answer output cycle. Obtain the total amount of answer information from the large language model across all question-and-answer sessions; The ratio of the amount of additional information provided by the large language model in each question-and-answer output to the total amount of answer information in each question-and-answer output is used to obtain the proportion of additional information provided by the large language model in each question-and-answer output. The knowledge point coverage value of the large language model under each question-and-answer output is obtained by comparing the number of knowledge points covered by the large language model under each question-and-answer output with the predefined number of knowledge point coverage fits under each question-and-answer output. The following data were extracted from the petroleum knowledge Q&A database: the percentage of additional information referenced in each Q&A output, the reference duration of the answer output in each Q&A output, and the number of user positive reviews. The question-answering output performance index of the large language model is obtained by comprehensively analyzing the knowledge point coverage value, the proportion of additional information, the answer output time, the number of positive user reviews within the question-answering output cycle, the usability assessment value of the original dataset in the petroleum field, and the application quality index of each question-answering pair in the petroleum field.
8. The petroleum knowledge question-answering method based on a large language model according to claim 7, characterized in that: The petroleum knowledge question-and-answer platform determines whether to perform repeated fine-tuning operations on the large language model. Specifically, the petroleum question-and-answer component verifies the output effect index of the large language model with the predefined output effect index, obtains the output effect verification result, and uploads it to the petroleum knowledge question-and-answer platform. The petroleum knowledge question-and-answer platform then determines whether to perform repeated fine-tuning operations on the large language model. The output effect verification result is either the first output effect verification result or the second output effect verification result. The first output effect verification result is specifically that the question-answering output effect index of the large language model is greater than or equal to the question-answering output effect index. The second output effect verification result is that the question-answering output effect index of the large language model is smaller than the question-answering output effect index. If the output effect verification result shows the first output effect verification result, the petroleum knowledge Q&A platform determines that there is no need to perform repeated fine-tuning operations on the large language model. If the output effect verification result shows the second output effect verification result, the petroleum knowledge Q&A platform determines that repeated fine-tuning operations on the large language model are required, and issues repeated fine-tuning instructions. The petroleum Q&A component performs repeated fine-tuning operations on the large language model according to the predefined repeated fine-tuning method.
9. The petroleum knowledge question-answering method based on a large language model according to claim 8, characterized in that: The predefined repeated fine-tuning method specifically involves the petroleum question-answering component using the built-in LoRA method to repeatedly fine-tune the large language model of the petroleum knowledge domain. The LoRA method is a technique that introduces an adaptation matrix into a large language model. Specifically, the repeated fine-tuning operation keeps the weights of the original model in the petroleum question-answering component unchanged, adds a low-rank matrix to adjust the output of the original model, and makes the original model a large language model in the petroleum knowledge domain, thereby completing the repeated fine-tuning operation of the large language model.
10. A system applying a petroleum knowledge question-answering method based on a large language model as described in any one of claims 1-9, characterized in that: include: The data availability preprocessing module is used by the petroleum knowledge Q&A platform to collect raw petroleum data from various data sources, which is referred to as the raw petroleum dataset. The data processing component performs availability preprocessing on the raw petroleum dataset, determines the availability assessment value of the raw petroleum dataset, verifies it with the predefined availability assessment threshold, determines the availability status of the raw petroleum dataset, and records the raw petroleum dataset after availability preprocessing as the available petroleum dataset. The question-and-answer pair application quality module is used by the petroleum knowledge question-and-answer platform to construct a petroleum field question-and-answer pair set from available datasets in the petroleum field according to a predefined construction method, perform calibration operations, determine the application quality index of each petroleum field question-and-answer pair, compare it with a predefined application quality reference index, and filter to obtain the petroleum field question-and-answer pair dataset. The model output effect module is used by the petroleum knowledge question-and-answer platform to store the petroleum-related question-and-answer pair dataset in the large language model to which the petroleum question-and-answer component belongs. The petroleum knowledge question-and-answer platform issues fine-tuning instructions, and the petroleum question-and-answer component performs fine-tuning operations on the large language model according to the predefined fine-tuning method. The question-and-answer output effect index of the large language model is evaluated and verified with the predefined question-and-answer output effect index. The petroleum knowledge question-and-answer platform determines whether to repeat the fine-tuning operation on the large language model, thereby completing the petroleum knowledge question-and-answer based on the large language model.
Citation Information
Patent Citations
Question answering method, question answering device and storage device based on Transformer model
CN111881279B
Question answering methods, training methods and devices for question answering models
CN114416953B
Instruction fine tuning data set construction method and device
CN118966379A
Knowledge question-answering system based on large language model
CN119396975A
Light field image spatial super-resolution reconstruction method and device, equipment and storage medium
CN120543384A
Cited By
A slope engineering knowledge intelligent interaction question and answer method and system
CN122364379A