Oil knowledge question and answer method and system based on large language model
By optimizing data processing and question-answer pair construction, and by fine-tuning the large language model, the accuracy and efficiency issues of petroleum knowledge question answering were resolved, achieving more accurate petroleum-related knowledge question answering.
Patent Information
- Application Number
- CN202511378527.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing model-based knowledge question answering technologies fail to reach their full potential when faced with specific tasks, resulting in significant biases in petroleum knowledge question answering.
This study optimizes a petroleum knowledge question-answering system based on a large language model by preprocessing data availability, judging the quality of question-answer pairs, and judging the model output effect. This includes data collection, preprocessing, question-answer pair construction, and model fine-tuning, ensuring data quality and question-answer accuracy.
It improved the efficiency and accuracy of petroleum knowledge Q&A, optimized the data processing flow, enhanced the relevance and relevance of Q&A components, and reduced vague or erroneous statements.
Smart Images

Figure CN120893584B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computing models, in particular to a petroleum knowledge question and answer method and system based on a large language model. BACKGROUND
[0002] A large language model (LLM) is a language model composed of a neural network with many parameters (usually billions of weights or more), which is trained on a large amount of unlabelled text using self-supervised learning or semi-supervised learning. With the continuous innovation of technology, such as deep sea exploration technology, shale oil exploitation technology, and the emergence of emerging fields, a large amount of new knowledge continues to flow in, while traditional petroleum knowledge has been accumulated for many years and has formed a deep deposit, which makes petroleum practitioners face great challenges in knowledge acquisition. Whether it is for new employees to quickly get familiar with the business, or for experienced experts to track the latest developments and solve cross-disciplinary problems, efficient and accurate knowledge support is needed.
[0003] For example, the invention patent with publication number CN111881279B discloses a question and answer method, a question and answer device and a storage device based on a Transformer model, the question and answer method comprising: obtaining a question text input by a user, processing the question text to obtain a question sequence; decoding the question sequence to obtain a plurality of candidate answers related to the question sequence; concatenating the question sequence with each candidate answer; scoring each concatenation result to select the candidate answer corresponding to the highest score as the optimal answer for the question sequence.
[0004] For example, the invention patent with publication number CN114416953B discloses a question and answer processing method, a training method, device, equipment, medium and product of a question and answer model, relating to the technical field of artificial intelligence, specifically to the technical field of natural language processing, deep learning and knowledge graph, the question and answer processing method comprising: obtaining processing data, wherein the processing data includes question data and candidate answers; performing general semantic understanding on the processing data to obtain general data features; selecting a target question and answer processing mode from candidate question and answer processing modes based on the general data features; processing the general data features using the target question and answer processing mode to obtain a target answer in the candidate answers for the question data.
[0005] In combination with the above technical solutions, it is found that the existing model-based knowledge question and answer technical solutions generally have general applicability, so that they may not be able to perform well when facing specific tasks, resulting in large deviations in knowledge question and answer, and ultimately affecting the use of petroleum knowledge question and answer. SUMMARY
[0006] In view of the deficiencies of the prior art, the present application provides a petroleum knowledge question and answer method and system based on a large language model, which can effectively solve the problems involved in the above background art.
[0007] To achieve the above object, the present application is implemented by the following technical solutions: The present application provides a petroleum knowledge question and answer method based on a large language model, comprising: data availability preprocessing: a petroleum knowledge question and answer platform collects original petroleum domain data from each data source, denoted as an original petroleum domain data set, a data processing component performs availability preprocessing operation on the original petroleum domain data set, determines the availability evaluation value of the original petroleum domain data set, checks the availability evaluation threshold value predefined, determines the availability state of the original petroleum domain data set, and records the original petroleum domain data set after the availability preprocessing operation as a petroleum domain available data set; question and answer pair application quality determination: the petroleum knowledge question and answer platform constructs a petroleum domain question and answer pair set according to a predefined construction method, and performs calibration operation, determines the application quality index of each petroleum domain question and answer pair, compares with the predefined application quality reference index, and screens to obtain a petroleum domain question and answer pair data set; model output effect determination: the petroleum knowledge question and answer platform stores the petroleum domain question and answer pair data set in a large language model to which a petroleum question and answer component belongs, the petroleum knowledge question and answer platform issues a fine-tuning instruction, and the petroleum question and answer component performs fine-tuning operation on the large language model according to a predefined fine-tuning method, evaluates the question and answer output effect index of the large language model, checks the predefined question and answer output effect index, and the petroleum knowledge question and answer platform determines whether to perform repeated fine-tuning operation on the large language model, thereby completing the petroleum knowledge question and answer based on the large language model.
[0008] The second aspect of the present application provides a large language model-based oil knowledge Q&A system, comprising: a data availability preprocessing module, used for an oil knowledge Q&A platform to collect oil field original data of each data source, denoted as an oil field original data set, a data processing component to perform availability preprocessing operation on the oil field original data set, to determine an availability evaluation value of the oil field original data set, to check the availability evaluation value with a predefined availability evaluation threshold, to determine the availability state of the oil field original data set, and to denote the oil field original data set after the availability preprocessing operation as an oil field available data set; a Q&A pair application quality module, used for the oil knowledge Q&A platform to construct an oil field Q&A pair set from the oil field available data set according to a predefined construction manner, to perform calibration operation, to determine an application quality index of each oil field Q&A pair, to compare the application quality index with a predefined application quality reference index, and to screen an oil field Q&A pair data set; and a model output effect module, used for the oil knowledge Q&A platform to store the oil field Q&A pair data set in a large language model to which the oil Q&A component belongs, for the oil knowledge Q&A platform to issue a fine-tuning instruction, for the oil Q&A component to perform fine-tuning operation on the large language model according to a predefined fine-tuning manner, for the oil knowledge Q&A platform to check an Q&A output effect index of the large language model with a predefined Q&A output effect index, for the oil knowledge Q&A platform to determine whether to perform repeated fine-tuning operation on the large language model, and for the oil knowledge Q&A platform to complete the large language model-based oil knowledge Q&A.
[0009] Compared with the prior art, the embodiments of the present application have at least the following advantages or beneficial effects:
[0010] (1) The present application provides a large language model-based oil knowledge Q&A method and system. The oil knowledge Q&A platform first collects oil field original data of each data source, and obtains an available data set after availability preprocessing and checking by a data processing component. Then, a Q&A pair set is constructed from the data set according to a predefined manner, a Q&A pair data set is screened out after calibration and comparison, more accurate information is provided for an oil Q&A component, and the efficiency and accuracy of the Q&A component are improved. Then, the Q&A pair data set is stored in a large language model of the oil Q&A component, a fine-tuning instruction is issued by the platform, the oil Q&A component fine-tunes the model according to a predefined manner, the Q&A output effect index of the model is evaluated and checked with a predefined index, it is determined whether to repeat fine-tuning, and the cycle is repeated until the large language model-based oil knowledge Q&A is finally realized.
[0011] (2) The present application performs availability preprocessing operation on the oil field original data set by a data processing component, determines a data preprocessing performance index of the data processing component, understands the resource requirements of the data processing component at different stages, optimizes the data processing process to improve the overall data processing speed, determines the stability of the data processing component according to the data preprocessing performance index, and helps to ensure the data processing quality.
[0012] (3) The present application evaluates the availability of the original data set in the oil field by determining the availability evaluation value, and corrects the errors and missing problems of the data in the oil field caused by various factors through availability preprocessing, so that the data is more accurate and reliable, and can provide higher quality data basis for subsequent knowledge question and answer work.
[0013] (4) The present application fine-tunes the large language model according to the pre-defined fine-tuning mode through the oil question and answer component, and through the fine-tuning operation, the oil question and answer component can better understand the oil professional terms, concepts and complex process flow, the fine-tuned component can more accurately combine the oil geology and mining knowledge to answer, reduce ambiguous or incorrect expressions, so as to accurately match the user's oil knowledge demand, evaluate the question and answer output effect index of the large language model, optimize the scientificity of the answer, and enhance the relevance and pertinence of the component. BRIEF DESCRIPTION OF DRAWINGS
[0014] The present application is further illustrated by the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present application. For ordinary skilled in the art, other drawings can be obtained without creative labor on the basis of the following drawings.
[0015] Figure 1 The present application is further illustrated by the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present application. For ordinary skilled in the art, other drawings can be obtained without creative labor on the basis of the following drawings.
[0016] Figure 2 The present application is further illustrated by the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present application. For ordinary skilled in the art, other drawings can be obtained without creative labor on the basis of the following drawings.
[0017] Figure 3 The present application is further illustrated by the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present application. For ordinary skilled in the art, other drawings can be obtained without creative labor on the basis of the following drawings. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] Referring to Figure 1 The first aspect of the present application provides a large language model-based oil knowledge question and answer method, which comprises: an oil knowledge question and answer platform collects original data in the oil field from various data sources, denoted as an original data set in the oil field, a data processing component performs availability preprocessing operation on the original data set in the oil field, determines the availability evaluation value of the original data set in the oil field, checks with the pre-defined availability evaluation threshold, determines the availability state of the original data set in the oil field, and the original data set in the oil field after availability preprocessing is denoted as an available data set in the oil field.
[0020] The above-mentioned oil knowledge Q&A platform is a digital platform focusing on knowledge exchange and answer in the field of oil, aiming to provide professional, accurate and convenient knowledge service for oil industry practitioners, researchers, students and people interested in oil knowledge. It relies on advanced information technology and rich oil field resources to realize the function of user asking and obtaining answers. Users can ask various oil-related questions on the platform, such as oil exploration technology, mining technology, refining process, market dynamics, industry policy, etc. The oil knowledge Q&A platform uses natural language processing technology to understand the user's problem intention, analyzes, semantically understands and classifies the natural language input by the user, and converts it into a form that the computer can understand, so as to accurately match the relevant knowledge and answers, and with the help of the powerful language understanding and generation ability of the large language model, the oil field knowledge is learned and inferred to generate high-quality answers. The data processing component is a key part of the oil knowledge Q&A platform, which is a tool for data processing in the oil knowledge Q&A platform, used to process, convert and optimize the collected oil field raw data to improve the usability and quality of the data.
[0021] The oil field raw data of each data source can be oil enterprise internal data such as production data, exploration data, management data, etc.; industry organization published data such as industry association data, etc.; academic research data such as academic journal paper data, research reports of colleges and research institutions, etc.; market data such as oil transaction data, market research agency data, etc.
[0022] The above-mentioned availability preprocessing operation specifically includes missing value processing: methods such as mean filling and median filling can be used for data perfection; error value correction: the data processing component can compare with historical data, associate with data of other related equipment, or make judgments and corrections according to physical principles, for example, in oil production, if the daily oil production data of an oil well suddenly fluctuates significantly and is inconsistent with the stable production trend in the past few months or years, further judgment is needed, for example, the daily oil production in a period of time is 100, 105, 101, 150, 106, 98, and the stable production interval in the past few months is [95, 110], then 150 indicates that the daily oil production data has a significant fluctuation, which may need to be corrected. Specifically, the slope, fluctuation range and other characteristics of the historical production curve can be compared to determine the anomaly; repeated value deletion: repeated records are identified and deleted by comparing key data fields such as oil product transaction time; data format unification: for example, date formats may have "year-month-day", "day / month / year" and other forms, which need to be unified into a standard format during preprocessing.
[0023] Specifically, the data processing component performs availability preprocessing operations on the original data set in the oil field, obtains preprocessing performance data of the data processing component, and obtains availability preprocessing operation data of the original data set in the oil field.
[0024] The preprocessing performance data of the data processing component specifically includes CPU occupancy of the data processing component at each preprocessing time point, memory occupancy of the data processing component at each preprocessing time point, preprocessing duration of the data processing component in a preprocessing period, and data processing amount of the data processing component in the preprocessing period.
[0025] The preprocessing time point is specifically a plurality of preprocessing time points obtained by dividing the preprocessing period by time points, and the division method can be 30 seconds. The preprocessing period is specifically a duration of the data processing component performing availability preprocessing operations on the original data set in the oil field. The preprocessing period is specifically a period of time from when the data processing component receives an availability preprocessing operation instruction to when the data processing component completes all original data operations and feeds back to the oil knowledge Q&A platform.
[0026] The preprocessing performance data can be extracted from the running report of the data processing component.
[0027] The CPU occupancy of the data processing component at each preprocessing time point is processed by averaging to obtain the CPU average occupancy of the data processing component in the preprocessing period.
[0028] The data processing amount of the data processing component in the preprocessing period is processed by ratio with the preprocessing duration of the data processing component in the preprocessing period to obtain the data processing throughput of the data processing component in the preprocessing period.
[0029] The CPU average occupancy limit value and the memory maximum occupancy reference value are extracted from the oil knowledge Q&A information base.
[0030] The CPU average occupancy of the data processing component in the preprocessing period, the memory occupancy of the data processing component at each preprocessing time point, the preprocessing duration of the data processing component in the preprocessing period, and the data processing throughput of the data processing component in the preprocessing period are comprehensively processed to obtain the data preprocessing performance index of the data processing component. The specific analysis method is as follows:
[0031] ;
[0032] In the formula, is the data preprocessing performance index of the data processing component, is the CPU average occupancy of the data processing component in the preprocessing period, is the CPU average occupancy limit value, is the memory occupation rate of the data processing component at the i-th pre-processing time point, i is the number of each pre-processing time point, M is the total amount of pre-processing time points, max represents the maximum value, is the maximum memory occupation rate reference value, is the pre-processing duration of the data processing component in the pre-processing period, is the data processing throughput of the data processing component in the pre-processing period, is the pre-processing performance influence parameter corresponding to the pre-processing duration predefined in the oil knowledge Q&A information base, is the pre-processing performance influence parameter corresponding to the data processing throughput predefined in the oil knowledge Q&A information base, and e is a natural constant.
[0033] The data preprocessing performance index of the data processing component in the embodiment is used to measure the efficiency and effect of the data processing component in the data preprocessing stage, such as data cleaning and transformation, and is also used to evaluate the ability and stability of the data preprocessing component in processing data.
[0034] The pre-processing performance influence parameter corresponding to the pre-processing duration and the pre-processing performance influence parameter corresponding to the data processing throughput are both extracted from the oil knowledge Q&A information base. For example, the pre-processing duration and the data processing throughput form a mapping set with the pre-processing performance influence parameter corresponding to the pre-processing duration and the pre-processing performance influence parameter corresponding to the data processing throughput predefined in the oil knowledge Q&A information base, respectively. The real-time pre-processing duration and the data processing throughput are brought into the mapping set to obtain the pre-processing performance influence parameter corresponding to the pre-processing duration and the pre-processing performance influence parameter corresponding to the data processing throughput.
[0035] In the embodiment, if the CPU average occupation rate is small within a suitable range, the pre-processing duration will be shortened, and the data preprocessing performance of the data processing component will be increased. When the CPU average occupation rate is too high and greater than the CPU defined occupation rate, the system performance may be reduced, and the phenomenon of lag may occur, which will prolong the pre-processing duration and have a negative impact on the data preprocessing performance of the data processing component. Similarly, appropriate memory occupation can improve the pre-processing speed and shorten the pre-processing duration. However, if the maximum memory occupation rate is too high and greater than the maximum memory occupation rate reference value, the system will frequently perform memory exchange, which will greatly reduce the data processing speed, prolong the pre-processing duration, and reduce the processing performance of the data processing component. The data processing throughput and the pre-processing duration are inversely proportional. The higher the data processing throughput, the more data is processed in a unit of time, and the shorter the time required to preprocess the same size of data, thereby increasing the processing performance of the data processing component.
[0036] The availability preprocessing operation data of the original data set in the oil field specifically includes the number of numerical missing values of each data group in the original data set in the oil field, the number of data garbled codes of the original data set in the oil field, the number of key elements of the original data set in the oil field, the total data amount of the original data set in the oil field, and the number of data formats of the original data set in the oil field. The above availability preprocessing operation data can be specifically extracted from the preprocessing execution report of the data processing component.
[0037] Further, the availability evaluation value of the original data set in the oil field is determined, and the specific determination process is as follows:
[0038] The number of key elements of the original data set in the oil field is subjected to ratio processing with the total data amount of the original data set in the oil field to obtain the key element proportion of the original data set in the oil field.
[0039] The number of numerical missing values of each data group in the original data set in the oil field is added to obtain the total number of numerical missing values of the original data set in the oil field. The number of data format references is extracted from the oil knowledge Q&A information base.
[0040] The total number of numerical missing values of the original data set in the oil field, the number of data garbled codes of the original data set in the oil field, the key element proportion of the original data set in the oil field, the number of data formats of the original data set in the oil field, and the data preprocessing performance index of the data processing component are comprehensively analyzed to obtain the availability evaluation value of the original data set in the oil field, and the specific analysis method is as follows:
[0041] ;
[0042] In the formula, is the availability evaluation value of the original data set in the oil field, is the total number of numerical missing values of the original data set in the oil field, is the number of data garbled codes of the original data set in the oil field, is the key element proportion of the original data set in the oil field, and is the number of data formats of the original data set in the oil field, is the number of data format references, is the data preprocessing performance index of the data processing component, is the availability influence parameter corresponding to the total number of numerical missing values predefined in the oil knowledge Q&A information base, is the availability influence parameter corresponding to the number of data garbled codes predefined in the oil knowledge Q&A information base, is the availability influence parameter corresponding to the key element proportion predefined in the oil knowledge Q&A information base, is the availability influence parameter corresponding to the data preprocessing performance index predefined in the oil knowledge Q&A information base, and e is a natural constant.
[0043] The availability evaluation value of the original data set in the oil field in this embodiment is a quantitative evaluation of the ability of the original data set in the oil field to meet the use requirements in subsequent oil field analysis, learning, etc. The higher the availability evaluation value, the more the initial data set can meet the needs of subsequent question answering.
[0044] It should be explained that the total number of numerical value missing refers to the total number of missing specific numerical values in each independent data group in the original data in the oil field; the number of data garbled refers to the number of data units that do not conform to the data format specification and cannot be normally recognized and analyzed in the original data in the oil field; the proportion of key elements refers to the proportion of the number of data elements that are of key importance to oil knowledge question answering, analysis, etc. in the original data set in the oil field, to the total number of elements in the data set, wherein the key elements can be oil exploration key elements such as formation lithology data, seismic wave reflection characteristic data, oil reservoir porosity and permeability data, oil production key elements such as oil well production data, working parameters of production equipment (such as stroke and number of strokes of a pumping unit), wellhead pressure data, etc. The non-key elements can be environmental data, management data, etc. in the original data set in the oil field that are not directly used for oil exploration, development or production.
[0045] The availability influence parameter corresponding to the total number of numerical value missing, the availability influence parameter corresponding to the number of data garbled, the availability influence parameter corresponding to the proportion of key elements, and the availability influence parameter corresponding to the data preprocessing performance index are all extracted in the oil knowledge question answering information library, for example, the total number of numerical value missing, the number of data garbled, the proportion of key elements, and the data preprocessing performance index are respectively mapped with the preset availability influence parameter corresponding to the total number of numerical value missing, the availability influence parameter corresponding to the number of data garbled, the availability influence parameter corresponding to the proportion of key elements, and the availability influence parameter corresponding to the data preprocessing performance index in the oil knowledge question answering information library to form a mapping set. The total number of numerical value missing, the number of data garbled, the proportion of key elements, and the data preprocessing performance index are brought into the mapping set to obtain the availability influence parameter corresponding to the total number of numerical value missing, the availability influence parameter corresponding to the number of data garbled, the availability influence parameter corresponding to the proportion of key elements, and the availability influence parameter corresponding to the data preprocessing performance index.
[0046] In this embodiment, if a large number of missing values exist in the key elements, it will directly affect the effectiveness of the key element proportion. When there are a large number of missing values in the key elements, the availability of these key elements for analysis or question answering is reduced, resulting in a decrease in the availability evaluation value of the original data set in the oil field. The more the number of data formats, the higher the risk of data garbled code. Different data formats have different encoding methods, storage rules and parsing requirements. When the data processing component processes data in multiple formats, it may not be familiar with some special formats or make mistakes during format conversion, resulting in data garbled code. Similarly, too many data formats may increase the total number of missing values, because different data formats may have different integrity requirements and data recording specifications. When processing data in multiple formats, some data may not be recorded correctly, thereby increasing the total number of missing values. Therefore, the number of data formats should be within an appropriate range to reduce the total number of missing values and data garbled codes, thereby improving the availability of the original data set in the oil field.
[0047] In this embodiment, the larger the data preprocessing performance index of the data processing component, the more effectively the missing values in the data set can be processed. The availability evaluation value increases due to the improvement of data integrity. At the same time, a high-performance data processing component performs well in error value recognition rate and correction accuracy, can accurately identify and correct error data, reduce the interference of error data on subsequent analysis, thereby improving the accuracy of the data. In addition, a good performance data processing component performs well in data standardization effect evaluation and normalization range compliance rate, and can compare and analyze different orders of magnitude of data (such as oil composition content, equipment operating parameters, etc.) in the original data set in the oil field on the same level, thereby improving the availability evaluation value of the data.
[0048] Specifically, the method for determining the availability state of the original data set in the oil field comprises: performing availability state verification on the availability evaluation value of the original data set in the oil field and a pre-defined availability evaluation threshold value to obtain an availability state verification result, and determining the availability state of the original data set in the oil field according to the availability state verification result.
[0049] The availability state verification result is a first availability state verification result or a second availability state verification result. The first availability state verification result is that the availability evaluation value of the original data set in the oil field is greater than or equal to the pre-defined availability evaluation threshold value in the oil knowledge question and answer information base. The second availability state verification result is that the availability evaluation value of the original data set in the oil field is less than the availability evaluation threshold value.
[0050] If the available state verification result is the first available state verification result, it is determined that the oil field original data set is in an available state, and if the available state verification result is the second available state verification result, it is determined that the oil field original data set is in an unavailable state, and data optimization is performed on the oil field original data set.
[0051] The above data optimization of the oil field original data set can be specifically that when the oil field original data set is in an unavailable state, the distribution of the data is observed, if there are a large number of blanks or obvious inconsistencies with the expected distribution on some variables, there may be missing values, then a multiple imputation method can be used, which is not simply filled with mean or median, but generates reasonable imputation values based on the distribution of the data and the relationship between variables; similar data records are compared one by one, for example, in the oil equipment procurement data, records with the same procurement time, equipment model, supplier and procurement amount are compared, which may be repeated, then a variety of duplicate value identification methods are used to comprehensively check and remove duplicate values; check if the values of key data fields have changed or been lost, for example, after converting the oil equipment maintenance records, if the values of key fields such as equipment model, maintenance time and maintenance content are incorrect or missing, it indicates that there is a format conversion problem, then a format conversion verification mechanism is used to check the integrity and accuracy of the data after each format conversion, to ensure that there is no data loss or distortion.
[0052] In this embodiment, it needs to be explained that the data in the oil field often has high complexity, a single parameter often cannot fully reflect the true situation of the data element, and there is usually a relationship between the variables in the oil data, if only looking at a single parameter of production, some internal relationships may be ignored, by considering multiple parameters comprehensively, the rationality of the data element can be more accurately grasped, the comprehensive formula can combine multiple related parameters to more comprehensively evaluate the data element, so as to more accurately judge whether the current production data is reasonable, avoid misjudgment caused by single parameter judgment, and improve the accuracy of the data.
[0053] The oil knowledge Q&A platform constructs the oil field Q&A pair set according to the pre-defined construction method according to the available oil field data set, and performs calibration operation to determine the application quality index of each oil field Q&A pair, and compares it with the pre-defined application quality reference index to screen the oil field Q&A pair data set.
[0054] The above constructing the oil field Q&A pair set according to the pre-defined construction method is specifically that the construction process is: first, constructing the oil field knowledge graph according to the available oil field data set, and second, constructing the oil field Q&A pair set according to the oil field knowledge graph.
[0055] The oil field knowledge graph is constructed, specifically, the key elements in the oil field such as oil resources, oil exploration and development, oil products, etc. are extracted from the available data set in the oil field as nodes; the elements representing the relationship between entities such as exploration, production, supply, etc. are extracted from the available data set in the oil field as the edges connecting the nodes; the key elements that may be attributes are extracted from the available data set in the oil field, and according to a plurality of named entities and a plurality of entity relationships, a plurality of knowledge triples are constructed, and according to the plurality of knowledge triples, a corresponding oil field knowledge graph is constructed, wherein the entity refers to the node in the knowledge graph, which represents the key element in the oil field, and the relationship refers to the connection and interaction between entities; extraction can be carried out by natural language processing technology for text mining, for example, using a keyword search algorithm, searching for keywords related to "oil resources" (such as "reserves", "distribution", "unconventional oil", etc.) in the available data set in the oil field, extracting data containing these keywords, and organizing and extracting them.
[0056] Specifically, the construction of the oil field knowledge graph is as shown in Figure 3 Figure 3 is a structural diagram of the knowledge graph, and Figure 3 It can be seen that the oil field knowledge graph is specifically composed of entities, attributes and relationships, through this structure, the oil knowledge question and answer platform can clearly define the data objects that need to be stored in the available data set in the oil field, the characteristics of these objects and their mutual relationship, thereby providing a data basis for the construction of the oil field question and answer set.
[0057] According to the oil field knowledge graph, the oil field question and answer set is constructed, first, the potential question and answer pair generation object is determined according to the entity and relationship in the knowledge graph, the question type is set, for example, the question: what are the main methods of oil exploration? The answer: the main methods of oil exploration include seismic exploration, electromagnetic exploration, drilling sampling, etc. Seismic exploration analyzes the reflection wave of underground rock stratum to infer the location of oil and gas reservoir, while drilling confirms the existence of oil and gas reservoir through actual drilling; the set question is used as a template to automatically generate and optimize the question by GPT-4, and the question and answer pair is obtained by using Rag.
[0058] Rag (Retrieval Augmentation Technology) is a technology that combines information retrieval with generation models, the core idea of Rag is to retrieve relevant information in a large text library, and then combine the retrieved information with a generation model to generate more accurate and contextually relevant answers.
[0059] According to the existing GPT-4 generated question, the answer is given by Rag to get the question and answer pair, the basic workflow of Rag is as follows: 1. Retrieval system - this system can retrieve relevant document fragments from the document library of the oil knowledge Q&A platform, using BM25 retrieval algorithm to retrieve relevant document fragments according to query matching; 2. Define query and retrieval - for a given question (query), the retrieval system will find the most relevant documents or paragraphs in the knowledge base, which will be used as input for the RAG model to generate more accurate answers; 3. Use Rag to generate answers - RAG model will retrieve documents and questions as input, pass into a GPT-4, generate answers according to RAG-Sequence or RAG-Token.
[0060] Example 1:
[0061] 1) User asks: What is the chemical formula of water?
[0062] 2) Generate query and perform information retrieval: The model reformulates "What is the chemical formula of water?" as a retrieval query, such as "water chemical formula". Retrieve the paragraph from the chemistry basics entry: "Water is composed of hydrogen and oxygen elements in a ratio of the number of atoms present in the molecule, and its chemical formula is H2O."
[0063] 3) Generate answer: Combine the retrieval content with the question to generate a complete answer, for example: "The chemical formula of water is H2O."
[0064] Example 2:
[0065] 1) User asks: What is the commonly used approximation of pi?
[0066] 2) Generate query and perform information retrieval: The model reformulates the question as "pi approximation", retrieves the paragraph: "Pi (π) is the ratio of the circumference of a circle to its diameter, and the commonly used approximation is 3.14159."
[0067] 3) Generate answer: For example: "The commonly used approximation of pi is 3.14159."
[0068] Example 3:
[0069] 1) User asks: What is the name of the natural satellite of the Earth?
[0070] 2) Generate query and perform information retrieval: Reformulate as "Earth natural satellite name", retrieve the paragraph: "The Earth has a natural satellite, commonly known as the 'Moon'."
[0071] 3) Generate answer: For example: "The natural satellite of the Earth is called the Moon."
[0072] The calibration operation, in the embodiment, is specifically performed by the oil knowledge Q&A platform on the oil field Q&A pairs. For example, the oil field Q&A pairs are subjected to semantic detection by using a semantic detection tool, and the Q&A pairs are subjected to re-calibration by using a manual evaluation method. The evaluation criteria usually involve evaluation of the accuracy, relevance, fluency and integrity of the Q&A pairs. The Q&A pairs with problems are deleted to obtain the oil field Q&A pair dataset.
[0073] Further, the application quality index of each oil field Q&A pair is determined, and the specific determination process is as follows:
[0074] The calibration operation data of the oil knowledge Q&A platform are obtained, specifically including the calibration duration of each oil field Q&A pair and the number of semantic error sentences of each oil field Q&A pair. The calibration operation data can be obtained in the calibration execution report of the oil knowledge Q&A platform.
[0075] The application data of each oil field Q&A pair are obtained, specifically including the number of literature citations of each oil field Q&A pair, the answer coding of each oil field Q&A pair and the knowledge update interval duration of each oil field Q&A pair. The application data can be obtained in the search report of the oil knowledge Q&A platform.
[0076] The number of literature citation adaptations of each oil field Q&A pair and the answer reference coding of each oil field Q&A pair are obtained from the oil knowledge Q&A information base. The answer coding of each oil field Q&A pair and the answer reference coding of each oil field Q&A pair are subjected to data processing to obtain and record the answer coding similarity of each oil field Q&A pair.
[0077] The data processing can be specifically performed by calculating the answer coding similarity by using a hash algorithm. In a specific embodiment, it is assumed that the answer coding of an oil field Q&A pair is The coding sequence is converted into MD5 (hash function) as Similarly, for the answer reference coding The coding sequence is converted into MD5 hash function as The two hash values are compared bit by bit. For the 128-bit MD5 hash value, the comparison starts from the first bit. and The number of same bits is counted. It is assumed that the number of same bits is K. The similarity can be represented as wherein H represents the hash value, and 128 represents the fixed number of MD5 bits.
[0078] The answer coding similarity of each oil field question and answer pair, the calibration duration of each oil field question and answer pair, the number of semantic error sentences of each oil field question and answer pair, the number of literature references of each oil field question and answer pair, the knowledge update interval duration of each oil field question and answer pair, and the availability evaluation value of the oil field original data set are comprehensively analyzed to obtain the application quality index of each oil field question and answer pair. The specific analysis method is as follows:
[0079] ;
[0080] In the formula, is the application quality index of the gth oil field question and answer pair, g is the number of each oil field question and answer pair, , G is the total amount of oil field question and answer pairs, is the answer coding similarity of the gth oil field question and answer pair, is the calibration duration of the gth oil field question and answer pair, is the number of semantic error sentences of the gth oil field question and answer pair, is the number of literature references of the gth oil field question and answer pair, is the number of literature reference adaptations of the gth oil field question and answer pair, is the knowledge update interval duration of the gth oil field question and answer pair, is the availability evaluation value of the oil field original data set, is the application quality impact parameter corresponding to the pre-defined answer coding similarity in the oil knowledge question and answer information library, is the application quality impact parameter corresponding to the pre-defined calibration duration in the oil knowledge question and answer information library, is the application quality impact parameter corresponding to the pre-defined number of semantic error sentences in the oil knowledge question and answer information library, is the application quality impact parameter corresponding to the pre-defined knowledge update interval duration in the oil knowledge question and answer information library, is the application quality impact parameter corresponding to the pre-defined availability evaluation value in the oil knowledge question and answer information library, and e is a natural constant.
[0081] In this embodiment, the application quality index of each oil field question and answer pair is used to measure the quality level of each group of question and answer pairs in the oil field in actual application. The higher the application quality index, the higher the answer quality of the question and answer pair in actual application, and the stronger the accuracy.
[0082] It needs to be explained that the above-mentioned answer coding similarity refers to the similarity between the coding obtained after coding processing of the answer of the question and answer pair in the oil field and the reference coding of the answer of the question and answer pair; the number of semantic error sentences refers to the number of sentences that are incorrect due to inaccurate semantic expression, do not conform to logic, violate oil professional knowledge or have ambiguity, etc. in the question and answer sentences in the oil field, for example, the question is "what should be paid attention to in the maintenance of oil pipeline", and the answer is "pay attention to that thing to avoid problems", where "that thing" is semantically ambiguous and unclear about what it specifically refers to, which may be corrosion of the pipeline, pressure change or other factors, such sentences are semantic ambiguity error sentences.
[0083] The application quality influence parameter corresponding to the answer coding similarity, the application quality influence parameter corresponding to the calibration time length, the application quality influence parameter corresponding to the number of semantic error sentences, the application quality influence parameter corresponding to the knowledge update interval time length and the application quality influence parameter corresponding to the availability evaluation value are all extracted from the oil knowledge question and answer information library, for example, the answer coding similarity, the calibration time length, the number of semantic error sentences, the knowledge update interval time length and the availability evaluation value are respectively mapped to the application quality influence parameter corresponding to the preset answer coding similarity of the oil knowledge question and answer information library, the application quality influence parameter corresponding to the calibration time length, the application quality influence parameter corresponding to the number of semantic error sentences, the application quality influence parameter corresponding to the knowledge update interval time length and the application quality influence parameter corresponding to the availability evaluation value, and the real-time answer coding similarity, the calibration time length, the number of semantic error sentences, the knowledge update interval time length and the availability evaluation value are brought into the mapping set to obtain the application quality influence parameter corresponding to the answer coding similarity, the application quality influence parameter corresponding to the calibration time length, the application quality influence parameter corresponding to the number of semantic error sentences, the application quality influence parameter corresponding to the knowledge update interval time length and the application quality influence parameter corresponding to the availability evaluation value.
[0084] In this embodiment, the high similarity of answer encoding generally means that the answer of the question and answer pair is consistent with the reference standard in content structure and expression of key knowledge points, and the number of semantic error sentences may be relatively small, thereby improving the application quality index of the question and answer pair in each oil field. On the contrary, a large number of semantic error sentences will lead to a large difference between the answer encoding and the reference encoding, thereby reducing the similarity of the answer encoding, and negatively affecting the application quality of the question and answer pair in each oil field. If the number of semantic error sentences increases, the calibration time will increase because more time is needed to identify and correct these errors, and the accuracy of the answer needs to be rechecked. When the knowledge update interval is too long, a large amount of knowledge needs to be updated and errors need to be corrected, thereby increasing the calibration time and reducing the application quality of the question and answer pair in each oil field. Reasonable literature references can reduce the number of semantic error sentences because the referenced literature has usually been strictly reviewed and verified, and can provide accurate knowledge support for the answer. However, when the number of literature references is too large, the readability of the content of the question and answer pair may be reduced, the focus may not be highlighted, and the application quality of the question and answer pair may be reduced. Too few literature references will reduce the credibility and timeliness of the question and answer pair, and also reduce the application quality of the question and answer pair.
[0085] In this embodiment, the greater the availability evaluation value of the oil field original data set, the higher the accuracy of the data, which can provide more accurate data support for the question and answer pair when constructing the question and answer pair in each oil field, thereby improving the application quality index. The data set with a large availability evaluation value has better integrity, which enables the provision of more comprehensive information when answering questions in the oil field, better meets the user's demand for comprehensive knowledge acquisition, and further improves the application quality index. At the same time, the high-availability original data set can ensure the timeliness of the data and provide the user with the latest knowledge to improve the application quality index.
[0086] Specifically, the screening obtains the oil field question and answer pair data set, and the specific analysis process is as follows:
[0087] If the application quality index of each oil field question and answer pair is greater than or equal to the application quality reference index predefined in the oil knowledge question and answer information library, the oil field question and answer pair is recorded as the oil field question and answer pair data set. If the application quality index of a certain oil field question and answer pair is less than the application quality reference index, the oil field question and answer pair is deleted, and the remaining oil field question and answer pair is recorded as the oil field question and answer pair data set.
[0088] The petroleum knowledge Q&A platform stores the petroleum field Q&A pair dataset in the large language model to which the petroleum Q&A component belongs, and issues fine-tuning instructions. The petroleum Q&A component fine-tunes the large language model according to the pre-defined fine-tuning manner, evaluates the Q&A output effect index of the large language model, and verifies the pre-defined Q&A output effect index. The petroleum knowledge Q&A platform determines whether to perform repeated fine-tuning operation on the large language model, thereby completing the petroleum knowledge Q&A based on the large language model.
[0089] The petroleum Q&A component is a key part of the petroleum knowledge Q&A platform, which is a module specially used for processing petroleum field Q&A related tasks. The component is mainly based on natural language processing technology and related petroleum field knowledge, and is responsible for receiving user-proposed petroleum field problems, searching and matching in the stored Q&A pair dataset, and then generating and returning appropriate answers. The petroleum Q&A component is optimized and customized for the petroleum field based on the large language model, making the Q&A service more professional and targeted, and better utilizing the petroleum field Q&A pair dataset to answer user questions.
[0090] The petroleum Q&A component fine-tunes the large language model according to the pre-defined fine-tuning manner. Specifically, the large language model to which the petroleum Q&A component belongs rewrites the user's input question according to the constructed knowledge graph through LoRA fine-tuning and converts it into natural language output using NLG capability.
[0091] Rewriting the user's input question is to optimize the question through the structure model of the knowledge graph, and to make the answer of the large language model more professional and comprehensive through reasoning of the knowledge graph. In a specific embodiment, when the user inputs "When was Apple Inc. founded?", based on the "Apple Inc." entity and its "founding time" attribute in the knowledge graph, the model can rewrite the question as "In what year was Apple Inc. founded?" or provide related supplementary questions such as "Who is the founder of Apple Inc.?" etc.
[0092] The rewriting of the question by the large language model is actually to structure and semantize the information in natural language, and then use the knowledge of entities, relationships, and attributes in the graph to understand and reconstruct the question. In simple terms, question rewriting is the process of the large model understanding the user's input question, and finding the optimal answer in the vast knowledge base and returning it to the user through NLP technology. This usually includes the following steps: 1. Analyze the meaning of the question; 2. Introduce the knowledge graph and map the parsed semantics to structured knowledge; 3. Rewrite the question according to the reasoning of the graph to generate a more accurate question; 4. Generate an answer and return the accurate result; 5. Convert it into natural language output with NLP technology.
[0093] Further, the evaluation of the large language model question and answer output effect index, the specific evaluation process is:
[0094] Obtaining the question and answer output data of the large language model, specifically including the number of knowledge point coverage of the large language model under each question and answer output, the amount of additional information provided by the large language model under each question and answer output, the answer output time length of the large language model under each question and answer output, and the number of user praise times of the large language model within the question and answer output period.
[0095] Each question and answer output above refers to a number of question and answer output processes within the question and answer output period. The determination of the question and answer output period is obtained by relevant management personnel according to the state of the oil question and answer component, the actual output demand, the specific knowledge problem and other factors. The above question and answer output data can be extracted from the output report of the oil question and answer component.
[0096] Obtaining the total amount of answer information of the large language model under each question and answer output, wherein the total amount of answer information can be extracted from the output report of the oil question and answer component. The amount of additional information provided by the large language model under each question and answer output is compared with the total amount of answer information of the large language model under each question and answer output to obtain the proportion of additional information of the large language model under each question and answer output.
[0097] The number of knowledge point coverage of the large language model under each question and answer output is compared with the pre-defined number of knowledge point coverage of each question and answer output in the oil knowledge question and answer information library to obtain the knowledge point coverage degree value of the large language model under each question and answer output. The additional information reference proportion under each question and answer output, the answer output reference time length under each question and answer output, and the user praise definition times are extracted from the oil knowledge question and answer information library.
[0098] The knowledge point coverage degree value of the large language model under each question and answer output, the proportion of additional information of the large language model under each question and answer output, the answer output time length of the large language model under each question and answer output, the number of user praise times of the large language model within the question and answer output period, the availability evaluation value of the oil field original data set, and the application quality index of each oil field question and answer pair are comprehensively analyzed to obtain the question and answer output effect index of the large language model, and the specific analysis method is as follows:
[0099] ;
[0100] In the formula, is the question and answer output effect index of the large language model, is the knowledge point coverage degree value of the large language model under the dth question and answer output, d is the number of each question and answer output, , D is the total amount of question and answer output, The proportion of additional information in the dth question and answer output of the large language model, The reference proportion of additional information in the dth question and answer output, The answer output time length of the large language model in the dth question and answer output, The reference answer output time length of the dth question and answer output, The number of user praises of the large language model in the question and answer output period, The number of user praise definitions, The application quality index of the gth petroleum field question and answer pair, g is the number of each petroleum field question and answer pair, G is the total number of petroleum field question and answer pairs, The availability evaluation value of the petroleum field original data set, The output effect influence parameter corresponding to the pre-defined knowledge point coverage value in the petroleum knowledge question and answer information base, The output effect influence parameter corresponding to the pre-defined application quality index in the petroleum knowledge question and answer information base, The output effect influence parameter corresponding to the pre-defined availability evaluation value in the petroleum knowledge question and answer information base, e is a natural constant.
[0101] The question and answer output effect index of the large language model in this embodiment is used to measure the performance and quality of the large language model in answering questions, helping to evaluate whether the model can accurately, completely, clearly and reasonably answer various questions raised by users, especially in the application scenario of petroleum knowledge question and answer in a specific field, reflecting the practicality and reliability of the model output answer.
[0102] It needs to be explained that the above-mentioned knowledge point coverage value refers to the proportional relationship between the petroleum field knowledge points involved in the answers output by the large language model in each question and answer output and the number of pre-defined adaptation-related knowledge points involved in the answers, wherein the knowledge points, such as the principle of seismic exploration technology, include the generation, propagation and reflection law of seismic waves, and how to infer the underground geological structure by receiving reflected waves; there are also the principles of logging technology, such as how different logging methods such as electrical logging and acoustic logging obtain formation information, when answering questions such as "how does seismic exploration find oil", these technical principles are knowledge points, such as the principles, application conditions and advantages and disadvantages of different exploitation methods such as self-flowing oil production, mechanical oil production (rod pump oil production, rodless pump oil production), when answering the question "in what circumstances should self-flowing oil production be selected", the related knowledge of exploitation method is the knowledge point.
[0103] The extra information proportion refers to the proportion of other information in addition to the core information required for directly answering the question in the output answer of each question and answer output of the large language model. The extra information may include relevant background knowledge, supplementary explanation, example, or other related but unnecessary content. The extra information reference proportion refers to the adaptive value corresponding to the extra information proportion of each question and answer output. The answer output duration refers to the time span from when the large language model receives the user's question to when the complete answer is output. This time span covers the entire process of the model's understanding of the question, retrieving relevant information in the knowledge base, generating an answer, and presenting the answer.
[0104] The output effect influence parameter corresponding to the knowledge point coverage degree value, the output effect influence parameter corresponding to the availability evaluation value, and the output effect influence parameter corresponding to the application quality index are extracted from the oil knowledge question and answer information library. For example, the knowledge point coverage degree value, the availability evaluation value, and the application quality index form a mapping set with the output effect influence parameter corresponding to the preset knowledge point coverage degree value in the oil knowledge question and answer information library, the output effect influence parameter corresponding to the availability evaluation value, and the output effect influence parameter corresponding to the application quality index. The real-time knowledge point coverage degree value, the availability evaluation value, and the application quality index are brought into the mapping set to obtain the output effect influence parameter corresponding to the knowledge point coverage degree value, the output effect influence parameter corresponding to the availability evaluation value, and the output effect influence parameter corresponding to the application quality index.
[0105] In this embodiment, when the knowledge point coverage degree value of the large language model in each question and answer output is high within a reasonable range, the comprehensive answer to the user's question can be met, thereby improving the number of user good comments and improving the question and answer output effect index of the large language model. Appropriate extra information can enrich the answer content and help users better understand the core knowledge, which can also improve the number of user good comments. However, when the extra information proportion is too high and greater than the extra information reference proportion, it will lead to negative problems such as unclear answer focus and information redundancy, thereby reducing the question and answer output effect of the large language model. At the same time, too high extra information reference proportion will generally lead to longer answer output duration. If the answer output duration is too long and deviates from the answer output reference duration, it will lead to a decrease in user good comments, thereby greatly reducing the question and answer output effect of the large language model.
[0106] The greater the availability evaluation value of the original data set in the oil field in this embodiment and the greater the application quality index of each oil field question and answer pair, the higher the accuracy of the original data in the oil field, the less the interference of the error data, and the more the output answer of the model conforms to the actual situation, thereby improving the accuracy in the question and answer output effect index; a good data set is also excellent in data consistency. Unified data can ensure that the output answer of the model is clear and accurate, and will not mislead users due to the confusion of units or definitions, which helps to improve the clarity in the question and answer output effect index. In addition, a high application quality index means that the accuracy of the question and answer pair itself is higher. When the large language model refers to these high-quality question and answer pairs to generate answers, it can more accurately answer the user's questions, also improving the accuracy in the question and answer output effect index.
[0107] Specifically, the oil knowledge question and answer platform determines whether to perform repeated fine-tuning operation on the large language model. Specifically, the oil question and answer component performs output effect verification on the question and answer output effect index of the large language model and the pre-defined question and answer output effect index, obtains an output effect verification result and uploads it to the oil knowledge question and answer platform. The oil knowledge question and answer platform determines whether to perform repeated fine-tuning operation on the large language model.
[0108] The output effect verification result is a first output effect verification result or a second output effect verification result. The first output effect verification result is specifically that the question and answer output effect index of the large language model is greater than or equal to the pre-defined question and answer output effect index in the oil knowledge question and answer information base. The second output effect verification result is specifically that the question and answer output effect index of the large language model is less than the question and answer output effect index. If the output effect verification result is the first output effect verification result, the oil knowledge question and answer platform determines that the large language model does not need to be fine-tuned. If the output effect verification result is the second output effect verification result, the oil knowledge question and answer platform determines that the large language model needs to be fine-tuned, and issues a repeated fine-tuning instruction. The oil question and answer component fine-tunes the large language model according to the pre-defined repeated fine-tuning manner.
[0109] Further, the pre-defined repeated fine-tuning manner is specifically that the oil question and answer component uses the built-in LoRA method to fine-tune the large language model in the oil knowledge field.
[0110] The LoRA method is a technology that introduces an adaptive matrix into the large language model. The repeated fine-tuning operation is to keep the weights of the original model in the oil question and answer component unchanged, add a low-rank matrix to adjust the output of the original model, and make the original model become a large language model in the oil knowledge field, thereby completing the repeated fine-tuning operation on the large language model.
[0111] In simple terms, it is to embed the above oil field data set into the original model of the oil question and answer component, change the parameter weight of the oil field part in the original model, and make the original large language model more proficient in the knowledge of the oil related field. The essence of model modification is the modification of model parameters - parameters (in natural language processing, it can be understood as converting user's questions into a set of numbers according to a specific method, and usually in the form of a matrix), the essence of fine tuning is to change the original parameters of the model from one state to another, so that the model is more proficient in a certain field, and LoRA fine tuning has lower storage requirements than full fine tuning, and limited computing resources can also be adapted.
[0112] Based on the embodiments of the present application, the oil knowledge question and answer system can also improve the operation instruction analysis accuracy of the oil pumping equipment, dynamic risk assessment and real-time adaptive control. The specific examples are shown as follows:
[0113] The oil knowledge question and answer system adopts a full-process design of multi-modal input-layered processing-dynamic execution-feedback iteration, including:
[0114] (1) Multi-modal command input and preprocessing module
[0115] Input collection:
[0116] Receive the instructions input by the operator through voice (supporting dialect) or keyboard (such as "adjust the pumping unit stroke to 5 times / minute, that is, the well with high downhole pressure").
[0117] Convert to text and remove redundancy:
[0118] Convert voice to text through Coze audio-to-text plug-in, call fine-tuned LLM (oil knowledge question and answer system of the embodiment) to filter redundant information (such as "that is" and "a little"), and retain core actions ("adjust stroke"), objects ("pumping unit"), parameters ("5 times / minute" and "high downhole pressure").
[0119] Feature extraction:
[0120] Generate a four-tuple of "action-object-parameter-working condition" (such as "adjust-pumping unit-stroke 5 times / minute-high downhole pressure"), which provides structured data for subsequent processing.
[0121] (2) Oil field specialized command generation module
[0122] Professional instruction conversion:
[0123] Input the four-tuple into the oil knowledge question and answer system of the embodiment, and convert it into a standardized control language (such as "pumping unit stroke 3 meters, stroke 5 times / minute, based on downhole pressure 1.2 MPa dynamic adjustment").
[0124] Instruction splitting and verification:
[0125] Splitting steps by logic (e.g., "1. Detect current pumping rate; 2. Adjust to 5 times per minute; 3. Monitor pressure changes in real-time"), supporting operators to modify instructions through voice or keyboard until meeting requirements.
[0126] (3) Code generation and multi-layer risk assessment module
[0127] Code generation:
[0128] Call Coze code generation plug-in to convert split instructions into executable code for embedded systems of oil pumping equipment (e.g., C++ scripts).
[0129] Expert model evaluation
[0130] Expert model 1 (safety rule verification):
[0131] Based on the oil engineering rule library (e.g., "sand content > 3% when pumping rate ≤ 4 times per minute") and real-time parameters of equipment (e.g., current sand content 3.5%), evaluate the risk of code execution, and trigger a warning and correction if the pumping rate setting violates the rules.
[0132] Expert model 2 (parameter rationality verification):
[0133] Check parameter matching in code (e.g., "whether stroke 3 meters and pumping rate 5 times per minute are suitable for pumping rod strength"), output optimization suggestions (e.g., "adjust pumping rate to 4 times per minute to match stroke").
[0134] Expert model 3 (syntax and logic verification):
[0135] Verify the correctness of code syntax and the coherence of execution logic to ensure no program errors.
[0136] (4) Control optimization and execution module
[0137] Dynamic parameter optimization:
[0138] Combine real-time data from downhole sensors (temperature, pressure, sand content) to generate adaptive control parameters (e.g., "when pressure rises to 1.5 MPa, automatically reduce pumping rate by 10%") through the oil knowledge Q&A system of the embodiment.
[0139] Execution and emergency stop mechanism:
[0140] Optimized code is transmitted to the embedded end for execution, starting the emergency stop trigger mechanism - when the sensor detects sudden abnormalities (e.g., sudden increase in pumping rod load, wellhead leakage), immediately terminate execution and trigger an alarm.
[0141] (5) Data feedback and model iteration module
[0142] Data collection and processing:
[0143] The embedded terminal collects device operation data (stroke, stroke frequency, energy consumption) and environmental data (downhole pressure, temperature) in real time, compresses them through the Coze data packaging plug-in, and uploads them to the cloud database.
[0144] Petroleum knowledge fusion feedback:
[0145] Compare the actual execution effect with the expected target (such as whether adjusting the stroke frequency reduces the risk of pump sticking), combine the petroleum domain knowledge graph (such as "the relationship between pumping rod wear and stroke frequency"), and generate model iteration direction to continuously optimize the instruction analysis and decision-making ability of the petroleum knowledge question and answer system in this embodiment.
[0146] It needs to be explained that in one specific embodiment, the petroleum domain LLM fine-tuning method of the present application based on the above embodiment can be as follows:
[0147] (1) Knowledge graph construction
[0148] Entity and relationship extraction:
[0149] Extract entities (such as "pumping unit", "stroke", "sand content") and relationships (such as "stroke frequency affects pumping efficiency" and "high sand content causes pump sticking") from petroleum engineering literature and pumping equipment manuals to construct knowledge triples.
[0150] Graph application:
[0151] Embed the knowledge graph into the LLM to enable the model to understand professional associations such as "reducing stroke frequency can reduce pumping rod wear" and improve the professionalism of instruction analysis.
[0152] (2) Question and answer pair generation and calibration
[0153] Knowledge graph-based question and answer pair design: generate professional petroleum domain question and answer pairs based on the knowledge graph, such as:
[0154] Question: "When the sand content exceeds 3%, how should the pumping unit stroke frequency be adjusted?"
[0155] Answer: "When the sand content exceeds 3%, the stroke frequency should be reduced to less than 4 times per minute to reduce the wear of the pumping rod and oil pipe and reduce the risk of pump sticking."
[0156] RAG enhanced generation: use the BM25 algorithm to search for petroleum industry standards, fault cases, and other documents, combine GPT-4o to generate high-precision question and answer pairs, and after manual calibration (accuracy and professionalism verification), form a fine-tuning data set.
[0157] (3) LLM fine-tuning implementation
[0158] Base model selection: Based on ChatGLM3-6B model, using LoRA (Low Rank Adaptation) method for efficient fine-tuning of parameters, only adjusting the model parameters related to the oil field, reducing the demand for computing resources.
[0159] Fine-tuning target: Make the model have the ability to analyze oil professional instructions, associate equipment parameters and downhole working conditions, and generate reasonable control logic.
[0160] 3. Question rewriting and intelligent decision-making
[0161] Question rewriting mechanism: When the user inputs a vague instruction (such as "this well is not pumping"), the model analyzes the semantics in combination with the knowledge graph and rewrites it into a precise question (such as "Is it due to high sand content that causes the pump to be stuck? Do you need to adjust the stroke rate?").
[0162] NLG output: Through natural language generation technology, the device state and decision-making suggestions are converted into colloquial language (such as "current sand content is 4%, suggest reducing the stroke rate from 5 times / min to 3 times / min to avoid pump sticking"), which is easy for operators to understand.
[0163] Reference Figure 2 The second aspect of the present application provides an oil knowledge question and answer system based on a large language model, comprising: a data availability preprocessing module, a question and answer pair application quality determination module, and a model output effect determination module.
[0164] The second aspect of the present application provides an oil knowledge question and answer system based on a large language model, further comprising: an oil knowledge question and answer information base for storing CPU average occupancy rate adaptation value, memory maximum occupancy rate reference value, data format reference number, literature reference adaptation number of each oil field question and answer pair, answer reference encoding of each oil field question and answer pair, additional information reference proportion under each question and answer output, answer output reference time under each question and answer output, user praise definition times, and pre-set values of various factors.
[0165] The data availability preprocessing module is connected to the question and answer pair application quality determination module, the question and answer pair application quality determination module is connected to the model output effect determination module, and the data availability preprocessing module, the question and answer pair application quality determination module and the model output effect determination module are all connected to the oil knowledge question and answer information base.
[0166] The data availability preprocessing module is used for the oil knowledge question and answer platform to collect original data of the oil field from each data source, denoted as an original data set of the oil field, and a data processing component performs an availability preprocessing operation on the original data set of the oil field, determines an availability evaluation value of the original data set of the oil field, checks the availability evaluation value with a predefined availability evaluation threshold, determines an availability state of the original data set of the oil field, and records the original data set of the oil field after the availability preprocessing operation as an available data set of the oil field.
[0167] The question and answer pair application quality module is used for the oil knowledge question and answer platform to construct a question and answer pair set of the oil field from the available data set of the oil field according to a predefined construction manner, and performs a calibration operation to determine an application quality index of each question and answer pair of the oil field, compare the application quality index with a predefined application quality reference index, and screen to obtain a question and answer pair data set of the oil field.
[0168] The model output effect module is used for the oil knowledge question and answer platform to store the question and answer pair data set of the oil field in a large language model to which the oil question and answer component belongs, and the oil knowledge question and answer platform issues a fine-tuning instruction, so that the oil question and answer component performs a fine-tuning operation on the large language model according to a predefined fine-tuning manner, evaluates a question and answer output effect index of the large language model, checks the question and answer output effect index with a predefined question and answer output effect index, and the oil knowledge question and answer platform determines whether to perform a repeated fine-tuning operation on the large language model, thereby completing oil knowledge question and answer based on the large language model.
[0169] The above content is merely an example and description of the structure of the present application, and those skilled in the art can make various modifications or supplements to the described specific embodiments or use similar ways to replace, as long as the modifications or supplements do not deviate from the structure of the present application or exceed the scope defined by the present application, and should belong to the protection scope of the present application.
Claims
1. A large language model-based oil knowledge question and answer method, characterized in that, Comprise: Data availability preprocessing: The oil knowledge Q&A platform collects original data in the oil field from various data sources, denoted as an original data set in the oil field, and a data processing component performs availability preprocessing operations on the original data set in the oil field to determine the availability evaluation value of the original data set in the oil field, checks it with the predefined availability evaluation threshold, determines the availability state of the original data set in the oil field, and records the original data set in the oil field after the availability preprocessing operation as the available data set in the oil field; Q&A pair application quality determination: The oil knowledge Q&A platform constructs the oil field Q&A pair set from the available data set in the oil field according to the predefined construction method, and performs calibration operations to determine the application quality index of each oil field Q&A pair, and compares it with the predefined application quality reference index to screen the oil field Q&A pair data set; Model output effect determination: The oil knowledge Q&A platform stores the oil field Q&A pair data set in a large language model to which the oil Q&A component belongs, and issues a fine-tuning instruction. The oil Q&A component fine-tunes the large language model according to the predefined fine-tuning method to evaluate the Q&A output effect index of the large language model, and checks it with the predefined Q&A output effect index. The oil knowledge Q&A platform determines whether to perform repeated fine-tuning operations on the large language model to complete the oil knowledge Q&A based on the large language model; The data processing component performs availability preprocessing operations on the original data set in the oil field, thereby obtaining preprocessing performance data of the data processing component and availability preprocessing operation data of the original data set in the oil field; The availability preprocessing operation specifically includes missing value processing, error value correction, and duplicate value deletion.
2. The oil knowledge Q&A method based on a large language model according to claim 1, characterized in that: The preprocessing performance data of the data processing component specifically includes the CPU occupancy rate of the data processing component at each preprocessing time point, the memory occupancy rate of the data processing component at each preprocessing time point, the preprocessing time length of the data processing component in the preprocessing period, and the data processing amount of the data processing component in the preprocessing period; The CPU occupancy rate of the data processing component at each preprocessing time point is processed by averaging to obtain the CPU average occupancy rate of the data processing component in the preprocessing period; The data processing amount of the data processing component in the preprocessing period and the preprocessing time length of the data processing component in the preprocessing period are processed by ratio to obtain the data processing throughput of the data processing component in the preprocessing period; The CPU average occupancy rate boundary value and the memory maximum occupancy rate reference value are extracted from the oil knowledge Q&A information library; The CPU average occupancy rate of the data processing component in the preprocessing period, the memory occupancy rate of the data processing component at each preprocessing time point, the preprocessing time length of the data processing component in the preprocessing period, and the data processing throughput of the data processing component in the preprocessing period are comprehensively processed to obtain the data preprocessing performance index of the data processing component; The availability of the original data set in the oil field is preprocessed operation data, specifically including the number of numerical missing values of each data group in the original data set in the oil field, the number of data garbled codes of the original data set in the oil field, the number of key elements of the original data set in the oil field, the total data amount of the original data set in the oil field, and the number of data formats of the original data set in the oil field.
3. The method according to claim 2, wherein the method is characterized in that: The availability evaluation value of the original data set in the oil field is determined, and the specific determination process is: The number of key elements of the original data set in the oil field is ratio-processed with the total data amount of the original data set in the oil field to obtain the proportion of key elements of the original data set in the oil field; The number of numerical missing values of each data group in the original data set in the oil field is added to obtain the total number of numerical missing values of the original data set in the oil field; The number of data format references is extracted from the oil knowledge Q&A information base; The total number of numerical missing values of the original data set in the oil field, the number of data garbled codes of the original data set in the oil field, the proportion of key elements of the original data set in the oil field, the number of data formats of the original data set in the oil field, and the data preprocessing performance index of the data processing component are comprehensively analyzed to obtain the availability evaluation value of the original data set in the oil field.
4. The method according to claim 3, wherein the method is characterized in that: The availability state of the original data set in the oil field is determined, specifically by checking the availability state of the original data set in the oil field with the availability evaluation threshold value to obtain the availability state checking result, so as to determine the availability state of the original data set in the oil field; The availability state checking result is a first availability state checking result or a second availability state checking result; The first availability state checking result is that the availability evaluation value of the original data set in the oil field is greater than or equal to the availability evaluation threshold value; The second availability state checking result is that the availability evaluation value of the original data set in the oil field is less than the availability evaluation threshold value; If the availability state checking result is the first availability state checking result, it is determined that the original data set in the oil field is in the available state, and if the availability state checking result is the second availability state checking result, it is determined that the original data set in the oil field is in the unavailable state, and the data optimization of the original data set in the oil field is performed.
5. The method of claim 1, wherein the method is based on a large language model. The application quality index of each oil field Q&A pair is determined, and the specific determination process is: Calibration operation data of the oil knowledge Q&A platform is obtained, specifically including calibration time length of each oil field Q&A pair and the number of semantic error sentences of each oil field Q&A pair; Application data of each oil field Q&A pair is obtained, specifically including the number of literature references of each oil field Q&A pair, the answer code of each oil field Q&A pair, and the knowledge update interval time length of each oil field Q&A pair; The number of literature reference adaptations of each oil field Q&A pair and the answer reference code of each oil field Q&A pair are extracted from the oil knowledge Q&A information base; The answer code of each oil field Q&A pair is data-processed with the answer reference code of each oil field Q&A pair to obtain and record the answer code similarity of each oil field Q&A pair; The application quality index of each petroleum field question and answer pair is obtained by comprehensively analyzing the answer coding similarity of each petroleum field question and answer pair, the calibration time length of each petroleum field question and answer pair, the number of semantic error sentences of each petroleum field question and answer pair, the number of literature references of each petroleum field question and answer pair, the knowledge update interval length of each petroleum field question and answer pair, and the availability evaluation value of the petroleum field original data set.
6. The method of claim 1, wherein the method is based on a large language model. The screening obtains the petroleum field question and answer pair data set, and the specific analysis process is as follows: If the application quality index of each petroleum field question and answer pair is greater than or equal to the application quality reference index, the data corresponding to each petroleum field question and answer pair is recorded as the petroleum field question and answer pair data set; If the application quality index of a certain petroleum field question and answer pair is less than the application quality reference index, the petroleum field question and answer pair is deleted, a number of petroleum field question and answer pairs with the application quality index greater than or equal to the application quality reference index are counted, and the data corresponding to the number of petroleum field question and answer pairs is recorded as the petroleum field question and answer pair data set.
7. The method of claim 1, wherein the method is based on a large language model. The evaluation of the question and answer output effect index of the large language model is as follows: The question and answer output data of the large language model is obtained, specifically including the number of knowledge point coverages under each question and answer output of the large language model, the amount of additional information provided under each question and answer output of the large language model, the answer output time length under each question and answer output of the large language model, and the number of user good comments in the question and answer output period of the large language model; The total amount of answer information under each question and answer output of the large language model is obtained; The amount of additional information provided under each question and answer output of the large language model is processed by ratio with the total amount of answer information under each question and answer output of the large language model to obtain the proportion of additional information under each question and answer output of the large language model; The number of knowledge point coverages under each question and answer output of the large language model is processed by ratio with the pre-defined number of knowledge point coverage adaptation under each question and answer output to obtain the knowledge point coverage degree value under each question and answer output of the large language model; The additional information reference proportion under each question and answer output, the answer output reference time length under each question and answer output, and the user good comment definition number are extracted from the petroleum knowledge question and answer information library; The question and answer output effect index of the large language model is obtained by comprehensively analyzing the knowledge point coverage degree value under each question and answer output of the large language model, the proportion of additional information under each question and answer output of the large language model, the answer output time length under each question and answer output of the large language model, the number of user good comments in the question and answer output period of the large language model, the availability evaluation value of the petroleum field original data set, and the application quality index of each petroleum field question and answer pair.
8. The method according to claim 7, wherein the method is characterized in that: The petroleum knowledge question and answer platform determines whether to perform repeated fine-tuning operation on the large language model, specifically, the petroleum question and answer component performs output effect verification on the question and answer output effect index of the large language model and the pre-defined question and answer output effect index to obtain an output effect verification result and upload the output effect verification result to the petroleum knowledge question and answer platform, so that the petroleum knowledge question and answer platform determines whether to perform repeated fine-tuning operation on the large language model; The output effect verification result is a first output effect verification result or a second output effect verification result; The first output effect verification result is specifically that the question and answer output effect index of the large language model is greater than or equal to the question and answer output effect index. The second output effect verification result is specifically that the question and answer output effect index of the large language model is less than the question and answer output effect index. If the output effect verification result is the first output effect verification result, the petroleum knowledge question and answer platform determines that the large language model does not need to be repeatedly fine-tuned, if the output effect verification result is the second output effect verification result, the petroleum knowledge question and answer platform determines that the large language model needs to be repeatedly fine-tuned, and issues a repeated fine-tuning instruction, and the petroleum question and answer component repeatedly fine-tunes the large language model according to the predefined repeated fine-tuning manner.
9. The method according to claim 8, wherein the method is characterized in that: The predefined repeated fine-tuning manner is specifically that the petroleum question and answer component uses the built-in LoRA method to repeatedly fine-tune the large language model in the petroleum knowledge field. The LoRA method is a technology that introduces an adaptive matrix into the large language model, and the specific repeated fine-tuning operation is to keep the weights of the original model in the petroleum question and answer component unchanged, add a low-rank matrix to adjust the output of the original model, and make the original model become a large language model in the petroleum knowledge field, so as to complete the repeated fine-tuning operation of the large language model.
10. A system for applying the method for querying petroleum knowledge based on a large language model according to any one of claims 1-9, characterized in that: It includes: A data availability preprocessing module is used to collect petroleum field original data from various data sources by the petroleum knowledge question and answer platform, which is denoted as a petroleum field original data set, and a data processing component performs availability preprocessing operation on the petroleum field original data set to determine the availability evaluation value of the petroleum field original data set, and checks it with the predefined availability evaluation threshold to determine the availability state of the petroleum field original data set, and the petroleum field original data set after the availability preprocessing operation is denoted as a petroleum field available data set; A question and answer pair application quality module is used to construct a petroleum field question and answer pair set from the petroleum field available data set according to a predefined construction method by the petroleum knowledge question and answer platform, and performs calibration operation to determine the application quality index of each petroleum field question and answer pair, and compares it with the predefined application quality reference index to screen the petroleum field question and answer pair data set; A model output effect module is used to store the petroleum field question and answer pair data set in the large language model to which the petroleum question and answer component belongs by the petroleum knowledge question and answer platform, and the petroleum knowledge question and answer platform issues a fine-tuning instruction, so that the petroleum question and answer component fine-tunes the large language model according to the predefined fine-tuning manner, evaluates the question and answer output effect index of the large language model, and checks it with the predefined question and answer output effect index, and the petroleum knowledge question and answer platform determines whether to repeatedly fine-tune the large language model, so as to complete the petroleum knowledge question and answer based on the large language model.
Citation Information
Patent Citations
Question answering method, question answering device and storage device based on Transformer model
CN111881279B
Question answering methods, training methods and devices for question answering models
CN114416953B
Instruction fine tuning data set construction method and device
CN118966379A
Light field image spatial super-resolution reconstruction method and device, equipment and storage medium
CN120543384A