Offshore oil professional intelligent question-answering system based on large model fine tuning and multi-mode RAG technology
By fine-tuning a large language model and using multimodal RAG technology, a multimodal knowledge base was constructed and combined with cosine distance retrieval, solving the problems of multimodal information processing and semantic understanding in marine oil engineering. This resulted in a highly accurate and sustainable intelligent question-answering system suitable for various complex scenarios in marine oil engineering.
Patent Information
- Application Number
- CN202510815871.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-11-14
AI Technical Summary
Existing offshore oil exploration and development field operations suffer from insufficient multimodal information processing capabilities, poor semantic understanding and professional adaptability, resulting in limited depth of knowledge utilization, lack of accuracy and credibility in responses, and the inability of the system to self-optimize.
By employing large language model fine-tuning and multimodal RAG technology, a multimodal knowledge base is constructed, and semantic segmentation and semantic vector encoding are performed. Combined with a two-stage semantic retrieval mechanism based on cosine distance, it supports unified retrieval and generation of text, images, and charts, and the system is optimized through a user feedback mechanism.
It achieves highly accurate, multimodal support, and continuously evolving intelligent question-answering capabilities, enhancing professional expression and user experience, and is applicable to various complex scenarios in offshore oil engineering.
Smart Images

Figure CN120952015A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence technology and oil and gas engineering information processing technology, specifically to an intelligent question-and-answer system for marine petroleum professionals based on large model fine-tuning and multimodal RAG technology. Background Technology
[0002] It is applicable to semantic understanding and intelligent question answering tasks in engineering scenarios such as marine drilling, logging, well logging and well completion, and belongs to the fields of natural language processing, intelligent information retrieval and oil and gas professional knowledge service technology. Current offshore oil exploration and development field operations require the acquisition of a large amount of professional knowledge in real time. However, these systems suffer from several problems: they only support text data and cannot process multimodal information such as images and charts, which limits the depth of knowledge utilization; the retrieval methods that rely on keyword matching lack semantic understanding capabilities, resulting in low matching accuracy; large language models are prone to generating illusionary content in specific domain applications; the answers lack professional control; the system cannot learn and evolve based on feedback; and it has poor professional adaptability, making it difficult to effectively support the needs of field applications in the long term. Summary of the Invention
[0003] This invention aims to address the shortcomings of existing technologies in terms of professional semantic understanding, multimodal information processing, contextual response capabilities, and continuous system optimization. It provides an intelligent question-answering method and system for marine drilling logging and completion based on large language model fine-tuning and multimodal retrieval-Augmented Generation (RAG) technology, so as to achieve highly accurate, highly professional, multimodal support, and continuously evolving intelligent question-answering capabilities for field applications in offshore oil engineering.
[0004] The method of the present invention includes the following steps: (1) Construction of multimodal knowledge base: Collect multi-source data such as text documents, images, charts and structured tables related to marine drilling, logging, well logging and completion, and perform format regularization and semantic segmentation. Text information is encoded by embedding model to construct semantic vectors, and image and table resources are linked to their descriptive paragraphs and uniformly included in the multimodal semantic retrieval space.
[0005] (2) Fine-tuning of large language models: Select pre-trained large language models as the basis (such as the Qwen series), use oil and gas engineering professional corpus for supervised fine-tuning, introduce a training strategy of alternating between general and professional corpus, and design a knowledge anti-forgetting mechanism to enhance the model's expressive ability and professional stability in the field of marine oil.
[0006] (3) Natural language query parsing: The system receives natural language questions input by users, extracts semantic keywords and key field information, and generates semantic vector representations for retrieval.
[0007] (4) Two-stage semantic retrieval mechanism: including two stages: index-level document recall and paragraph-level matching backoff, both of which are based on cosine distance for semantic matching calculation.
[0008] During the indexing phase, the system calculates the cosine distance between the user query vector and the document index vector. The smaller the distance, the closer the two vectors are in direction, meaning the higher the content relevance.
[0009] If the user's question contains fields such as hash symbols, then further execute the field matching rules to ensure that the recalled document is consistent with the key fields; During the paragraph matching rollback phase, when the index does not match a document, the system performs cosine distance matching on all paragraphs in the database and selects the paragraph with the smallest distance for subsequent processing.
[0010] Using cosine distance as a semantic difference metric is suitable for approximate retrieval in high-dimensional semantic spaces, and can effectively improve recall accuracy and matching robustness in professional question-answering scenarios.
[0011] (5) Knowledge generation mechanism trigger: When no valid content is found in the semantic retrieval stage, the system generates a controlled answer based on the large language model. The generation is strictly limited to the corpus and context already learned by the model, and the generation of information without source support is prohibited in order to control the accuracy and credibility of the content.
[0012] (6) Multimodal output generation: Automatically embed image, chart or table resources according to the type of answer content. The system supports text output, mixed text and image layout and structured result display, improving professional expression and visualization capabilities.
[0013] (7) Multi-round semantic interaction support: The system retains context information to support continuous follow-up questions and context-aware question-and-answer interaction, improving user experience and query continuity.
[0014] (8) User feedback collection and trustworthy tag generation: The system collects user behavior feedback information (including likes, follow-up questions, exits, etc.) to generate knowledge trustworthy tags, drive the iteration of knowledge base content and continuous fine-tuning and optimization of the large language model, and improve the system's evolution capability.
[0015] Furthermore, the system supports hierarchical explanations of technical terms, including term definitions, structural components, principle explanations, and application examples, which are suitable for engineering training and assisting non-professional users in understanding.
[0016] Furthermore, the system can be extended to access real-time data sources on-site, enabling linkage with marine drilling operation systems (such as well site sensors, construction records, etc.) to generate personalized decision-making suggestions based on the current working conditions.
[0017] In summary, this invention, by constructing a structured, multimodal, and semantically unified knowledge base, and combining a large language model finely tuned with professional corpus with a two-stage RAG semantic retrieval mechanism based on cosine distance, achieves accurate question answering, text-image response, multi-turn interaction, and feedback optimization in the professional context of marine petroleum engineering. It has broad engineering application prospects and practical value.
[0018] This invention provides an intelligent question-answering system for marine petroleum professionals based on large model fine-tuning and multimodal RAG technology. It offers the following advantages:
[0019] (1) The present invention has strong multimodal knowledge fusion capability: the system supports unified regularization and semantic encoding of multi-source data such as text, images, charts and structured tables, and constructs a unified knowledge representation space across modalities, breaking through the limitation of traditional question answering systems that only process text information, and significantly improving the breadth of knowledge coverage and query flexibility.
[0020] (2) The present invention achieves high accuracy and efficiency through semantic retrieval: The present invention adopts a two-stage semantic matching mechanism based on cosine distance, which performs refined retrieval at the document level and paragraph level respectively. It has high matching accuracy and strong retrieval recall capability, and is particularly suitable for professional technical question and answer scenarios with sensitive fields such as hash symbols.
[0021] (3) The present invention highlights the professional expression ability: by introducing oil and gas engineering corpus for supervised fine-tuning and combining knowledge anti-forgetting strategy, the expression ability of large language model in terms of terminology use, operation norms, and scene understanding is significantly improved, ensuring that the output content of the system has high professionalism and practical value.
[0022] (4) The present invention provides flexible multimodal visualization output: the system supports multiple output formats such as mixed text and graphics, structured data display and original chart embedding, which helps to improve the clarity, intuitiveness and engineering applicability of the answer content and meet various usage scenarios such as technical support, learning and training and remote decision-making.
[0023] (5) The present invention has strong multi-round contextual interaction capabilities: The present invention supports continuous follow-up questions from users and contextual semantic tracking, enabling the system to maintain topic consistency and logical coherence in real-life question-and-answer processes, thereby enhancing the interactive experience and system practicality.
[0024] (6) The present invention has self-optimization capability through the system: through the collection of user feedback behavior and the generation of credibility labels, the system realizes the identification of low credibility answers and dynamic updates of knowledge base content, and can be used for continuous fine-tuning training of the model, with good adaptive evolution capability.
[0025] (7) The present invention has high engineering adaptability and scalability: the method and system can be flexibly connected to the well site real-time data and production record system, supporting application needs such as on-site auxiliary decision-making, construction parameter analysis, and instrument usage consultation. It has broad engineering application value and deployment prospects. In summary, the present invention has made breakthroughs in key technologies such as multimodal fusion, semantic retrieval, professional answers, graphic and text display, contextual interaction and system evolution. It is applicable to various complex scenarios such as technical Q&A, intelligent consultation and operation training in marine oil engineering, and has significant practical application value and promotion significance. Attached Figure Description
[0026] Figure 1 This is a schematic diagram of the question-and-answer system of the present invention. Detailed Implementation
[0027] To make the technical solution of the present invention clearer, the present invention will be described in detail below with reference to specific embodiments.
[0028] This embodiment provides a method and system for knowledge-based question answering in marine drilling logging and completion based on a large language model. It involves the construction and encoding of multimodal data, fine-tuning of the language model, semantic vector generation, retrieval-enhanced generation (RAG) algorithm, multimodal output display, and the integrated implementation of a user feedback mechanism. The invention will be further described below with reference to specific embodiments.
[0029] Example 1 S1. Construction of a multimodal knowledge base Collect professional data covering all operational stages of offshore drilling, logging, well completion, etc. Data sources include, but are not limited to: technical operation manuals, textbooks, instrument and equipment manuals, historical well site construction data, and record documents containing images and tables.
[0030] The document content is extracted using OCR (if it is an image PDF), converted into a structured format, and then semantically segmented so that each paragraph becomes an independent semantic unit.
[0031] For example, a total of 7405 segments were actually extracted. The specific document and segment statistics are shown in the table below:
[0032] The final segmentation example is as follows:
[0033] The specific processing procedure is as follows: (1) Document parsing and formatting: Convert unstructured documents (such as PDFs and image files) into parsable text formats. If the document is a scanned document, use an OCR tool to extract the text content. For PDFs with partially embedded images, extract and label the positions and descriptions of the charts.
[0034] (2) Semantic segmentation and metadata annotation: The extracted text is manually analyzed and semantically segmented using natural language processing tools. Each paragraph is bound to metadata such as document name, page number, paragraph number, technical category, and stage label (e.g., "well logging interpretation" or "well logging process"), and records whether it contains images or table links.
[0035] (3) Semantic encoding and vectorization: All text paragraphs are semantically vectorized to form a unified semantic expression space.
[0036] (4) Image and complex table linking processing: Images and complex tables are not directly vectorized, but are associated with their corresponding paragraphs through linking, and the paragraphs participate in unified semantic retrieval.
[0037] (5) Entry construction and knowledge base entry: The paragraph semantic vector, related links and their metadata are encapsulated into a unified knowledge entry structure, stored in the multimodal knowledge management system, and a paragraph-level index structure is constructed for subsequent use by the semantic retrieval mechanism based on cosine distance.
[0038] Through the above steps, a high-quality, well-structured multimodal semantic knowledge base was constructed, providing a unified data foundation and retrieval support for professional question answering and semantic generation of large language models.
[0039] S2, Model Fine-tuning This step aims to perform targeted optimization of the pre-trained large language model in the field of marine oil engineering, enabling it to understand professional terminology, express complex questions and answers, and output structured knowledge.
[0040] The large language model can be a Qwen series model, such as Qwen-7B, Qwen2.5-72B-Instruct, or other open-source large models with multi-turn question answering and knowledge generation capabilities, such as InternLM2.5-7B. This embodiment preferably uses the Qwen2.5-72B-Instruct model as the basic language model.
[0041] (1) Training data preparation: Semantic paragraphs are extracted from multiple publicly available or authorized professional materials to construct supervised training samples. Each sample includes three parts: instruction, input, and output, as shown in the example below:
[0042] { "instruction": "Based on the user's input of a natural language question, and combined with CNOOC's relevant knowledge and information, generate an accurate answer."
[0043] "input": "The Importance of Electrical Logging in Geophysical Logging", "output": "Electrical logging is a specific application and manifestation of macroscopic electromagnetic field theory in the field of geophysical logging..." (2) Training strategy design: Supervised fine-tuning (SFT) is adopted and training is performed on a multi-GPU cluster (such as 8×A100). The following parameters are set: ① The learning rate is set to 2e-5; ② Batch size is 16; ③ The training cycle is set to 3-5 rounds; ④ The optimizer used is AdamW; ⑤ The maximum truncation length of the input sequence is 2048 tokens.
[0044] The training corpus adopts an alternating mixed training strategy of "general corpus + industry corpus" to maintain the language model's inheritance of general language capabilities, while enhancing its expressive ability in oil and gas engineering terminology, processes, rules, etc.
[0045] (3) Introduction of knowledge anti-forgetting mechanism: In order to avoid the model forgetting general knowledge during the professional fine-tuning process, a knowledge retention loss term is introduced during training. The performance of professional tasks and general question answering tasks is optimized in the training objective function at the same time to improve the generalization stability and knowledge integrity of the model.
[0046] (4) Training process monitoring and verification evaluation: To ensure a stable and controllable training process, a training management system that includes multi-dimensional evaluation and monitoring is constructed: Loss function monitoring: Real-time recording of the loss change trend of the training set and validation set, smoothing the curve with a moving average, and judging whether there is a risk of overfitting; BLEU and ROUGE metrics are used to measure phrase accuracy and information coverage, respectively, and are applicable to definition-based and explanation-based question answering. Manual review mechanism: After each round of training, a sample of verification samples is selected and reviewed by engineers with many years of field experience in oil and gas for terminology, logic, accuracy and usability.
[0047] (5) Model Deployment and Loading: After training, the fine-tuned large language model is integrated into the system's language model processing module, and question-answering capabilities are provided to the outside world via API. The system supports subsequent incremental fine-tuning and online learning mechanisms to adapt to on-site knowledge updates and changes in assignments.
[0048] Through the above steps, this invention achieves fine-tuning and expression control of industry knowledge for offshore oil engineering, enhances the language model's ability to understand and answer questions about professional domain terms, scenarios, structures, and processes, and provides support for semantic retrieval, knowledge generation, and interactive question answering in subsequent steps.
[0049] S3, Query Parsing and Vector Generation The purpose of this step is to transform the natural language question input by the user into a vector representation suitable for the semantic retrieval process, thereby enabling efficient semantic matching and question answering responses in the future.
[0050] (1) Natural language question input: Users input natural language questions through the system front-end interface, such as: "What instruments are used in integrated logging?" The system receives the question text and enters the semantic processing flow.
[0051] (2) Domain semantic analysis: The system calls a named entity recognition model trained on an oil and gas engineering corpus to identify key semantic elements in the query, including but not limited to: ① Structured fields such as well number and well location; ② Parameter items (such as "resistivity" and "gamma rays"); ③ The object of the operation (e.g., "well logging" or "well surveying"); ④ Action type (e.g., "use", "analyze", "explain"); ⑤ Name of the tool or equipment (such as "instrument" or "sensor").
[0052] The aforementioned semantic elements are used to enhance the representation of the query vector and are used in subsequent field matching logic.
[0053] (3) Query semantic vector generation: The cleaned and structured natural language question is input into the unified semantic coding model for embedding calculation to obtain a fixed-length semantic vector representation. This vector serves as the query input of the two-stage semantic retrieval module (RAG) and participates in the subsequent calculation of the cosine distance with the semantic vectors of documents and paragraphs.
[0054] (4) Field information marking: If the user's question contains structural field information such as well number, measurement point depth, and operation stage, the system will extract the information as an independent query field and perform field consistency judgment in the rule verification stage of the subsequent retrieval process.
[0055] (5) Context state maintenance: For multi-round continuous questioning scenarios, the system retains the user's question, the system's answer and related semantic context from the previous round, and considers the context information together when generating the query vector in the current round, so as to improve the logical consistency and response relevance of continuous question and answer.
[0056] Through the above processing steps, the system transforms unstructured natural language questions into searchable structured semantic vector representations and establishes semantic scene context, providing accurate input for two-stage cosine distance matching and subsequent large language model question answering generation.
[0057] S4, Two-stage RAG retrieval mechanism The purpose of this step is to accurately retrieve the most relevant document paragraphs or knowledge units to the user's query from the constructed multimodal knowledge base through a two-stage semantic matching mechanism, thereby providing basic support for question answer generation. The retrieval mechanism is based on the Retrieval-Augmented Generation (RAG) structure and is divided into an index matching stage and a paragraph matching backoff stage, both of which use cosine distance as the semantic matching metric.
[0058] (1) Index matching stage: The system first calculates the cosine distance between the semantic vector of the user query and the pre-generated index vector of each document in the knowledge base. The calculation formula is as follows: ; Here, represents the query vector, and represents the document fragment vector. The smaller the distance, the closer the two vectors are in direction, meaning the higher the content relevance. The system sorts the documents by distance from smallest to largest and selects the Top-N most similar document fragments as a reference for answer generation.
[0059] (2) Index rule verification stage: If the user's question contains structure fields (such as hash number, well segment, depth range, etc.), the system will perform field consistency verification in the candidate documents: ① Compare whether the hash number information recorded in the candidate document metadata is consistent with the query field; ② If there are inconsistencies or missing information, the candidate document will be excluded; ③ Retain documents that meet the consistency requirements for hash number, stage, or device fields, ensuring that both semantics and structure match.
[0060] This step can effectively filter irrelevant content that is semantically similar but does not match the fields, thus improving the accuracy of the matching results.
[0061] (3) Paragraph matching rollback stage: If no documents that meet the conditions are recalled during the indexing stage, the system starts the paragraph-level rollback mechanism and performs cosine distance matching on all semantic paragraphs in the knowledge base.
[0062] ① The system calculates the cosine distance between the query vector and the semantic vector of each paragraph; ② Combined with metadata such as field information, document category, and technical tags for comprehensive filtering; ③ Finally, the paragraph with the smallest distance is selected as the search result for the subsequent question-and-answer module to generate the answer.
[0063] This stage offers finer-grained matching, making it suitable for precise problem localization in long document structures. It ensures that even if the index is not hit, it can still return the knowledge unit most relevant to the query content.
[0064] (4) Matching result determination: If no valid matching result is obtained in the paragraph back-off stage (i.e. all cosine distances are higher than the set threshold), the system considers it as "no search hit" and enters the next step of knowledge generation stage (see step S5), where the large language model performs controlled generative answers based on context.
[0065] Through the above two-stage structural design, this invention maintains a high recall rate while enhancing the semantic accuracy and structural consistency of the results, and improving the fault tolerance and adaptability of the retrieval module to business scenarios.
[0066] S5. Content Generation and Output The purpose of this step is to use the finely tuned large language model, based on the user's question and the aforementioned search results, to generate professional, accurate, and logically rigorous answers. The system strictly follows preset instructions to ensure that the answers are reliable, semantically clear, and meet the requirements of the job scenario.
[0067] (1) Domain Judgment and Strategy Branching: The system first determines whether the user's input question belongs to the petroleum industry or related fields (such as oil and gas geology, well logging engineering, drilling and completion technology, etc.). The processing logic is as follows:
[0068] If the problem is related to petroleum engineering or geology, then proceed to the data matching and generation process; If the question is irrelevant, the system will not reject the answer and will call the fine-tuning model to provide a concise response, avoiding unnecessary extensions or repetition of background information. All generation processes must ensure that there is no hallucination content, no fabrication, and that the logic is consistent.
[0069] (2) Generation strategy when data is matched: If a paragraph or chart link related to the question is successfully retrieved in the multimodal knowledge base, the system executes the following strategy: Text paragraph matching: The system takes the matching paragraph and the question as input, calls the model to generate a non-repetitive and non-reiterative answer to the question, ensuring that the information is directly quoted, not deleted, and the language is clear; Chart paragraph matching: The system does not generate text, but directly displays the original image, chart or table resources, along with their associated text descriptions, to form a mixed text and image layout.
[0070] (3) Generation strategy when no data is found: If no relevant data is found during the search, the system will strictly follow the instructions and respond as follows: Fixed output: We're sorry, we haven't found any relevant information at the moment.
[0071] Following this suggestion: Although there is no direct answer at the moment, the following information may be helpful to you.
[0072] The model generates knowledge in a limited way based on training corpus and similar contexts, ensuring logical rigor and avoiding the inclusion of fictitious content.
[0073] This approach ensures that valuable professional advice can still be provided to users even when information is missing.
[0074] (4) Cross-domain problem handling: For questions that do not belong to the petroleum or related fields, the system does not refuse to answer, but calls the general capabilities of the model to answer concisely, outputs no more than necessary information, and does not introduce irrelevant background or terminology.
[0075] (5) Formula generation specifications: When the system's responses involve parameter calculations, well logging models, physical relationship expressions, etc., the output should use LaTeX syntax. All formula expressions must have a clear source and ensure that the physical meaning, units, and actual engineering context are consistent.
[0076] (6) Multi-round question and answer and context tracking: The system supports users to ask supplementary or follow-up questions after a question and answer session. The system will automatically recognize the semantic content of the previous round and retain the context state to construct a new round of complete input vector, ensuring that the answers are coherent, the background is consistent, and the logic does not jump in multiple rounds of dialogue.
[0077] S6, User Feedback and Trusted Tagging Mechanism This step aims to collect feedback information generated by users during the use of the question-and-answer system, establish a knowledge credibility tagging system, provide data support for subsequent knowledge base optimization and continuous model fine-tuning, and build the system's adaptive evolution capability.
[0078] (1) Feedback behavior collection: The system records user interaction behavior with each question and answer result in real time through the front-end interface or API. The main feedback types include, but are not limited to:
[0079] A "like" indicates that the user is satisfied with the answer or considers the answer valid; A "dislike" indicates that the user believes the answer is inaccurate, incomplete, or has no reference value.
[0080] All feedback information is recorded in a structured log format and bound to the following metadata: Issue ID; Source of generation (hit paragraph / generation mode); User behavior types; Non-sensitive statistical data such as timestamps and IP addresses (optionally encrypted).
[0081] (2) Knowledge base optimization and content review suggestions: When a knowledge paragraph or model output has multiple negative feedbacks (such as continuous downvotes or high bounce rate), the system marks the content as a low credibility item and adds it to the manual review queue.
[0082] Administrators or knowledge engineers can periodically review low-trust content; Mark, replace, or supplement outdated, incorrect, or unclear paragraphs; All optimization actions are performed manually to ensure system stability and knowledge accuracy.
[0083] The system does not directly perform automatic update operations to ensure control and content review security.
[0084] (3) Model fine-tuning data selection strategy: The credibility label is also used in the subsequent data sample selection process for incremental fine-tuning of the language model: High-quality response samples can be included in the fine-tuning dataset; Samples with low credibility or those frequently questioned by users will be excluded or reconstructed; The system can combine BLEU, ROUGE and user behavior scores to dynamically adjust the priority of training data.
[0085] This mechanism enables the question-answering system to evolve over the long term, supporting fine-tuning and iteration without compromising current performance and stability.
[0086] (4) Feedback Privacy and System Audit Protection: User feedback is collected only for technical optimization and does not contain sensitive information such as user identity. The system provides feedback log export and administrator audit interface to ensure that data use is controllable and the process is traceable.
[0087] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0088] This invention provides an intelligent question-answering system applicable to professional fields such as marine drilling and oil and gas exploration by combining the innovative application of large language model fine-tuning and multimodal RAG technology. The system has significant advantages in knowledge base construction, professional terminology processing, query and retrieval accuracy, answer generation and user feedback mechanism.
[0089] The technical solutions of this invention are not limited to the specific implementations in the embodiments. Any person skilled in the art can modify or improve the implementation details without departing from the core technical concept and scope of protection of this invention. Such modifications or equivalent substitutions should be considered within the technical scope of this invention. Therefore, the scope of protection of this invention should be determined by the appended claims.
Claims
1. A method for intelligent question answering in marine drilling logging and completion based on large language model fine-tuning and multimodal retrieval enhanced generation (RAG) technology, characterized in that... Includes the following steps: S1. Collect multi-source data such as text, images, charts and tables related to drilling, logging, well logging and well completion, perform format regularization, semantic segmentation and vectorization processing, and build a unified coding multimodal knowledge base. Among them, images and complex tables are linked to descriptive text paragraphs through links, and text paragraphs participate in semantic coding. S2. Supervised fine-tuning of the pre-trained large language model based on corpus in the field of oil and gas engineering is carried out. A mixed training strategy of general and professional corpus is adopted, and a knowledge anti-forgetting mechanism is introduced. S3. Receive natural language questions input by the user, parse semantic information and generate query vectors; S4. Perform two-stage semantic retrieval, including: a. Index recall phase: Based on the cosine distance between the query vector and the document index vector, documents with a distance lower than a preset threshold are selected as candidates; b. Index rule verification phase: If the problem contains fields such as hash numbers, then hash number consistency verification and other field matching rules will be further performed on the candidate documents, and only the relevant documents will be retained; c. Paragraph matching fallback phase: If no document is found during the indexing phase, cosine distance matching and rule judgment are performed on all paragraphs in the database to select the most relevant paragraph; d. Knowledge generation stage when no match is found: When no valid match is found in the indexing stage and the paragraph backtracking stage, the system triggers the knowledge generation mode of the large language model and provides a reasoned answer to the question based on the learned domain knowledge. The generation process is constrained by the knowledge boundary control mechanism and prohibits the free generation of unverified information. S5. Input the hit paragraph or generated content into the large language model to generate question-and-answer results; based on whether the paragraph contains image, chart, or table links, automatically call relevant resources for multimodal mixed display, supporting text, graphic, or structured output; support multi-round semantic follow-up interaction, and the system automatically retains the context for continuous answers; S6. Collect user feedback data to generate knowledge credibility tags, which are used to drive knowledge base updates and continuous model optimization.
2. The marine petroleum professional intelligent question-answering system based on large model fine-tuning and multimodal RAG technology according to claim 1, characterized in that: Images, charts, and complex tables are not directly vector-encoded. Instead, they are attached to their corresponding text paragraphs via links. These text paragraphs are then encoded using a language model and used for semantic retrieval.
3. The marine petroleum professional intelligent question-answering system based on large model fine-tuning and multimodal RAG technology according to claim 1, characterized in that: During the index rule verification phase, the system performs a hash field matching on questions containing hash symbols, retaining only document content that matches or is related to the hash symbol in the query.
4. The marine petroleum professional intelligent question-answering system based on large model fine-tuning and multimodal RAG technology according to claim 1, characterized in that: The paragraph matching fallback phase is triggered only if no document is hit during the indexing phase, and performs the same field matching and metadata filtering.
5. The marine petroleum professional intelligent question-answering system based on large model fine-tuning and multimodal RAG technology according to claim 1, characterized in that: The knowledge generation phase is triggered only when there is no match in the entire retrieval process, and the generation process is limited to the scope of model training, prohibiting the generation of untrusted content.
6. The marine petroleum professional intelligent question-answering system based on large model fine-tuning and multimodal RAG technology according to claim 1, characterized in that: The system output can automatically embed charts or tables based on paragraph type, enhancing visualization and professional expression.
7. The marine petroleum professional intelligent question-answering system based on large model fine-tuning and multimodal RAG technology according to claim 1, characterized in that: The user feedback includes behaviors such as liking, asking follow-up questions, and exiting the app. The feedback information is used to build knowledge credibility tags for subsequent content optimization.
8. The marine petroleum professional intelligent question-answering system based on large model fine-tuning and multimodal RAG technology according to claim 1, characterized in that: Knowledge credibility tags influence retrieval ranking strategies and the selection of training sample weights, enabling the system to dynamically and adaptively evolve.
9. The marine petroleum professional intelligent question-answering system based on large model fine-tuning and multimodal RAG technology according to claim 1, characterized in that: The system supports hierarchical explanations of terms, including term definitions, structural components, principle explanations, and application examples, making it suitable for engineering training scenarios.
10. A marine drilling logging and completion intelligent question-and-answer system implementing the method of any one of claims 1 to 9, characterized in that, include: S1, Natural Language Processing Module, is used to parse user questions and extract semantic vectors and key fields; S2, Multimodal Knowledge Construction Module, is used to construct a unified retrieval structure from information such as text vectors and chart links; S3, a two-stage semantic retrieval module, including index recall, rule verification, paragraph rollback, and a fallback generation mechanism; S4, Large Language Model Question Answering Module, is used to generate professional answers based on controlled content; S5, a multimodal display module, is used to output text, charts, or structured results, and supports context-preserving multi-round follow-up questions; S6, User Feedback Collection Module, is used to record behavioral data and drive the continuous optimization of knowledge and models.
Citation Information
Cited By
Intelligent teaching assisting system base LLM training method for well drilling simulator
CN121414554A
Drilling simulator-oriented intelligent teaching assistant system base LLM training method
CN121414554B
Multi-level well drilling knowledge base intelligent construction method and device and computer equipment
CN121902941A