Hydrofracture question-answering system and method based on cross-language retrieval enhanced generation

By using cross-language retrieval enhancement generation technology, a dynamic text vector library is constructed, which solves the problems of data lag and terminology adaptation in large language models in the field of hydraulic fracturing. This enables efficient, accurate, and real-time acquisition of professional knowledge and supports multilingual retrieval and dynamic updates.

CN121029935APending Publication Date: 2025-11-28ZHEJIANG UNIV

Patent Information

Application Number
CN202511145650.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing large language models in the field of hydraulic fracturing suffer from problems such as data lag, answer illusions, and difficulties in personalized adaptation. Traditional text vector libraries have low update efficiency and poor adaptability to domain terminology, making it difficult to meet the real-time knowledge acquisition needs of specialized fields.

Method used

Employing cross-language retrieval enhancement generation technology, a dynamic text vector library is constructed. By encoding document blocks and user questions into high-dimensional text vectors through a pre-trained model, and combining the approximate nearest neighbor algorithm and a pre-trained reordering model, fast matching and accurate answers are achieved. It also supports incremental updates of the knowledge base and mapping of Chinese and English terms.

Benefits of technology

It achieves efficient, accurate, and real-time knowledge acquisition in the field of hydraulic fracturing, improves the reliability and accuracy of professional Q&A, supports multilingual retrieval and dynamic knowledge updates, and reduces the need for manual browsing of unstructured documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029935A_ABST
    Figure CN121029935A_ABST
Patent Text Reader

Abstract

The invention discloses a system and a method for enhancing generation of hydrofracture questions and answers based on cross-language retrieval. The knowledge question-answering system aiming at the hydraulic fracturing field is developed by utilizing a cross-language retrieval enhancement generation technology, so that a user can obtain an optimal answer and a corresponding source from a multi-language knowledge base by matching only by using Chinese retrieval; the complex process that a user needs to find useful information from numerous and jumbled papers and reports is changed. According to the system, a dynamic text vector library framework is adopted, multi-language and complex document formats are supported, the analysis function of tables and formulas is integrated, documents in the hydraulic fracturing field are stored in the same knowledge base, the knowledge base is continuously updated along with document expansion in the professional field, a user can inquire related problems in the field in the system, and the user experience is improved. By integrating and updating a large amount of document information, required professional knowledge in the field can be efficiently obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of natural language processing and information retrieval technology, specifically to a hydraulic fracturing question-and-answer system and method based on cross-language retrieval enhancement generation, which is particularly applicable to document processing and question-and-answer scenarios in the field of hydraulic fracturing engineering technology. Background Technology

[0002] With the rapid development of information technology, textual data in the field of hydraulic fracturing has experienced explosive growth, covering multiple dimensions such as technical documents, engineering practices, and scientific research papers. In recent years, the scale of oil and gas operations has expanded rapidly, and the iteration of fracturing technology has accelerated. Between 2018 and 2022, global demand for hydraulic fracturing increased by 60%, and the number of fracturing segments increased by 80%. The expansion of the scale of operations has led to a leap in the volume of textual data such as technical documents and construction reports.

[0003] At the same time, text data sources are diversified. For example, the FracFocus knowledge base contains more than 39,000 fracturing fluid composition disclosure documents, and more than one million fracturing-related studies have been published globally from 1946 to 2022. Documents that are mainly in natural language lack a unified semantic standard, making it difficult for traditional knowledge bases to parse them efficiently.

[0004] While existing Large Language Models (LLMs) perform well in general question-answering tasks, they still have limitations in highly specialized fields such as hydraulic fracturing. LLMs, on the other hand, can understand and follow human natural language instructions. They utilize massive text datasets, are trained using the Transformer architecture and self-attention mechanisms, and achieve multi-task learning by predicting the next word in a sequence, thus representing a significant technological breakthrough in the field of artificial intelligence.

[0005] However, LLMs still face several challenges in their use: First, they rely on static training data from the past, making it difficult to acquire the latest knowledge; second, in the absence of sufficient domain knowledge as training data, models are prone to confusing questions or answers, leading to "illusory facts"; third, insufficient local knowledge adaptation makes it difficult to customize personalized LLMs, such as optimizing a specific solution based on a particular project. Therefore, the development of LLMs urgently needs a technology that can address the issues of data lag, answering illusions, and personalized adaptation, further improving the accuracy and reliability of responses.

[0006] To address these challenges, Retrieval Enhanced Generation (RAG) technology offers a groundbreaking solution for knowledge-based question answering in the field of hydraulic fracturing by integrating retrieval and generation mechanisms. RAG effectively solves the knowledge gap problem in LLMs by dynamically retrieving the latest information from external knowledge bases (such as patent databases, academic papers, and engineering reports). Simultaneously, by anchoring answers to search results, it significantly reduces the probability of fabricated content and supports answer tracing, thereby enhancing credibility. Furthermore, RAG can be customized to access professional knowledge bases, using a search engine to accurately filter highly relevant documents, ensuring that the generated content aligns with the application context. Several typical cases (such as MedGraphRAG in the medical field) have demonstrated that RAG outperforms fine-tuned LLMs in professional question answering, providing a technological paradigm for the hydraulic fracturing field.

[0007] Currently, the application scope of enhanced retrieval generation technology in the field of hydraulic fracturing is still insufficient. Traditional text vector libraries have problems such as low efficiency of static index updates, poor adaptability of domain terms, and lack of support for real-time incremental maintenance. Especially in the engineering field, new terms, formulas and standards continue to emerge, and the existing architecture is difficult to achieve dynamic evolution of the knowledge base.

[0008] Chinese patent CN120144720A, entitled "Multi-source Data Fusion Knowledge Enhancement Retrieval Generation Method, Device, and System," proposes a multi-source data fusion knowledge enhancement retrieval generation method. Compared to the former, this application aims to build a specialized question-and-answer system in the field of hydraulic fracturing, improving the reliability and accuracy of responses from LLMs (Locally Multi-source Machines) in this field. Summary of the Invention

[0009] To address the shortcomings of existing technologies, this invention provides a hydraulic fracturing question-and-answer system and method based on cross-language retrieval enhancement generation.

[0010] In a first aspect, the present invention provides a hydraulic fracturing question-answering method based on cross-language retrieval enhancement generation, comprising the following steps:

[0011] The client inputs collected hydraulic fracturing documents, and the server generates a text vector knowledge base.

[0012] Based on the collected hydraulic fracturing documents, hydraulic fracturing-related questions are retrieved on the client side, and the server generates text vectors.

[0013] The text vector knowledge base is initially matched with the text vector to obtain coarse-grained results;

[0014] The coarse-grained results are then finely sorted using a pre-trained re-sorting process, and the final sorted results are output.

[0015] The client generates a final response based on the server's preset prompt word rules and the final sorting results.

[0016] Secondly, the present invention provides a hydraulic fracturing question-and-answer system based on cross-language retrieval enhancement generation, including a client and a server. The client inputs collected documents, and the server generates a text vector knowledge base.

[0017] The client includes a user operation unit and a system output module;

[0018] The user operation unit includes a knowledge base loading mode and a user retrieval mode;

[0019] The knowledge base loading mode provides users with the option to select how to load the knowledge base.

[0020] The user search mode allows users to search for relevant issues and system interactions in the field of hydraulic fracturing.

[0021] The server includes a document processing module, a vectorization module, a vector storage module, a vector retrieval module, a ranking module, and a generative question answering module;

[0022] The document processing module intelligently divides multi-format documents in the same folder into document blocks, extracts terms and formulas, and stores them independently.

[0023] The vectorization module, based on a pre-trained embedding model, encodes document blocks and user questions into high-dimensional text vectors, generating semantic numerical representations to support similarity retrieval.

[0024] The vector storage module normalizes and persistently stores the text vector, supporting dynamic addition, deletion, modification, and querying of files;

[0025] The vector retrieval module quickly matches the user's retrieval vector with the vector library using an approximate nearest neighbor algorithm, and returns a coarse-grained candidate set.

[0026] The fine-ranking module returns coarse-grained results from the vector retrieval module by combining a pre-trained fine-ranking model with user questions and document content.

[0027] The generative question-answering module uses a pre-trained large language model to generate a structured answer based on domain-customized prompt word templates, taking the user's question and the refined document fragment as input.

[0028] The beneficial effects of this invention are as follows: This invention utilizes retrieval enhancement generation technology to develop a hydraulic fracturing question-and-answer system based on cross-language retrieval enhancement generation. After a user searches for a question, the system can accurately extract relevant information and literature sources from massive amounts of complex multilingual documents, significantly improving efficiency compared to traditional manual retrieval methods. This invention adopts a dynamic text vector library architecture, and as users continuously add new question-and-answer pairs, the system's knowledge coverage and answer precision continuously improve. Through standardized output templates, previously scattered professional knowledge is transformed into structured information, greatly improving the efficiency of knowledge acquisition in the field of hydraulic fracturing. This invention completely changes the traditional working mode that relies on manually browsing unstructured documents, realizing intelligent acquisition of professional technical knowledge. Attached Figure Description

[0029] Figure 1 This is a system architecture diagram of an embodiment of this application;

[0030] Figure 2 This is a system flowchart of an embodiment of this application;

[0031] Figure 3 Detailed diagram of document loading for embodiments of this application;

[0032] Figure 4 This is a flowchart illustrating the intelligent document segmentation process in an embodiment of this application.

[0033] Figure 5 This is a flowchart illustrating the formula processing in an embodiment of this application.

[0034] Figure 6 This is a flowchart illustrating the text vector knowledge base and text vector generation process in an embodiment of this application.

[0035] Figure 7 This is a flowchart illustrating the incremental update process in an embodiment of this application.

[0036] Figure 8 This is a further search flowchart of an embodiment of this application;

[0037] Figure 9 This is an overall flowchart of an embodiment of this application;

[0038] Figure 10 This is a classification diagram of popular topics in hydraulic fracturing according to embodiments of this application. Detailed Implementation

[0039] This application provides a hydraulic fracturing question-and-answer method and system based on cross-language retrieval enhancement generation.

[0040] For the field of hydraulic fracturing, this invention employs a dynamic text vector library architecture to protect the integrity of terminology, automatically associating Chinese and English terms to solve the problem of mixed terminology in literature and improve the accuracy of bilingual retrieval. Through intelligent parsing of professional formulas, it accurately identifies mathematical expressions in documents and treats them as independent semantic units, ensuring accurate parsing of key parameters. This system supports incremental updates to the knowledge base. As knowledge in the field of hydraulic fracturing continues to expand, the system rapidly iterates on domain knowledge without disrupting the existing knowledge base. This allows users to simply search for relevant questions in the field of hydraulic fracturing on the client side, and the system quickly matches the most suitable response from the existing multilingual knowledge base.

[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in further detail below with reference to the accompanying drawings.

[0042] This invention provides a knowledge-based question-answering method in the field of hydraulic fracturing based on retrieval enhancement generation, see [link to documentation]. Figure 9 This includes the following steps:

[0043] S1: Upon initial use, the user collects documents in a specific hydraulic fracturing field. The knowledge base is built in the knowledge base loading mode of the user operation unit on the client. The knowledge base is generated by the document processing module and vectorization module of the server. The knowledge base loading mode also supports loading existing knowledge bases and incrementally updating knowledge bases.

[0044] S2: Users select the user search mode in the user operation unit of the client to search for relevant issues in the field of hydraulic fracturing, and generate text vectors through the vectorization module of the server;

[0045] S3: Based on the text vector knowledge base generated above, perform preliminary matching with the text vector generated for the user's search question to obtain several coarse-grained results related to the user's search question;

[0046] S4: The coarse-grained results obtained above are used for fine-grained re-ranking through pre-training and re-ranking to calculate fine-grained relevance scores. The server prioritizes the results with high relevance, generates preliminary responses, and outputs the final ranking results.

[0047] S5: Based on the preset prompts and the final sorting results, generate the final response on the client side. The response includes the model's reasoning process, the answer, and relevant evidence.

[0048] Based on the same concept as the above method, this application also provides a knowledge question-answering system in the field of hydraulic fracturing based on retrieval-enhanced generation, including a client and a server, such as... Figure 1 As shown.

[0049] The client is configured with a user operation unit and a system output module, which are responsible for user interaction;

[0050] The user operation unit includes a knowledge base loading mode and a user retrieval mode;

[0051] Among them, the knowledge base loading mode is a mode that allows users to choose how to load the knowledge base;

[0052] User search mode is a mode in which users interact with the system by inputting relevant questions in the field of hydraulic fracturing.

[0053] The server is equipped with a document processing module, a vectorization module, a vector storage module, a vector retrieval module, a ranking module, and a generative question-answering module, which are responsible for core data processing, retrieval, and generation tasks.

[0054] The document processing module is used to intelligently segment multi-format documents in the same folder into document blocks, extract relevant terms and formulas, and store them separately.

[0055] The vectorization module is used to generate text vectors from document blocks (English) and user questions (Chinese) through a pre-trained embedding model, including an embedding encoding unit, a vector normalization unit, and a cosine similarity calculation unit.

[0056] The embedding coding unit is a unit that uses a multilingual pre-trained embedding model to map Chinese and English to the same high-dimensional semantic vector space, thereby enabling cross-language retrieval capabilities.

[0057] The vector normalization unit is a unit that sets the modulus of the vectors generated by the above modules to 1 to facilitate subsequent processing.

[0058] The unit for calculating cosine similarity is a unit that calculates the cosine similarity of normalized vectors, thus narrowing the distance between similar semantic vectors in the high-dimensional semantic vector space.

[0059] The vector storage module is used to normalize text vectors and store them in a specific file, supporting dynamic addition, deletion, modification and query, which is convenient for users to use repeatedly;

[0060] The vector retrieval module is based on the approximate nearest neighbor search algorithm, which quickly matches the user's retrieval vector with the stored vector library and returns coarse-grained candidate results. This is used to quickly find the results most similar to the user's question from a large number of documents, and also supports adding new text vectors.

[0061] The fine-ranking module is used to combine the coarse-grained results (candidate documents) obtained from the vector retrieval module with the user's original question and the complete content of the document through a pre-trained fine-ranking model to calculate more accurate relevance and perform a more surprising ranking.

[0062] The generative question-answering module is used to combine user questions and finely formatted document fragments to generate structured answers based on domain-customized prompt word templates using a pre-trained large language model.

[0063] Specifically, such as Figure 1 As shown, the entire system supports two working modes: knowledge base loading mode and user retrieval mode.

[0064] In this embodiment of the application, the knowledge base loading mode process is as follows (process ①):

[0065] 1. Users select the knowledge base loading mode through the user operation unit;

[0066] 2. When a user chooses to build a new knowledge base, the request is sent to the document processing module on the server side for document preprocessing;

[0067] 3. The processed document undergoes vector representation conversion via the vectorization module;

[0068] 4. The generated vectors are used to build or update the knowledge base through the vector storage module and saved on the local disk.

[0069] In this embodiment of the application, the user retrieval mode process is as follows (process ②):

[0070] 1. Users can enter the user search mode through the user operation unit, and users can use Chinese to search for English documents in the field of hydraulic fracturing;

[0071] 2. The query is sent to the server's vectorization module for vector conversion;

[0072] 3. The generated query vector is saved in the vector storage module and executed by the vector retrieval module;

[0073] 4. The vectors of relevant documents or fragments retrieved are sorted and optimized through the fine-ranking module;

[0074] 5. Finally, the "generative question-answering module" generates the answer based on the ranking results and a predetermined format;

[0075] 6. The generated answer is returned to the client's "system output module" and presented to the user.

[0076] Among them, such as Figure 2 As shown, this application provides the overall process for system startup and knowledge base construction, loading, and updating:

[0077] 1. Start the system;

[0078] 2. Configure the configuration file: Users can choose the embedded model, the fine-ranked model, the large language model, and the temperature setting of the model output. In order to meet the multilingual search query, the model that supports multilingual embedding is used, so that users can accurately search for relevant content in English documents even when they input Chinese.

[0079] 3. Before querying a question, users need to select a method for loading the knowledge base. This system provides three methods for loading the knowledge base, such as... Figure 3 As shown: (1) Construct a new knowledge base, (2) Load the knowledge base, (3) Incrementally update the knowledge base;

[0080] (1) Building a new knowledge base: After the user selects to build a new knowledge base, the document processing, text vectorization, vector storage construction, and knowledge base saving operations are performed to save the generated text vector knowledge base locally, such as Figure 4 As shown:

[0081] ① During document processing, multilingual documents are intelligently segmented into blocks, prioritizing paragraph division to enhance the semantic meaning of words in fixed sentences;

[0082] ② The system sets a segmentation threshold to judge the length of paragraphs: if it does not exceed the set threshold, the paragraph is retained as a whole; if it exceeds the set threshold, it is split into sentences.

[0083] ③ After the system segments the sentence, it judges the length of the sentence: if it does not exceed the set threshold, the complete sentence is retained; if it exceeds the set threshold, it is segmented according to technical terms.

[0084] After completing document segmentation, to enable users to effectively conduct Chinese and English searches, the system added a Chinese-English glossary for hydraulic fracturing terminology, as shown in the following example:

[0085] The original English document stated, "Hydraulic fracturing improves...".

[0086] The enhanced text reads: "Hydraulic fracturing improves...".

[0087] Specifically, the document will typically contain formulas, and the system will process the formulas according to a set procedure, such as... Figure 5 As shown:

[0088] ① The user selects the original document for processing, and the system extracts LaTeX formulas: It identifies and extracts mathematical formulas in LaTeX format from the document;

[0089] ②System stored formula list: The extracted formulas are stored separately, which may include their location information in the original text;

[0090] ③ Replace with Formula in text: In the original text, replace the extracted LaTeX formula with a general placeholder, such as "Formula", to simplify the text vectorization process and avoid the formula complexity from affecting semantic understanding;

[0091] ④ Assign formulas when dividing text: When dividing a document into blocks, ensure that each text block can be associated with its original formula information;

[0092] ⑤ Restore formulas during response: When generating the final response, placeholders are replaced with the original LaTeX formulas based on the previously stored formula list to ensure the completeness and accuracy of the response.

[0093] Specifically, after completing document segmentation and formula processing, the segmented documents are further vectorized into text. Through a pre-trained local embedding model, the user query text and the processed document are transformed into vectors.

[0094] In this embodiment, the generation process of the text vector knowledge base is completed in the vectorization module, and the generation process of the text vector knowledge base is as follows: Figure 6 As shown:

[0095] ① The system processes specific documents submitted by users, including text and formulas, and generates text vectors;

[0096] ② The system processes user queries and generates text vectors;

[0097] ③ The vectorization module processes the text: First, the embedding encoding module maps English text to a high-dimensional semantic vector space using a pre-trained multilingual embedding model; second, the vector normalization module normalizes the generated embedding vectors; finally, an efficient vector index is constructed. During the retrieval phase, the system first performs keyword pre-filtering, then calculates the cosine similarity between the query vector and the document vector, and finally uses a re-ranking model to optimize the results. Text content and corresponding vectors are stored separately, saved as document block data and vector index files respectively.

[0098] (2) Loading an existing knowledge base: The system allows users to choose to load an existing knowledge base.

[0099] (3) Incremental update of the knowledge base: As the content in the field of hydraulic fracturing is continuously updated, the system allows users to continue adding new text vectors to the existing text vector knowledge base, such as... Figure 7 As shown, it includes:

[0100] ① The user adds a processed new document to the existing vector knowledge base, and the system processes the document block;

[0101] ② Generate embeddings: Generate vector embeddings for the processed document blocks;

[0102] ③ Update the vector index: Add the newly generated vector to the index of the vector storage for subsequent retrieval;

[0103] ④ Update the keyword index: If the system also maintains a keyword index, update the index to reflect the content of the new document;

[0104] ⑤ Save the new knowledge base: Save the updated knowledge base to ensure data persistence.

[0105] 4. Interactive Q&A: The system enters a state of waiting for user queries and retrieval;

[0106] 5. Receiving User Questions: The system receives user input questions. To improve the user experience of retrieving multilingual document content using a specific language, this system selects models pre-trained on multilingual datasets when choosing embedding models, ranking models, and large language models.

[0107] 6. Semantic Cache Check: If a matching result is found in the semantic cache for the current question, the cached result is returned; if no matching result is found in the semantic cache for the current question, the following subsequent steps are performed, such as... Figure 8 As shown, it includes:

[0108] ① Generate question vectors: Transform user questions into vector representations;

[0109] ② Vector retrieval: Searching for documents or fragments similar to the question vector in the vector storage;

[0110] ③ Re-ranking: Use the pre-trained fine-ranking model to re-rank the retrieved results and optimize relevance;

[0111] ④ Constructing the context: Based on the sorting results, construct the context information used to generate the answer;

[0112] ⑤ Generate prompts: Combine user questions and context to generate prompts to be fed into the large language model;

[0113] ⑥ Call the large language model: Send the prompt to the large language model and request the generation of an answer;

[0114] ⑦ Generate Chinese answers: The large language model generates Chinese answers, which include the model's reasoning process, the answer and its corresponding source, and the answer's reference source;

[0115] ⑧ Cache results: Store the results of this question and answer in the semantic cache so that the same or similar questions can be returned directly in the future;

[0116] ⑨ Return Result: Return the final Chinese answer to the user.

[0117] also, Figure 10 This section introduces some popular topic categories covered by this system in the field of hydraulic fracturing:

[0118] 1. Geomechanics and fracture mechanisms: Focusing on rock mechanical response, fracture propagation mechanisms, and stress field changes during hydraulic fracturing;

[0119] 2. Digital Technology and Intelligentization: Involving digital modeling, intelligent control, data analysis and optimization of hydraulic fracturing;

[0120] 3. New Materials and Chemical Additives: Focus on the research and application of new materials such as fracturing fluids and proppants, as well as the performance and function of various chemical additives;

[0121] 4. Shale Gas and Unconventional Oil and Gas Development: Focusing on the application technologies and challenges of hydraulic fracturing in the development of unconventional oil and gas reservoirs such as shale gas and tight oil;

[0122] 5. Other hydraulic fracturing topics.

[0123] Through the above hierarchical description, the knowledge question-answering system and method in the field of hydraulic fracturing based on retrieval enhancement not only possesses powerful knowledge base management and updating capabilities, but also provides efficient, accurate, and professional question-answering services through vector retrieval and large language models. The innovation of this invention lies in:

[0124] 1. It fills the gap in hydraulic fracturing without the need for enhanced generation.

[0125] This system innovatively introduces RAG technology into the field of hydraulic fracturing, constructs a professional knowledge base, and integrates large language model generation capabilities, solving the adaptation problem of traditional static knowledge bases in complex scenarios involving multiple disciplines (geomechanics, fluid mechanics, etc.). By calling the latest technical data in real time, the system's timeliness and professionalism are significantly improved.

[0126] 2. Enhanced cross-linguistic competence and mapping of Chinese and English terminology

[0127] This system pioneers a bidirectional mapping and collaborative retrieval system for Chinese and English knowledge. It utilizes multilingual vector embedding technology to construct a unified knowledge graph, resolving the coexistence of English literature and Chinese standards in the field of hydraulic fracturing. Users can input Chinese queries (such as "fracturing fluid additive selection"), and the system can automatically retrieve and integrate Chinese and English materials, improving the accuracy and reliability of cross-language question answering.

[0128] 3. Multimodal parsing capability for complex documents

[0129] This system employs document structuring (such as table recognition and formula conversion) to transform unstructured text, such as hydraulic fracturing technical reports and experimental data, into searchable vectors, enabling precise parsing of parameters and formulas. For example, when a user queries "fracturing fluid viscosity calculation formula," the system can automatically extract the mathematical expressions from the PDF and generate executable logic, reducing human error.

[0130] 4. Dynamic knowledge base update and lossless incremental mechanism

[0131] This system employs a vector-index-based dynamic update mechanism, supporting lossless incremental updates of new technical data without disrupting the existing knowledge base. By synchronizing industry databases, user feedback, and expert review in real time, the system ensures that the knowledge base content remains up-to-date with the latest technological developments. For example, when a fracturing technology is banned due to new regulations, the system can automatically mark relevant entries and prioritize searching for alternative solutions. Simultaneously, a caching mechanism avoids redundant searches, improving efficiency.

[0132] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for generating hydraulic fracturing question-and-answer based on cross-language retrieval enhancement, characterized in that, Includes the following steps: The client inputs collected hydraulic fracturing documents, and the server generates a text vector knowledge base. Based on the collected hydraulic fracturing documents, hydraulic fracturing-related questions are retrieved on the client side, and the server generates text vectors. The text vector knowledge base is initially matched with the text vector to obtain coarse-grained results; The coarse-grained results are then finely sorted using a pre-trained re-sorting process, and the final sorted results are output. The client generates a final response based on the server's preset prompt word rules and the final sorting results.

2. The method for generating hydraulic fracturing question-and-answer based on cross-language retrieval enhancement as described in claim 1, characterized in that, The client includes a user operation unit and a system output module, which are responsible for user interaction.

3. The method for generating hydraulic fracturing question-and-answer based on cross-language retrieval enhancement according to claim 2, characterized in that, The user operation unit includes a knowledge base loading mode and a user retrieval mode; The knowledge base loading mode provides users with the option to select how to load the knowledge base, including three methods: building a new knowledge base, loading an existing knowledge base, and incrementally updating the knowledge base. The user search mode retrieves hydraulic fracturing-related questions and system interactions.

4. The method for generating hydraulic fracturing question-and-answer based on cross-language retrieval enhancement according to claim 3, characterized in that, The system's startup and knowledge base loading include: Start the system; Configure the configuration file to support multiple speech embedding models; Choose the method of loading the knowledge base to search for the question; Interactive question and answer; Receiving user questions; Semantic cache check.

5. The method for generating hydraulic fracturing question-and-answer based on cross-language retrieval enhancement according to claim 4, characterized in that, The construction of the new knowledge base involves converting multi-format documents into a text vector library for storage, including: The original document is divided into blocks to form document blocks; Set a segmentation threshold; if the threshold is not exceeded, the segmentation is retained; otherwise, it is split into sentences. Determine the sentence length; if it is within the limit, retain it; otherwise, split it according to terms.

6. The method for generating hydraulic fracturing question-and-answer based on cross-language retrieval enhancement according to claim 5, characterized in that, The original document is processed using a flow formula, including: Identify and extract mathematical formulas in LaTeX format; The system stores a list of formulas; Replace LaTeX formulas with generic placeholders; Assign formulas when dividing text into blocks; When answering, restore the formula to LaTeX.

7. The method for generating hydraulic fracturing question-and-answer based on cross-language retrieval enhancement according to claim 6, characterized in that, The processed document blocks are vectorized into text to generate a text vector knowledge base, including: Process user-submitted documents; Handling user queries; The vectorization module processes the text, including embedding encoding, vector normalization, and cosine similarity calculation.

8. The method for generating hydraulic fracturing question-and-answer based on cross-language retrieval enhancement according to claim 4, characterized in that, The incremental update database refers to adding new text vectors to an existing knowledge base, including: Add new documents to the existing text vector knowledge base and process document blocks; Generate vector embeddings from the processed document blocks; Update vector index; Update the keyword index; Save the new text vector knowledge base.

9. The method for generating hydraulic fracturing question-and-answer based on cross-language retrieval enhancement according to claim 4, characterized in that, The semantic cache check: if a match is successful, the cache result is returned; Otherwise, proceed with subsequent processing, including: Transform user questions into vector representations; Find documents or fragments similar to the question vector; The retrieved results are reordered using the pre-trained fine-ranking model that was set up. Contextual information for generating responses is constructed based on the sorting results; Based on the user's question and context, generate suggestions to be fed into the large language model; Send the prompt to the large language model and request it to generate a response; Generate Chinese answers, including the reasoning process, the answer and its corresponding source, and a reference to the answer. Store the answer in a semantic cache so that it can be returned directly for the same or similar questions in the future; Return the result.

10. A hydraulic fracturing question-and-answer system based on cross-language retrieval enhancement, characterized in that, It includes a client and a server; the client inputs the collected documents, and the server generates a text vector knowledge base. The client includes a user operation unit and a system output module; The user operation unit includes a knowledge base loading mode and a user retrieval mode; The knowledge base loading mode provides users with the option to select how to load the knowledge base. The user search mode allows users to search for relevant issues and system interactions in the field of hydraulic fracturing. The server includes a document processing module, a vectorization module, a vector storage module, a vector retrieval module, a ranking module, and a generative question answering module; The document processing module intelligently divides multi-format documents in the same folder into document blocks, extracts terms and formulas, and stores them independently. The vectorization module, based on a pre-trained embedding model, encodes document blocks and user questions into high-dimensional text vectors, generating semantic numerical representations to support similarity retrieval. The vector storage module normalizes and persistently stores the text vector, supporting dynamic addition, deletion, modification, and querying of files; The vector retrieval module quickly matches the user's retrieval vector with the vector library using an approximate nearest neighbor algorithm, and returns a coarse-grained candidate set. The fine-ranking module returns coarse-grained results from the vector retrieval module by combining a pre-trained fine-ranking model with user questions and document content. The generative question-answering module uses a pre-trained large language model to generate a structured answer based on domain-customized prompt word templates, taking the user's question and the refined document fragment as input.

Citation Information

Patent Citations

  • Knowledge enhancement retrieval generation method, device and system for multivariate data fusion

    CN120144720A

Cited By

  • RAG intelligent retrieval question-answering system and method based on enhanced metadata

    CN121256007A

  • AI test question generation method and system based on real-time LaTeX conversion

    CN121580993A

  • An ai test question generation method and system based on real-time latex conversion

    CN121580993B

  • Large language model and knowledge graph fusion method based on semantic embedding model

    CN122311394A