Retrieval enhancement generation method and system oriented to chemical corpus data
By building a special vocabulary and weighted map for chemical knowledge and timely updating terms, the problem that big models have difficulty understanding professional terms and new terms in the chemical field is solved, and the accuracy and adaptability of the answers are improved.
Patent Information
- Application Number
- CN202510173012.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-17
AI Technical Summary
In the field of chemical industry, it is difficult for existing large models and RAG technologies to accurately understand and identify professional terms, resulting in insufficient accuracy and relevance of information retrieval and answer generation, and the inability to learn and apply emerging chemical terms in a timely manner.
Build a special and common lexicon for chemical knowledge, provide accurate term interpretation and semantic understanding for large models, design a weighted graph of chemical knowledge, prioritize key documents through weight allocation mechanisms, and introduce dynamic adaptive mechanisms to update terms in a timely manner.
The quality and accuracy of the answers of large models in the chemical field are improved, ensuring that key information is not missed, and timely adapting to changes in the chemical field knowledge, providing more accurate and timely answers.
Smart Images

Figure CN120104741A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of chemical data processing, and in particular to a retrieval enhancement generation method and system for chemical corpus data. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] In recent years, large models have shown rapid development momentum in the field of natural language processing. With their powerful computing power and massive parameters, they have learned a rich variety of knowledge and demonstrated impressive performance in various tasks. However, large models also have inherent limitations. Their knowledge reserves are frozen at the moment when training is completed. They cannot acquire and update new knowledge generated after training in real time. In addition, they are limited by model capacity and training data selection, and it is impossible to include all knowledge. This leads to the fact that in practical applications, when the knowledge queried by users exceeds the scope of the large model training data, the large model often finds it difficult to give accurate and satisfactory answers, and the performance will drop significantly, which cannot meet the user's needs for the most accurate and comprehensive knowledge.
[0004] In order to meet this challenge, the Retrieval Enhanced Generation (RAG) algorithm came into being. RAG innovatively integrates external data into the learning and answer generation process of the big model. It is like an intelligent search engine. It can quickly find the latest data or related historical conversation records in external data sources according to the user's query needs, and then generate more informative prompts (Prompt), so as to guide the big model to generate answers that are more in line with user expectations. This method has to some extent broken through the bottleneck of lagging knowledge updates of big models, improved the accuracy and timeliness of questions and answers, and brought new ideas and methods for knowledge acquisition and information interaction. However, RAG technology still faces many technical difficulties in practical applications. For example, due to the imperfect retrieval mechanism, documents that are ranked high and contain key information may be missed, thus affecting the quality of answers. Moreover, for professional terms in the industry, RAG technology's ability to understand is still relatively weak, and it is difficult to accurately grasp their meaning and contextual relationships. This is particularly prominent in highly professional fields, which limits its in-depth application and promotion in specific fields.
[0005] Faced with a huge amount of chemical data, it is a very time-consuming and labor-intensive task to quickly and accurately find the required information from it. The emergence of large models has provided assistance for the retrieval and information mining of chemical data. However, in the actual application process, there are a series of problems that need to be solved. First, the chemical industry has a large number of unique and highly professional terms, which have specific meanings and complex conceptual systems. When facing these professional terms, large models often cannot accurately understand and identify their connotations, resulting in a large deviation between the output information and actual needs, and the accuracy is greatly reduced. Secondly, although RAG technology can broaden the breadth and depth of knowledge in the chemical industry to a certain extent, its classic retrieval method is prone to miss documents containing key information, making it impossible to effectively query and utilize some important chemical knowledge, affecting the decision-making process of research and production. Furthermore, new terms and concepts are constantly emerging in the chemical industry. Due to the lag in the training and updating of large models, these new terms are often not incorporated into their knowledge system in a timely manner, resulting in the inability of large models to give effective answers when facing related questions, further restricting the application effect of large models in the chemical industry. Summary of the invention
[0006] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a retrieval enhancement generation method and system for chemical corpus data. First, the present invention additionally builds a special word library and a general word library for chemical knowledge, and provides a comprehensive and accurate chemical terminology resource for the large model, so that the large model and RAG technology can better extract and understand the professional terms of the chemical industry, and effectively solve the problem that the chemical industry terms are difficult to be accurately extracted and identified. Secondly, by designing a weighted graph of chemical knowledge, different weights are given to all entity words and text blocks in the graph, and different degrees of attention can be given to each knowledge element according to these weights during the retrieval process, so as to ensure that those key and important documents can be retrieved first, and the problem of missing key documents in RAG technology is effectively solved. Finally, a dynamic adaptive mechanism is introduced, which can dynamically enrich the content of the chemical knowledge special word library according to the update and change of chemical field knowledge, and timely incorporate newly emerging chemical terms into it, so as to ensure that the large model can quickly learn and master the latest knowledge, and solve the problem that new terms cannot be learned and applied in time, thereby significantly improving the application performance and effect of the large model in the chemical field, and providing strong technical support for research, production and innovation in the chemical field.
[0007] In order to achieve the above object, the present invention adopts the following technical solution:
[0008] A first aspect of the present invention provides a retrieval enhancement generation method for chemical engineering corpus data.
[0009] A retrieval enhancement generation method for chemical engineering corpus data, comprising:
[0010] Construct a chemical knowledge vocabulary and a general vocabulary, cut the pre-processed chemical corpus into blocks, obtain several chemical text blocks, and perform vectorization processing on the chemical knowledge vocabulary, general vocabulary and chemical text blocks;
[0011] Calculate the similarity between each entity word and the chemical text block in the vectorized chemical knowledge vocabulary, assign different weights to the association relationship between each entity word and the chemical text block according to the size of the similarity, and construct a first mapping relationship; for the entity words that form a mapping relationship with the chemical text block, calculate the similarity between each entity word, assign different weights to the association relationship between each entity word according to the size of the similarity, and construct a second mapping relationship; based on the first mapping relationship and the second mapping relationship, construct a chemical knowledge weighted graph;
[0012] According to the questions asked by the user, a number of question entity words are extracted in combination with the chemical knowledge special word library and the general word library; according to the question entity words, the corresponding chemical text is searched in the chemical knowledge weighted graph, if retrieved, the answer is output according to the first mapping relationship and the second mapping relationship, otherwise, the question entity words are added to the chemical knowledge weighted graph and the chemical knowledge special word library.
[0013] Furthermore, the first mapping relationship is: a mapping relationship of <entity word chemical text block weight>, and the second mapping relationship is: a mapping relationship of <entity word entity word weight>.
[0014] Furthermore, according to the question entity word, the corresponding chemical text is searched in the chemical knowledge weighted graph, and if found, the answer is output according to the first mapping relationship and the second mapping relationship; the method includes:
[0015] According to the problem entity word, the corresponding chemical text is retrieved in the chemical knowledge weighted graph, sorted according to the weight in the first mapping relationship, and all chemical text blocks that meet the extraction threshold of the text block that fully matches the entity word are selected to obtain the text block that fully matches the entity word;
[0016] According to the problem entity word, entity words with associated relationships are retrieved in the chemical knowledge weighted graph according to the second mapping relationship, and the retrieved entity words are sorted according to the weights, and several entity words that meet the associated entity word extraction threshold are selected to obtain associated entity words; for each associated entity word, the corresponding chemical text block is retrieved in the chemical knowledge weighted graph, and the retrieved chemical text blocks are sorted according to the weights in the first mapping relationship, and all chemical text blocks that meet the associated entity word text block extraction threshold are selected to obtain the associated entity word text block;
[0017] Combine the text blocks that fully match the entity words with the text blocks that are related to the entity words to form a contextual text block that matches the user's question, input it into the big model, and get the answer.
[0018] Furthermore, the method of adding the problem entity words to the chemical engineering knowledge weighted graph and the chemical engineering knowledge special word library includes:
[0019] Calculate the similarity between each problem entity word and the chemical text block, and for the chemical text block associated with the problem entity word, re-rank the similarities between all existing entity words in the chemical text block and the chemical text block, and the similarities between the problem entity word and the chemical text block, and reallocate the weights;
[0020] After assigning the weights, the problem entity words are added to the chemical knowledge vocabulary, and the new weights and mapping relationships are saved back into the chemical knowledge weighted graph.
[0021] Furthermore, the preprocessing of the chemical engineering corpus also includes completing missing values of the chemical engineering corpus, deleting interfering data, checking whether the data is correct, and converting the chemical engineering data into a plain text format.
[0022] Furthermore, the missing value completion includes filling the missing values using an extrapolation method and filling the missing values using a smoothing filling method.
[0023] A second aspect of the present invention provides a retrieval enhancement generation system for chemical engineering corpus data.
[0024] A retrieval enhancement generation system for chemical engineering corpus data, comprising:
[0025] The construction and preprocessing module is configured to: construct a chemical knowledge vocabulary and a general vocabulary, cut the preprocessed chemical corpus into blocks to obtain a number of chemical text blocks, and perform vectorization processing on the chemical knowledge vocabulary, the general vocabulary and the chemical text blocks;
[0026] A graph construction module is configured to: calculate the similarity between each entity word and the chemical text block in the vectorized chemical knowledge special word library, assign different weights to the association relationship between each entity word and the chemical text block according to the size of the similarity, and construct a first mapping relationship; for the entity words that form a mapping relationship with the chemical text block, calculate the similarity between each entity word, assign different weights to the association relationship between each entity word according to the size of the similarity, and construct a second mapping relationship; based on the first mapping relationship and the second mapping relationship, construct a chemical knowledge weighted graph;
[0027] The output module is configured as follows: according to the questions asked by the user, a number of question entity words are extracted in combination with the chemical knowledge special vocabulary and the general vocabulary; according to the question entity words, the corresponding chemical text is searched in the chemical knowledge weighted graph, if retrieved, the answer is output according to the first mapping relationship and the second mapping relationship, otherwise, the question entity words are added to the chemical knowledge weighted graph and the chemical knowledge special vocabulary.
[0028] A third aspect of the present invention provides a computer-readable storage medium.
[0029] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the retrieval enhancement generation method for chemical engineering corpus data as described in the first aspect above.
[0030] A fourth aspect of the present invention provides a computer device.
[0031] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the retrieval enhancement generation method for chemical engineering corpus data as described in the first aspect above are implemented.
[0032] A fifth aspect of the present invention provides a computer program product or a computer program.
[0033] The present invention provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the retrieval enhancement generation method for chemical engineering corpus data as described in the first aspect above.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1. In view of the large number and complexity of professional terms in the chemical industry, it is difficult for large models and RAG technology to accurately understand the meaning and semantics of these terms, resulting in the inability to accurately extract knowledge related to professional terms during information retrieval and answer generation, affecting the accuracy and relevance of output information and other technical problems. The present invention provides an accurate and comprehensive term interpretation and semantic understanding basis for large models by building a special word library for chemical knowledge, so that it can effectively identify and process chemical professional terms, thereby improving the quality and accuracy of answers to questions in the chemical industry.
[0036] 2. In view of the technical problems such as the lack of an effective weight allocation mechanism in the retrieval process of the classic RAG technology, it is easy to miss the top-ranked documents containing key information, and it is impossible to provide the most valuable knowledge reference for the large model. The present invention designs a weighted graph of chemical knowledge, assigns reasonable weights to each entity word and text block, and can differentiate the different knowledge elements according to the weights during retrieval, and preferentially retrieves documents with higher relevance and greater importance to the question, ensuring that key information is not missed, improving the quality and effectiveness of the retrieval results, and optimizing the basis for the large model's answers.
[0037] 3. With the continuous development of the chemical industry, new terms continue to emerge, but the training data update of the large model is lagging, which leads to the inability to learn and master new chemical terms in time, and the inability to give effective answers to questions involving new terms. The present invention constructs a dynamic adaptive mechanism to automatically incorporate newly emerging terms into the chemical knowledge vocabulary and make corresponding adjustments to the weighted graph, so that the large model can quickly adapt to changes in knowledge, maintain the ability to learn and apply the latest chemical knowledge, and ensure that accurate and timely answers can always be provided when facing constantly updated problems in the chemical field. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0039] Figure 1 It is a flow chart of the retrieval enhancement generation method for chemical engineering corpus data shown in the present invention. DETAILED DESCRIPTION
[0040] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0041] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0042] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0043] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods and systems according to various embodiments of the present disclosure. It should be noted that each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the flowchart and / or block diagram, and the combination of boxes in the flowchart and / or block diagram can be implemented using a dedicated hardware-based system that performs a specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0044] Embodiment 1
[0045] like Figure 1 As shown, the present embodiment provides a retrieval enhancement generation method for chemical corpus data. The present embodiment uses the method applied to a server as an example for illustration. It is understandable that the method can also be applied to a terminal, and can also be applied to a terminal, a server, and a system, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and this application is not limited here. In the present embodiment, the method comprises the following steps:
[0046] Step 1: Construct a chemical knowledge-specific vocabulary and a general vocabulary, cut the preprocessed chemical corpus into blocks to obtain several chemical text blocks, and vectorize the chemical knowledge-specific vocabulary, general vocabulary and chemical text blocks.
[0047] The process of constructing a chemical knowledge special word library includes: collecting and arranging professional terms or entity terms commonly used in the chemical industry, and constructing a chemical knowledge special word library in the format of <entity word interpretation>, wherein interpretation refers to a detailed explanation of professional terms or entity discourse. The process of constructing a general word library includes: collecting and arranging professional terms or entity terms in authoritative dictionaries such as the "Modern Chinese Dictionary", and constructing a general word library in the format of <entity word interpretation>, wherein interpretation refers to a detailed explanation of professional terms or entity discourse.
[0048] The present invention provides a large model with rich chemical terminology explanations and semantic association information through specially built chemical knowledge word libraries and general word libraries, so that it can deeply understand the connotation and extension of professional terms in the chemical industry. When facing user questions, the large model can accurately extract the professional terms and accurately grasp the core points of the problem, thereby generating answers that are more in line with the professional knowledge system and practical application scenarios in the chemical industry, greatly improving the fit between the answer and the user's expectations. Whether it is a complex consultation on the principle of chemical reactions or a production process problem of a specific chemical product, more accurate and professional answers can be obtained, which effectively solves the problem of answer deviation caused by inaccurate understanding of terms and improves the user's satisfaction and trust in knowledge acquisition in the chemical industry.
[0049] In one or more embodiments, the process of preprocessing the chemical engineering corpus includes:
[0050] (1) There are missing data in the original data
[0051] ① Use extrapolation to fill in. For missing data in the chemical industry, such as reaction temperature, reaction time, reactant concentration, etc., treat the missing values as dependent variables, select other related variables (such as reaction yield) as independent variables, establish a regression model, and then use the regression model to extrapolate and fill in the missing values; if the necessary data is missing to establish a regression model, find other complete data similar to the data where the missing data is located, and estimate and fill in the missing values based on the corresponding values of the similar data; if the corresponding values of similar data cannot be found, randomly select a value with similar key characteristics (such as similar raw material ratio characteristics) from all complete data to fill in the missing values.
[0052] ② Use the smoothing filling method to fill in. Calculate the mean of the missing value in the data class (e.g. if the absorbance value of a certain band is missing, calculate the mean of the missing value in the spectral data class of the same substance), and then use the mean to fill in the missing value; if the mean cannot be calculated, arrange the data in order of size, take the value in the middle as the median, and use the median to fill in the missing value; you can also assign different weights to different data points based on certain characteristics or weights of the data, and then calculate the weighted average to fill in the missing value.
[0053] (2) Delete headers, footers, special characters, garbled characters, redundant fields, HTML tags and other interfering data.
[0054] (3) With the help of automated tools and manual inspection, check whether the actual data is consistent with the pre-defined data type of the chemical data field; set a reasonable value range for the data and check whether the data is within the range; check based on the logical relationship between the data, check for typos and incorrect data, such as in the chemical process flow, check whether the raw material input and product output should conform to a certain chemical stoichiometric relationship; check whether there are typos and incorrect data (such as mistakenly writing "sodium hydroxide" as "sodium hydride").
[0055] (4) Convert chemical data into plain text format; use existing software to convert chemical data into plain text; convert text data into simplified Chinese, convert digital data into Arabic numerals, etc.
[0056] In one or more embodiments, the process of cutting the chemical corpus into blocks to obtain a number of chemical text blocks specifically includes: cutting the chemical corpus into blocks to form chemical text blocks with fixed sizes. Define parameter thresholds such as chunk_size (the maximum length of each block) and chunk_overlap (the size of the overlapping parts between blocks), use the split_text method, and recursively split the text according to the "\n\n" paragraph segmentation rule; if the text block does not meet the chunk_size, then cut the text block again according to the "\n" line break segmentation rule; if the text block still does not meet the chunk_size, then cut the text block again according to the "." period segmentation rule; if the text block still does not meet the chunk_size, then cut the text block again according to the "," comma segmentation rule. Finally, use the merge_splits method to merge the cut text blocks into chemical text blocks that meet the chunk_size and the overlapping parts between blocks meet the chunk_overlap.
[0057] In some embodiments, the process of vectorizing the chemical knowledge-specific vocabulary and the general vocabulary includes: converting the chemical knowledge-specific vocabulary and the general vocabulary into vector representation using the M3E vector embedding model.
[0058] More specifically, the process of vectorizing the chemical text block includes: converting the chemical text block into a vector representation using the M3E vector embedding model.
[0059] Step 2: Calculate the similarity between each entity word and the chemical text block in the vectorized chemical knowledge vocabulary, assign different weights to the association relationship between each entity word and the chemical text block according to the size of the similarity, and construct a first mapping relationship; for the entity words that form a mapping relationship with the chemical text block, calculate the similarity between each entity word, assign different weights to the association relationship between each entity word according to the size of the similarity, and construct a second mapping relationship; based on the first mapping relationship and the second mapping relationship, construct a weighted graph of chemical knowledge.
[0060] The process of calculating the similarity between each entity word in the vectorized chemical knowledge word library and the chemical text block includes: in the vectorized chemical knowledge word library and the chemical text block, according to the entity words and definitions in the chemical knowledge word library, respectively calculating the similarity between each entity word Block with chemical text The cosine similarity of (the formula is as follows).
[0061]
[0062] In one or more implementations, the method of assigning different weights to the association relationship between each entity word and the chemical text block according to the size of the similarity to construct a first mapping relationship includes:
[0063] For each chemical text block, the cosine similarity between the entity words in the chemical knowledge word library and the chemical text block can be obtained through the above steps. These similarities are sorted from large to small, and different weights are assigned to the association relationship between each entity word and the chemical text block according to the similarity (high weight is assigned to high similarity, and low weight is assigned to low similarity). The first mapping relationship, i.e., the mapping relationship of <entity word chemical text block weight>, is formed.
[0064] In some embodiments, for entity words that form a mapping relationship with a chemical text block, the similarity between each entity word is calculated, and different weights are assigned to the association relationship between each entity word according to the size of the similarity, to construct a second mapping relationship; the method includes: for entity words that form a mapping relationship with a chemical text block, the cosine similarity between each entity word is calculated. Different weights are assigned to the association relationship between each entity word according to the size of the cosine similarity (high weights are assigned to high similarities, and low weights are assigned to low similarities). A second mapping relationship is formed, i.e., a mapping relationship of <entity word entity word weight>.
[0065] Specifically, the method constructs a weighted graph of chemical knowledge based on the first mapping relationship and the second mapping relationship; including: combining the two mapping relationships of <entity word entity word weight> and <entity word chemical text block weight>, storing them in the neoj4 graph database, and constructing a weighted graph of chemical knowledge for subsequent data retrieval.
[0066] Step 3: Based on the questions asked by the user, a number of question entity words are extracted in combination with the chemical knowledge special vocabulary and the general vocabulary; based on the question entity words, the corresponding chemical text is searched in the chemical knowledge weighted graph. If retrieved, the answer is output according to the first mapping relationship and the second mapping relationship. Otherwise, the question entity words are added to the chemical knowledge weighted graph and the chemical knowledge special vocabulary.
[0067] Among them, according to the questions asked by the users, a number of problem entity words are extracted by combining the chemical engineering knowledge special word library and the general word library; the specific process includes: according to the questions asked by the users, the two word libraries, the chemical engineering knowledge special word library and the general word library, are combined to extract entity words, and a number of problem entity words can be obtained, some of which can be found in the chemical engineering knowledge special word library, while others cannot be found in the chemical engineering knowledge special word library.
[0068] In one or more embodiments, according to the problem entity word, the corresponding chemical text is searched in the chemical knowledge weighted map, if retrieved, the answer is output according to the first mapping relationship and the second mapping relationship, otherwise, the problem entity word is added to the chemical knowledge weighted map and the chemical knowledge special word library; the method includes:
[0069] 1. For the problem entity words, they can be found in the chemical engineering knowledge database:
[0070] (1) According to the problem entity word, the corresponding chemical text block is retrieved in the chemical knowledge weighted graph, and the retrieved chemical text blocks are sorted in descending order according to the weight in the mapping relationship of <entity word chemical text block weight>, and then the top_c_k (extraction threshold of text blocks that fully match the entity word) chemical text blocks are extracted, which are called "text blocks that fully match the entity word".
[0071] (2) The entity words with associated relationships are searched in the chemical knowledge weighted graph according to the mapping relationship of <entity word entity word weight>, and the retrieved entity words are sorted in descending order according to the weight, and then the first nc_k (associated entity word extraction threshold) entity words are selected as "associated entity words". For each associated entity word, the corresponding chemical text block is searched in the chemical knowledge weighted graph, and the retrieved chemical text blocks are sorted in descending order according to the weight in the mapping relationship of <entity word chemical text block weight>, and then the top top_nc_k (associated entity word text block extraction threshold) chemical text blocks are extracted, which are called "associated entity word text blocks".
[0072] (3) Combine the "completely matching entity word text block" with the "related entity word text block" to form a "context text block" that matches the user's question.
[0073] In the chemical industry, the method of "completely matching entity word text block" + "associated entity word text block" is used simultaneously to retrieve background knowledge blocks, and the retrieved knowledge blocks are rearranged so that knowledge that is more in line with user questions is selected into the context.
[0074] Input the "context block" into the big model and output the answer.
[0075] The construction of a weighted graph of chemical knowledge provides a more comprehensive and in-depth knowledge association network for the retrieval process. During the search, not only can directly related documents be found based on entity word matching, but also the semantic relationship of entity words in the graph can be expanded to other related entity word nodes to dig out potential and indirectly related information, which greatly broadens the breadth and depth of the search and ensures that important information that may be related to the problem is not missed. At the same time, according to the weight of the attention to the knowledge elements, the search results can be more accurately focused on the most relevant and critical parts of the problem, avoiding the interference of irrelevant or low-quality information, significantly improving the accuracy of the search, and providing strong information support for research, production and decision-making in the chemical field, effectively improving work efficiency and decision-making quality.
[0076] 2. For problem entity words that cannot be found in the chemical engineering knowledge database:
[0077] (1) In all chemical text blocks, a combination of fuzzy query and precise query is used to calculate the cosine similarity between each problem entity word and the chemical text block.
[0078] (2) For the chemical text block associated with the problem entity word, the cosine similarities between all existing entity words in the chemical text block and the chemical text block and the cosine similarities between the problem entity word and the chemical text block are reordered and the weights are reallocated.
[0079] (3) After assigning the weights, the problem entity words are added to the chemical engineering knowledge vocabulary, and the new weights and mapping relationships are saved back to the neoj4 graph database to achieve automatic updating of the vocabulary and weighted graph.
[0080] Targeted supplements and optimizations have been made to address the shortcomings of the classic RAG technology. On the one hand, by designing a weighted graph of chemical knowledge, a weight distribution mechanism has been introduced into the RAG retrieval process, so that the retrieval no longer relies solely on traditional text matching rankings, but can be sorted according to the importance and relevance of knowledge elements, avoiding the omission of key documents and improving the quality and reliability of retrieval results. On the other hand, the addition of a dynamic adaptive mechanism enables RAG technology to adapt to the dynamic changes in knowledge in the chemical field in real time, realize the automatic addition of chemical professional vocabulary, and ensure that the retrieved information is always the latest and most comprehensive, overcoming the problem of delayed answers caused by the untimely knowledge update of traditional RAG technology. These technical supplements and optimization measures have enhanced the applicability and effectiveness of RAG technology in the chemical field, and provided better support for the widespread application of large models in the chemical industry.
[0081] Embodiment 2
[0082] This embodiment provides a retrieval enhancement generation system for chemical engineering corpus data.
[0083] A retrieval enhancement generation system for chemical engineering corpus data, comprising:
[0084] The construction and preprocessing module is configured to: construct a chemical knowledge vocabulary and a general vocabulary, cut the preprocessed chemical corpus into blocks to obtain a number of chemical text blocks, and perform vectorization processing on the chemical knowledge vocabulary, the general vocabulary and the chemical text blocks;
[0085] A graph construction module is configured to: calculate the similarity between each entity word and the chemical text block in the vectorized chemical knowledge special word library, assign different weights to the association relationship between each entity word and the chemical text block according to the size of the similarity, and construct a first mapping relationship; for the entity words that form a mapping relationship with the chemical text block, calculate the similarity between each entity word, assign different weights to the association relationship between each entity word according to the size of the similarity, and construct a second mapping relationship; based on the first mapping relationship and the second mapping relationship, construct a chemical knowledge weighted graph;
[0086] The output module is configured as follows: according to the questions asked by the user, a number of question entity words are extracted in combination with the chemical knowledge special vocabulary and the general vocabulary; according to the question entity words, the corresponding chemical text is searched in the chemical knowledge weighted graph, if retrieved, the answer is output according to the first mapping relationship and the second mapping relationship, otherwise, the question entity words are added to the chemical knowledge weighted graph and the chemical knowledge special vocabulary.
[0087] In some embodiments, the first mapping relationship is: a mapping relationship of <entity word chemical text block weight>, and the second mapping relationship is: a mapping relationship of <entity word entity word weight>.
[0088] In some embodiments, the output module is specifically configured as follows: according to the question entity word, the corresponding chemical text is retrieved in the chemical knowledge weighted map, sorted according to the weight in the first mapping relationship, and all chemical text blocks that meet the threshold for extracting the text block that fully matches the entity word are selected to obtain the text block that fully matches the entity word; according to the question entity word, the entity word with an associated relationship is retrieved in the chemical knowledge weighted map according to the second mapping relationship, and the retrieved entity words are sorted according to the size of the weight, and several entity words that meet the threshold for extracting the associated entity word are selected to obtain the associated entity word; for each associated entity word, the corresponding chemical text block is retrieved in the chemical knowledge weighted map, and the retrieved chemical text blocks are sorted according to the weight in the first mapping relationship, and all chemical text blocks that meet the threshold for extracting the text block of the associated entity word are selected to obtain the associated entity word text block; the fully matched entity word text block is combined with the associated entity word text block to form a context text block that matches the user's question, which is input into the large model to obtain the answer.
[0089] In some embodiments, the output module is further configured to: calculate the similarity between each problem entity word and the chemical text block, and for the chemical text block that is associated with the problem entity word, reorder the similarities between all existing entity words in the chemical text block and the chemical text block, and the similarities between the problem entity word and the chemical text block, and reallocate the weights; after allocating the weights, add the problem entity word to the chemical knowledge special vocabulary, and save the new weights and mapping relationships back into the chemical knowledge weighted graph.
[0090] In some embodiments, the preprocessing of the chemical corpus also includes completing missing values of the chemical corpus, deleting interfering data, checking whether the data is correct, and converting the chemical data into a plain text format.
[0091] In some embodiments, the missing value completion includes filling the missing values using an extrapolation method and filling the missing values using a smoothing filling method.
[0092] Embodiment 3
[0093] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the retrieval enhancement generation method for chemical engineering corpus data described in the first embodiment are implemented.
[0094] Embodiment 4
[0095] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the retrieval enhancement generation method for chemical engineering corpus data described in the first embodiment are implemented.
[0096] Embodiment 5
[0097] This embodiment provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the retrieval enhancement generation method for chemical engineering corpus data described in the first embodiment.
[0098] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0099] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0100] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1The steps for the functions specified in one or more boxes.
[0102] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A retrieval enhancement generation method for chemical engineering corpus data, characterized in that: include: Construct a chemical knowledge vocabulary and a general vocabulary, cut the pre-processed chemical corpus into blocks, obtain several chemical text blocks, and perform vectorization processing on the chemical knowledge vocabulary, general vocabulary and chemical text blocks; Calculate the similarity between each entity word and the chemical text block in the vectorized chemical knowledge vocabulary, assign different weights to the association relationship between each entity word and the chemical text block according to the size of the similarity, and construct a first mapping relationship; for the entity words that form a mapping relationship with the chemical text block, calculate the similarity between each entity word, assign different weights to the association relationship between each entity word according to the size of the similarity, and construct a second mapping relationship; Based on the first mapping relationship and the second mapping relationship, a weighted graph of chemical engineering knowledge is constructed; According to the questions asked by users, a number of question entity words are extracted by combining the chemical knowledge special word library and the general word library; According to the question entity word, the corresponding chemical text is searched in the chemical knowledge weighted graph. If retrieved, the answer is output according to the first mapping relationship and the second mapping relationship. Otherwise, the question entity word is added to the chemical knowledge weighted graph and the chemical knowledge special vocabulary.
2. The retrieval enhancement generation method for chemical engineering corpus data according to claim 1 is characterized in that: The first mapping relationship is: a mapping relationship of <entity word chemical text block weight>, and the second mapping relationship is: a mapping relationship of <entity word entity word weight>.
3. The retrieval enhancement generation method for chemical engineering corpus data according to claim 2 is characterized in that: According to the problem entity word, the corresponding chemical text is searched in the chemical knowledge weighted graph, and if found, the answer is output according to the first mapping relationship and the second mapping relationship; the method includes: According to the problem entity word, the corresponding chemical text is retrieved in the chemical knowledge weighted graph, sorted according to the weight in the first mapping relationship, and all chemical text blocks that meet the extraction threshold of the text block that fully matches the entity word are selected to obtain the text block that fully matches the entity word; According to the problem entity word, entity words with associated relationships are retrieved in the chemical knowledge weighted graph according to the second mapping relationship, and the retrieved entity words are sorted according to the weights, and several entity words that meet the associated entity word extraction threshold are selected to obtain associated entity words; for each associated entity word, the corresponding chemical text block is retrieved in the chemical knowledge weighted graph, and the retrieved chemical text blocks are sorted according to the weights in the first mapping relationship, and all chemical text blocks that meet the associated entity word text block extraction threshold are selected to obtain the associated entity word text block; Combine the text blocks that fully match the entity words with the text blocks that are related to the entity words to form a contextual text block that matches the user's question, input it into the big model, and get the answer.
4. The retrieval enhancement generation method for chemical engineering corpus data according to claim 2 is characterized in that: The method of adding the problem entity words to the chemical engineering knowledge weighted graph and the chemical engineering knowledge special word library comprises: Calculate the similarity between each problem entity word and the chemical text block, and for the chemical text block associated with the problem entity word, re-rank the similarities between all existing entity words in the chemical text block and the chemical text block, and the similarities between the problem entity word and the chemical text block, and reallocate the weights; After assigning the weights, the problem entity words are added to the chemical knowledge vocabulary, and the new weights and mapping relationships are saved back into the chemical knowledge weighted graph.
5. The retrieval enhancement generation method for chemical engineering corpus data according to claim 2 is characterized in that: The preprocessing of the chemical engineering corpus also includes completing missing values of the chemical engineering corpus, deleting interfering data, checking whether the data is correct, and converting the chemical engineering data into a plain text format.
6. The retrieval enhancement generation method for chemical engineering corpus data according to claim 5 is characterized in that: The missing value filling includes filling the missing values using an extrapolation method and filling the missing values using a smoothing filling method.
7. A retrieval enhancement generation system for chemical engineering corpus data, characterized in that: include: The construction and preprocessing module is configured to: construct a chemical knowledge vocabulary and a general vocabulary, cut the preprocessed chemical corpus into blocks to obtain a number of chemical text blocks, and perform vectorization processing on the chemical knowledge vocabulary, the general vocabulary and the chemical text blocks; A graph construction module is configured to: calculate the similarity between each entity word and the chemical text block in the vectorized chemical knowledge special word library, assign different weights to the association relationship between each entity word and the chemical text block according to the size of the similarity, and construct a first mapping relationship; for the entity words that form a mapping relationship with the chemical text block, calculate the similarity between each entity word, assign different weights to the association relationship between each entity word according to the size of the similarity, and construct a second mapping relationship; based on the first mapping relationship and the second mapping relationship, construct a chemical knowledge weighted graph; The output module is configured as follows: according to the questions asked by the user, a number of question entity words are extracted in combination with the chemical knowledge special vocabulary and the general vocabulary; according to the question entity words, the corresponding chemical text is searched in the chemical knowledge weighted graph, if retrieved, the answer is output according to the first mapping relationship and the second mapping relationship, otherwise, the question entity words are added to the chemical knowledge weighted graph and the chemical knowledge special vocabulary.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the retrieval enhancement generation method for chemical engineering corpus data as described in any one of claims 1 to 6 are implemented.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the retrieval enhancement generation method for chemical engineering corpus data as described in any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in the retrieval enhancement generation method for chemical corpus data according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Construction method, system and device of retrieval enhancement generation system and medium
CN118797060A
Retrieval enhancement generation-based retrieval method, product, equipment and medium
CN119003795A
Response enhancement method and device based on RAG technology
CN119005330A
Method and system for generating enhanced knowledge questions and answers for mixed retrieval of heterogeneous database
CN119311831A
Generating free text representing semantic relationships between linked entities in a knowledge graph
US20200218988A1
Cited By
Federal learning-based multi-modal general-purpose model cooperative training method
CN120611772A
Collaborative training method for multi-modal general-specialized model based on federated learning
CN120611772B