Local knowledge base RAG method and device based on Bayesian reasoning
By applying Bayesian inference RAG method in the knowledge base search system, the existing system's precise extraction and context-dependent distinction problems in multimodal documents and professional field knowledge retrieval are solved, and efficient and accurate information retrieval effect is achieved.
Patent Information
- Application Number
- CN202510343520.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-21
AI Technical Summary
When existing knowledge base search systems deal with multimodal documents and professional field knowledge, it is difficult to achieve accurate extraction and distinguish context dependencies, resulting in redundant and low correlation of search results.
The local knowledge base RAG method based on Bayesian inference is used to obtain and encode knowledge data, calculate the paragraph semantic probability and professional vocabulary condition probability, and generate target search information using Bayesian formulas.
It improves the relevance and accuracy of the search results, reduces redundant information, improves the search efficiency, and meets the needs of professional fields for high-precision information retrieval.
Smart Images

Figure CN120218247A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of information retrieval, and particularly relates to a local knowledge base RAG method and device based on Bayesian inference. Background Art
[0002] When using large language models for information retrieval, their retrieval accuracy and response efficiency directly affect the quality of knowledge services and user experience. With the rapid iteration of the industry knowledge system, professional field knowledge bases are facing challenges such as a sharp increase in data scale and the complication of semantic associations. In particular, it is difficult to accurately extract the professional terms and context logical relationships hidden in multi-modal documents through traditional methods. Therefore, building retrieval enhancement based on in-depth semantic understanding has become the key breakthrough point for improving the efficiency of knowledge services.
[0003] Existing knowledge base retrieval systems mainly adopt keyword matching and vector similarity calculation modes. Although they can achieve basic semantic associations, they lack directional parsing of professional terms and cannot effectively distinguish the context dependence of domain knowledge. In addition, for the hierarchical features of paragraph semantics in the knowledge base (such as technical concept nesting and polysemy ambiguity elimination), the knowledge base retrieval system is easily interfered by the text sparsity problem, resulting in a high redundancy and weak relevance of retrieval results. In recent years, although the retrieval methods based on deep learning can enhance semantic representation capabilities, their black-box models are not optimized for the characteristics of local knowledge bases (such as data heterogeneity and dynamic update requirements). When existing knowledge base retrieval systems perform retrieval for problems, there are problems such as a low relevance and high redundancy between the background information generated by the knowledge base system and the problem, and a low accuracy of retrieval answers. Summary of the Invention
[0004] This application aims to solve at least one of the technical problems existing in the prior art. For this purpose, this application proposes a local knowledge base RAG method and device based on Bayesian inference, which improves the accuracy and reliability of problem retrieval.
[0005] In the first aspect, this application provides a local knowledge base RAG method based on Bayesian inference, and the method includes:
[0006] Obtain all knowledge data of the local knowledge base, perform paragraph classification and encoding on the knowledge data to obtain a paragraph set, and perform professional vocabulary extraction and encoding on the knowledge data to obtain a professional vocabulary data set;
[0007] Calculate the semantic probability of each paragraph in the paragraph set;
[0008] Obtain the target professional vocabulary set of the target problem, and calculate the semantic probability of each vocabulary in the target professional vocabulary set in the paragraph set;
[0009] Based on the semantic probability of each word in the target professional vocabulary set in the paragraph set, obtain the semantic probability of all words in the target professional vocabulary set in the paragraph set;
[0010] Calculate the occurrence frequency of each professional word in the target professional vocabulary set in each paragraph of the paragraph set, and based on the occurrence frequency, calculate the conditional probability of each word in the target professional vocabulary set in each paragraph;
[0011] Based on the conditional probability of each word in the target professional vocabulary set in each paragraph, the semantic probability of all words in the target professional vocabulary set in the paragraph set, and the semantic probability of each paragraph, use Bayes' formula to calculate and sort the conditional probability of the target professional vocabulary set in each paragraph, and obtain the conditional probability sorting result;
[0012] Based on the conditional probability sorting result, select paragraphs with a total number of words not exceeding a preset threshold as the target paragraph set, and the number of paragraphs in the target paragraph set is at least 1;
[0013] Generate target retrieval information based on the target paragraph set and the target question.
[0014] According to an embodiment of the present application, calculating the semantic probability of each paragraph in the paragraph set includes:
[0015] Use the BERT model to generate word vectors for each professional word in the professional vocabulary dataset and semantic vectors for each paragraph in the paragraph set;
[0016] Perform an average calculation on the word vectors of all professional words to obtain the average word vector of the professional vocabulary dataset;
[0017] Calculate the similarity between the semantic vector of each paragraph in the paragraph set and the average word vector;
[0018] Based on the similarity, obtain the semantic probability of each paragraph in the paragraph set.
[0019] According to an embodiment of the present application, the similarity between the semantic vector of each paragraph in the paragraph set and the average word vector is calculated by the cosine similarity formula, and the calculation formula of the cosine similarity is as follows:
[0020]
[0021] Where VA is the average word vector, VB j is the semantic vector of the j-th paragraph, and SimB j is the similarity between the semantic vector of the j-th paragraph and the average word vector.
[0022] According to an embodiment of the present application, obtaining the semantic probability of each paragraph in the paragraph set based on the similarity includes:
[0023] Normalize the similarity to obtain the semantic probability of each paragraph in the paragraph set. The normalization calculation formula is as follows:
[0024]
[0025] where SimB j is the similarity between the semantic vector of the j-th paragraph and the average word vector, and P(B j ) is the semantic probability of the j-th paragraph.
[0026] According to an embodiment of the present application, classifying and encoding the knowledge data to obtain a paragraph set includes:
[0027] Convert the format of the knowledge data to obtain the first knowledge text data;
[0028] Preprocess the first knowledge text data to obtain the second knowledge text data;
[0029] Based on the content logic and the word count threshold, classify and encode the second knowledge text data to obtain a paragraph set.
[0030] According to an embodiment of the present application, the calculation formula for the conditional probability of the target professional vocabulary set in each paragraph is as follows:
[0031]
[0032] where P(B j / (A1,A2...A n )) is the conditional probability of the target professional vocabulary set in the j-th paragraph, P((A1,A2...A n ) / B j ) is the semantic probability of all the words in the target professional vocabulary set in the j-th paragraph, P(B j ) is the semantic probability of the j-th paragraph in the paragraph set, and P(A1,A2...A n ) is the semantic probability of all the words in the target professional vocabulary set in the paragraph set.
[0033] According to an embodiment of the present application, selecting paragraphs with a total word count not exceeding a preset threshold as the target paragraph set based on the conditional probability sorting result includes:
[0034] Based on the conditional probability sorting result, select the paragraph with the maximum conditional probability as the first paragraph set;
[0035] Calculate the total number of words in the first paragraph set, and determine whether the total number of words in the first paragraph set is greater than or equal to a preset threshold;
[0036] When the total number of words in the first paragraph set is greater than or equal to the preset threshold, use the first paragraph set as the target paragraph set;
[0037] When the total number of words in the first paragraph set is less than the preset threshold, select the paragraph with the highest conditional probability from the remaining paragraphs, add it to the first paragraph set, and jump to the step of calculating the total number of words in the first paragraph set.
[0038] In a second aspect, the present application provides a local knowledge base RAG device based on Bayesian inference, and the device includes:
[0039] An acquisition module, configured to acquire all knowledge data of the local knowledge base, perform paragraph classification and encoding on the knowledge data to obtain a paragraph set, and perform professional vocabulary extraction and encoding on the knowledge data to obtain a professional vocabulary data set;
[0040] A first processing module, configured to calculate the semantic probability of each paragraph in the paragraph set;
[0041] A second processing module, configured to acquire a target professional vocabulary set of a target problem, and calculate the semantic probability of each vocabulary in the target professional vocabulary set in the paragraph set;
[0042] A third processing module, configured to obtain the semantic probability of all vocabulary in the target professional vocabulary set in the paragraph set based on the semantic probability of each vocabulary in the target professional vocabulary set in the paragraph set;
[0043] A fourth processing module, configured to calculate the occurrence frequency of each professional vocabulary in the target professional vocabulary set in each paragraph in the paragraph set, and calculate the conditional probability of each vocabulary in the target professional vocabulary set in each paragraph based on the occurrence frequency;
[0044] A fifth processing module, configured to calculate and sort the conditional probability of the target professional vocabulary set in each paragraph using Bayes' formula based on the conditional probability of each vocabulary in the target professional vocabulary set in each paragraph, the semantic probability of all vocabulary in the target professional vocabulary set in the paragraph set, and the semantic probability of each paragraph, to obtain a conditional probability sorting result;
[0045] A sixth processing module, configured to select paragraphs with a total number of words not exceeding a preset threshold as the target paragraph set based on the conditional probability sorting result, and the number of paragraphs in the target paragraph set is at least 1;
[0046] A generation module, configured to generate target retrieval information based on the target paragraph set and the target problem.
[0047] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for the local knowledge base RAG based on Bayesian inference as described in the first aspect above is implemented.
[0048] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for the local knowledge base RAG based on Bayesian inference as described in the first aspect above is implemented.
[0049] In a fifth aspect, the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the method for the local knowledge base RAG based on Bayesian inference as described in the first aspect.
[0050] In a sixth aspect, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for the local knowledge base RAG based on Bayesian inference as described in the first aspect above is implemented.
[0051] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application.
[0052] The method for the local knowledge base RAG based on Bayesian inference provided by the present invention has the following beneficial effects compared with the prior art:
[0053] (1) By obtaining all knowledge data in the local knowledge base, classifying paragraphs, extracting and encoding professional vocabulary, combining the semantic probability of each paragraph and the conditional probability of the target professional vocabulary, and calculating and sorting the conditional probability of the target professional vocabulary set based on Bayes' formula, the present invention realizes extracting background information related to the target problem from the local knowledge base as the target retrieval information, improves the relevance and accuracy of the retrieval results, reduces redundant information, improves the retrieval efficiency, and meets the requirements of the professional field for high-precision information retrieval.
[0054] (2) By using the BERT model to generate word vectors of the professional vocabulary data set and semantic vectors of the paragraph set, and calculating the average of the word vectors of the professional vocabulary, and combining the similarity analysis of the semantic vectors of the paragraphs and the average word vectors, the present invention improves the matching degree between the paragraph semantics and the professional vocabulary, enhances the accuracy and relevance of the retrieval results, and through the semantic probability based on the similarity calculation, helps to better understand and extract knowledge information highly relevant to the target problem, improves the automation level of the knowledge base retrieval, and reduces the interference of redundant information on the retrieval quality.
[0055] (3) By calculating the cosine similarity between the semantic vector of each paragraph in the paragraph set and the average word vector, the present invention can effectively measure the semantic matching degree between the paragraph and the professional vocabulary. Through the cosine similarity formula, the association between the semantic information of the paragraph and the average word vector becomes more accurate, improving the calculation accuracy of the semantic probability, enhancing the extraction efficiency of the information in the paragraph set, and reducing the interference of redundant information. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:
[0057] Figure 1 is one of the flow diagrams of the local knowledge base RAG method based on Bayesian inference provided by the embodiments of the present application;
[0058] Figure 2 is the flow chart for calculating the conditional probability of professional vocabulary in each paragraph provided by the embodiments of the present application;
[0059] Figure 3 is the flow diagram for calculating the semantic probability of each paragraph provided by the embodiments of the present application;
[0060] Figure 4 is the flow diagram of data preprocessing provided by the embodiments of the present application;
[0061] Figure 5 is the second flow diagram of the local knowledge base RAG method based on Bayesian inference provided by the embodiments of the present application;
[0062] Figure 6 is the structural diagram of the local knowledge base RAG device based on Bayesian inference provided by the embodiments of the present application;
[0063] Figure 7 is the structural diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application fall within the scope of protection of the present application.
[0065] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or more. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.
[0066] The following will, with reference to the accompanying drawings, through specific embodiments and their application scenarios, elaborate in detail on the Bayesian inference-based local knowledge base RAG method, the Bayesian inference-based local knowledge base RAG device, the electronic device, and the readable storage medium provided by the embodiments of this application.
[0067] Among them, the Bayesian inference-based local knowledge base RAG method can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.
[0068] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablet computers having a touch-sensitive surface (for example, a touch screen display and / or a touchpad). It should also be understood that in some embodiments, the terminal may not be a portable communication device, but a desktop computer having a touch-sensitive surface (for example, a touch screen display and / or a touchpad).
[0069] In the following various embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and a joystick.
[0070] The Bayesian inference-based local knowledge base RAG method provided by the embodiments of this application, the execution subject of the Bayesian inference-based local knowledge base RAG method can be an electronic device or a functional module or functional entity in the electronic device that can implement the Bayesian inference-based local knowledge base RAG method. The electronic devices mentioned in the embodiments of this application include, but are not limited to, mobile phones, tablet computers, computers, cameras, and wearable devices, etc. The following takes an electronic device as the execution subject to illustrate the Bayesian inference-based local knowledge base RAG method provided by the embodiments of this application.
[0071] Figure 1 is one of the flow diagrams of the Bayesian inference-based local knowledge base RAG method provided by the embodiments of this application, as Figure 1As shown in the figure, the local knowledge base RAG method based on Bayesian inference includes: Step 110, Step 120, Step 130, Step 140, Step 150, Step 160, Step 170, and Step 180.
[0072] Step 110: Obtain all the knowledge data of the local knowledge base, perform paragraph classification and encoding on the knowledge data to obtain a paragraph set, and perform professional vocabulary extraction and encoding on the knowledge data to obtain a professional vocabulary data set;
[0073] It is easy to understand that in the local knowledge base, the knowledge data is composed of professional vocabulary in a certain ordered manner. Since large language models mainly process text data, it is necessary to uniformly convert various forms of content in the local knowledge base (such as Word documents, PDF files, Excel tables, pictures, videos, etc.) into semantic text in paragraph form.
[0074] Exemplarily, a picture can be converted into a descriptive text, and an Excel table can be converted into corresponding descriptive text row by row. The conversion process can be completed either manually or with the help of automated tools.
[0075] In some embodiments, the input word count of the large language model is limited. For ease of processing, the length of each text paragraph should not be too long. After performing paragraph classification and encoding on the knowledge data to obtain a paragraph set, each paragraph should not exceed 1000 words.
[0076] Furthermore, perform professional vocabulary extraction and encoding on the paragraph set to obtain a professional vocabulary data set. The extraction of professional vocabulary can be done using automated tools (such as spaCy, TF-IDF, etc.) or manually.
[0077] Step 120: Calculate the semantic probability of each paragraph in the paragraph set;
[0078] Exemplarily, let {Ai} be the professional vocabulary data set, where i corresponds to the serial number of the professional vocabulary, starting from 1 to n, P(Ai) is the probability of the professional vocabulary Ai appearing in the entire knowledge base, and let {Bj} be the paragraph set, which represents all the text paragraphs in the local knowledge base, and the range of j starts from 1 to m. P(Bj) is the semantic probability of the j-th paragraph.
[0079] Optionally, the semantic probability of the j-th paragraph can be calculated by a method based on uniform paragraph distribution. Assuming that the professional semantics of each paragraph is uniformly distributed, in this case, the calculation formula for the semantic probability of the j-th paragraph is as follows:
[0080]
[0081] It should be noted that this assumption has limitations because the amount of professional information contained in different paragraphs varies greatly and cannot reflect the actual situation.
[0082] Optionally, the semantic probability of the j-th paragraph can be calculated by a uniform distribution based on the number of characters, assuming that the professional information contained in each character in each paragraph is uniformly distributed. Since a paragraph is composed of all its characters, the probability can be calculated according to the total number of characters. The calculation formula for the semantic probability of the j-th paragraph is as follows:
[0083]
[0084] Optionally, the semantic probability of the j-th paragraph can also be calculated by a distribution based on semantic similarity. The word vector technology is used to capture the semantic features of words and paragraphs. By matching the semantic information of the j-th paragraph with that of professional vocabulary, P(Bj) is calculated according to the degree of matching. Generally speaking, the closer the semantics of the j-th paragraph is to that of the professional vocabulary, the higher the value of P(Bj). By calculating the similarity between the paragraph and the professional vocabulary in the semantic space, a semantic-based weight assignment method for P(Bj) is provided, which can more reasonably reflect the relevance between the paragraph and professional knowledge, thus realizing semantic-based probability calculation.
[0085] Step 130: Obtain the target professional vocabulary set of the target problem, and calculate the semantic probability of each vocabulary in the target professional vocabulary set in the paragraph set;
[0086] It is easy to understand that the professional vocabulary contained in the target problem is obtained to get the target professional vocabulary set, and the semantic probability P(Ai) of each vocabulary in the target professional vocabulary set in the paragraph set is calculated.
[0087] Step 140: Based on the semantic probability of each vocabulary in the target professional vocabulary set in the paragraph set, obtain the semantic probability of all the vocabulary in the target professional vocabulary set in the paragraph set;
[0088] In some embodiments, the target professional vocabulary set is {A1, A2,..., Ai,..., An}. Based on the semantic probability of each vocabulary in the paragraph set, the semantic probability of all the vocabulary in the target professional vocabulary set in the paragraph set is obtained through the following formula:
[0089] P(A1, A2...A n ) = P(A1)P(A2)...P(A n )
[0090] Step 150: Calculate the occurrence frequency of each professional vocabulary in the target professional vocabulary set in each paragraph in the paragraph set, and based on the occurrence frequency, calculate the conditional probability of each vocabulary in the target professional vocabulary set in each paragraph;
[0091] It is easy to understand that the conditional probability of each word in the target professional vocabulary set in each paragraph is calculated. The overall probability distribution of these professional words in the local knowledge base is summarized and counted. Figure 2 is a flowchart for calculating the conditional probability of professional words in each paragraph provided by the embodiments of the present application, as Figure 2 shown. This process includes the following steps:
[0092] (1) Select a paragraph Bj: Select a paragraph Bj from the paragraph set in sequence as the current processing object. If it is the first selection, select the paragraph with serial number 1. If it is not the first selection, select the next paragraph. After completion, go to step 2.
[0093] (2) Select a professional word Ai: Select a professional word Ai from the target professional vocabulary set in sequence as the current processing object. If it is the first selection, select the professional word with serial number 1. If it is not the first selection, select the next professional word. After completion, go to step 3.
[0094] (3) Count the word frequency of the current word in paragraph Bj: Use the professional word Ai and adopt a regular expression to count its occurrence frequency in the current paragraph Bj. If the professional word Ai does not appear in paragraph Bj, set its word frequency to 0. After completion, go to step 4.
[0095] (4) Determine whether all professional words have been traversed: Check whether the word frequency statistics of all words in the target professional vocabulary set have been completed. If not all have been traversed, return to step 2, select the next professional word and continue the statistics. If all have been traversed, go to step 5.
[0096] (5) Calculate P(Ai / Bj) under each paragraph Bj: After completing the word frequency statistics of all professional words in the current paragraph Bj, use a table for statistics, summarize the word frequencies of all professional words in the current paragraph Bj, calculate the total word frequency, and then divide the word frequency of each professional word in paragraph Bj by the total word frequency to obtain the probability P(Ai / Bj) under this paragraph Bj. After completion, go to step 6.
[0097] Exemplarily, the current paragraph is the second paragraph, and the statistical results for n professional words are shown in Table 1.
[0098] Table 1
[0099] Paragraph number Technical term Word frequency 2 Technical term 1 4 2 Technical term 2 6 2 Technical term 3 0 2 … … 2 Technical term i 1 2 Technical term n 1
[0100] (6) Determine whether all paragraphs have been traversed: Check whether the word frequency statistics of all paragraphs in the paragraph set have been completed. If not all have been traversed, return to step 1, select the next paragraph and continue the statistics. If all have been traversed, go to step 7.
[0101] (7) Calculate the semantic probability P(Ai) of professional vocabulary: Summarize the total word frequency of each professional vocabulary in all paragraphs, and calculate the total word frequency of the target professional vocabulary set. Based on these data, calculate the semantic probability P(Ai) of each professional vocabulary. After completion, proceed to step 8.
[0102] (8) Update the corresponding P(Ai / Bi) with zero using the semantic P(Ai): For professional vocabulary that does not appear in some paragraphs, its conditional probability P(Ai / Bi) is zero. To reduce the abnormal results caused by the numerator being zero in subsequent calculations, use the semantic probability P(Ai) to replace these values with zero. This replacement operation will not affect the final calculation result and can ensure the selection of appropriate paragraph text information. After completion, proceed to step 9.
[0103] (9) Store the probability data: Store the conditional probability P(Ai / Bj) of each professional vocabulary in each paragraph and its semantic probability P(Ai) into the database respectively for subsequent analysis and processing.
[0104] Step 160: Based on the conditional probability of each vocabulary in the target professional vocabulary set in each paragraph, the semantic probability of all vocabulary in the target professional vocabulary set in the paragraph set, and the semantic probability of each paragraph, use Bayes' formula to calculate the conditional probability of the target professional vocabulary set in each paragraph and sort it to obtain the conditional probability sorting result.
[0105] In some embodiments, the calculation formula for the conditional probability of the target professional vocabulary set in each paragraph is as follows:
[0106]
[0107] where P(B j / (A1,A2...A n )) is the conditional probability of the target professional vocabulary set in the j-th paragraph, P((A1,A2...A n ) / B j ) is the semantic probability of all vocabulary in the target professional vocabulary set in the j-th paragraph, P(B j ) is the semantic probability of the j-th paragraph in the paragraph set, and P(A1,A2...A n ) is the semantic probability of all vocabulary in the target professional vocabulary set in the paragraph set.
[0108] In some embodiments, the target set of specialized vocabulary includes two specialized vocabulary words, A1 and A2. The goal is to find the most relevant background information using these two words. Based on the conditional probability of each word in the target set of specialized vocabulary in each paragraph, the semantic probability of all words in the target set of specialized vocabulary in the paragraph set, and the semantic probability of each paragraph, the Bayesian formula is used to calculate the conditional probability P(Bj / (A1,A2)) of the target set of specialized vocabulary in each paragraph. The calculation formula is as follows:
[0109]
[0110] Among them, P((A1,A2) / Bj) is the semantic probability of all words in the target set of specialized vocabulary in the j-th paragraph, P(Bj) is the semantic probability of the j-th paragraph, and P(A1,A2) is the semantic probability of all words in the target set of specialized vocabulary in the paragraph set.
[0111] Since the specialized vocabulary words are conditionally independent of each other, we have:
[0112] P((A1,A2) / B j )=P(A1 / B j )·P(A2 / B j )
[0113] Among them, P(A1 / Bj) is the probability of observing A1 in the j-th paragraph, and P(A2 / Bj) is the probability of observing A2 in the j-th paragraph.
[0114]
[0115] Among them, Tj is the total number of specialized vocabulary words in the j-th paragraph.
[0116] P(A1,A2)=P(A1)·P(A2)
[0117]
[0118] Among them, Ttotal is the total number of specialized vocabulary words in the local knowledge base.
[0119] In this embodiment, by combining the semantic probability of each word in the paragraph and the overall semantic probability in the paragraph set to calculate the conditional probability of the target set of specialized vocabulary in each paragraph, the relevance and matching degree of the target specialized vocabulary and the paragraph set can be better evaluated, the relevance and accuracy of the retrieval results are improved, redundant information is reduced, and the retrieval efficiency is improved.
[0120] Step 170: Based on the conditional probability ranking result, select paragraphs with a total number of words not exceeding a preset threshold as the target paragraph set, and the number of paragraphs in the target paragraph set is at least 1;
[0121] In some embodiments, the preset threshold is 2000 words. All paragraphs in the paragraph set are sorted according to the magnitude of the conditional probability. The number of words in each paragraph in the paragraph set is 1000 words, and the paragraphs with the first and second conditional probabilities are selected as the target paragraph set.
[0122] Step 180, generate target retrieval information based on the target paragraph set and the target question.
[0123] Finally, input the target paragraph set and the target question into the large language model to generate target retrieval information, and the large language model generates the answer corresponding to the target question according to the target retrieval information.
[0124] According to the local knowledge base RAG method based on Bayesian inference provided by the embodiments of the present application, by obtaining all knowledge data in the local knowledge base, classifying paragraphs, extracting and encoding professional vocabulary, combining the semantic probability of each paragraph and the conditional probability of the target professional vocabulary, calculating and sorting the conditional probability of the target professional vocabulary set based on Bayes' formula, it realizes extracting background information related to the target question from the local knowledge base as the target retrieval information, improves the relevance and accuracy of the retrieval results, reduces redundant information, improves the retrieval efficiency, and meets the requirements of the professional field for high-precision information retrieval.
[0125] In some embodiments, calculating the semantic probability of each paragraph in the paragraph set includes:
[0126] Using the BERT model to generate the word vectors of each professional vocabulary in the professional vocabulary dataset and the semantic vectors of each paragraph in the paragraph set;
[0127] Performing an average calculation on the word vectors of all professional vocabularies to obtain the average word vector of the professional vocabulary dataset;
[0128] Calculating the similarity between the semantic vector of each paragraph in the paragraph set and the average word vector;
[0129] Based on the similarity, obtaining the semantic probability of each paragraph in the paragraph set.
[0130] It is easy to understand that the BERT model is used to generate the word vectors of all professional vocabularies in the professional vocabulary dataset, these word vectors are averaged to obtain the average vector in the professional vocabulary dataset, then the similarity is calculated one by one using this average vector and the semantic vectors of all paragraphs, and finally all similarity results are normalized. The normalized value is the semantic probability of each paragraph. Figure 3 is the flow schematic diagram of calculating the semantic probability of each paragraph provided by the embodiments of the present application. As Figure 3 shown, this flow includes the following steps:
[0131] (1) Select a professional term Ai: Sequentially select a professional term Ai from the professional term dataset as the current processing object. If it is the first selection, select the professional term with serial number 1; if not, select the next professional term. After completion, proceed to step 2.
[0132] (2) Obtain the professional term vector VAi: Use the BERT model to encode the currently selected professional term Ai to generate its corresponding word vector VAi (for example, the dimension of the word vector is 768). After completion, proceed to step 3.
[0133] (3) Determine whether all professional terms have been traversed: Check whether word vectors have been generated for all terms in the professional term dataset. If not, return to step 1, continue to select the next professional term and calculate its word vector. If completed, proceed to step 4.
[0134] (4) Calculate the average vector VA of professional terms: Calculate the average of the word vectors VAi of all professional terms to obtain the overall average vector VA of the professional term dataset. After completion, proceed to step 5.
[0135] (5) Select a paragraph Bj: Sequentially select a paragraph Bj from the paragraph set as the current processing object. If it is the first selection, select the paragraph with serial number 1; if not, select the next paragraph. After completion, proceed to step 6.
[0136] (6) Obtain the vector VBj of the paragraph: Use the BERT model to encode the currently selected paragraph Bj to generate its corresponding semantic vector VBj (the dimension is the same as that of the word vector of the professional term). After completion, proceed to step 7.
[0137] (7) Calculate SimBj: Use the cosine similarity calculation formula to compare the vector VBj of paragraph Bj with the average vector VA of professional terms and calculate the similarity SimBj between them. After completion, proceed to step 8.
[0138] (8) Determine whether all paragraphs have been traversed: Check whether similarities have been calculated for all paragraphs in the paragraph set. If not, return to step 5, continue to select the next paragraph and calculate its vector and similarity; if completed, proceed to step 9.
[0139] (9) Calculate P(Bj) for all paragraphs: Normalize the similarities SimBj of all paragraphs and calculate the semantic probability P(Bj) of each paragraph to ensure that the sum of the semantic probabilities of all paragraphs is 1. After completion, proceed to step 10.
[0140] (10) Store P(Bj): According to the paragraph number, store its corresponding semantic probability P(Bj) in the database to support subsequent semantic analysis and processing.
[0141] In this embodiment, by using the BERT model to generate word vectors of the professional vocabulary dataset and semantic vectors of the paragraph set, and calculating the average of the word vectors of the professional vocabulary, combined with the similarity analysis of the semantic vectors of the paragraphs and the average word vectors, the matching degree between the paragraph semantics and the professional vocabulary is improved, the accuracy and relevance of the retrieval results are enhanced. Through the semantic probability based on similarity calculation, it helps to better understand and extract knowledge information highly relevant to the target problem, improves the automation level of the knowledge base retrieval, and reduces the interference of redundant information on the retrieval quality.
[0142] In some embodiments, the similarity between the semantic vector of each paragraph in the paragraph set and the average word vector is calculated by the cosine similarity formula, and the calculation formula of the cosine similarity is as follows:
[0143]
[0144] Where VA is the average word vector, VB j is the semantic vector of the j-th paragraph, and SimB j is the similarity between the semantic vector of the j-th paragraph and the average word vector.
[0145] It should be noted that by using the same model such as BERT to obtain the corresponding word vectors {VAI, VA2,..., VAi,..., VAn} of each professional vocabulary and the word vectors {VB1, VB2,..., VBj,..., VBm} of each paragraph, if each element in these two groups of word vectors is not in the interval [0, 1], normalization methods such as Min-Max can be used to uniformly normalize these two groups of word vectors to ensure that each element in each word vector is within the range of [0, 1].
[0146] Through the cosine similarity formula, the similarity between each paragraph and the professional vocabulary dataset can be calculated. If a certain paragraph is more professional, the similarity value is larger; otherwise, the similarity value is smaller.
[0147] In this embodiment, by calculating the cosine similarity between the semantic vector of each paragraph in the paragraph set and the average word vector, the semantic matching degree between the paragraph and the professional vocabulary can be effectively measured. Through the cosine similarity formula, the association between the semantic information of the paragraph and the average word vector is made more accurate, improving the calculation accuracy of the semantic probability, enhancing the extraction efficiency of the information in the paragraph set, and reducing the interference of redundant information.
[0148] In some embodiments, obtaining the semantic probability of each paragraph in the paragraph set based on the similarity includes:
[0149] Normalize the similarity to obtain the semantic probability of each paragraph in the paragraph set. The normalization calculation formula is as follows:
[0150]
[0151] where SimB j is the similarity between the semantic vector of the j-th paragraph and the average word vector, and P(B j ) is the semantic probability of the j-th paragraph.
[0152] In this embodiment, by normalizing the similarity of each paragraph in the paragraph set, the semantic probability of each paragraph can be effectively calculated. Using normalization improves the calculation accuracy of the semantic probability, thereby improving the accuracy of paragraph semantic analysis and reducing the interference of redundant information on the results.
[0153] In some embodiments, classifying and encoding the knowledge data to obtain a paragraph set includes:
[0154] Convert the format of the knowledge data to obtain the first knowledge text data;
[0155] Preprocess the first knowledge text data to obtain the second knowledge text data;
[0156] Based on the content logic and the word count threshold, classify and encode the second knowledge text data to obtain a paragraph set.
[0157] It is easy to understand that sorting out the knowledge data in all formats in the local knowledge base, forming a unified text format, and performing professional vocabulary extraction and encoding to obtain a professional vocabulary data set, Figure 4 is the schematic flow chart of data preprocessing provided by the embodiments of the present application. As Figure 4 shown, this process includes the following steps:
[0158] (1) Data collection: Determine the scope of the local knowledge base, list all file formats that need to be sorted out (such as docx, pdf, xlsx, pictures, videos, etc.), classify and manage the files, file them according to file types. After completion, proceed to step 2.
[0159] (2) Format conversion: Convert content in various formats into a unified text format. For word documents, python-docx can be used to extract the body content and remove format information, retaining only plain text. For PDF documents, pdfplumber can be used for extraction. For scanned versions, OCR tools can be used to recognize text. For Excel tables, pandas can be used to read them, and each row of data can be converted into a descriptive statement based on column names. For pictures, OCR can be used for recognition. For videos, speech recognition can be used to extract the text in the audio and convert it into text. After completion, proceed to step 3.
[0160] (3) Text preprocessing: Remove irrelevant information such as headers, footers, page numbers, blank lines, special characters, etc. Segment the text and split long text into paragraphs not exceeding a preset threshold (e.g., 1000 words) according to the content logic. After completion, proceed to step 4.
[0161] (4) Professional vocabulary extraction: Extract professional vocabulary from the text to construct a professional vocabulary dataset. Use natural language processing tools for word segmentation and part-of-speech tagging, and use word frequency statistics to extract high-frequency professional vocabulary. Combine with a domain thesaurus or glossary to filter out common vocabulary and retain professional terms. After completion, proceed to step 5.
[0162] (5) Information encoding: Starting from 1, assign unique codes to the organized text and professional vocabulary in sequence to obtain a professional vocabulary dataset and a paragraph set.
[0163] In this embodiment, by converting the knowledge data in format and performing preprocessing, the second knowledge text data is obtained. Based on the content logic and word count threshold, it is classified into paragraphs and encoded to obtain a paragraph set, which improves the accuracy of paragraph classification and the standardization of encoding, reduces the complexity in subsequent analysis, and improves the retrieval efficiency.
[0164] In some embodiments, based on the conditional probability sorting result, selecting paragraphs with a total word count not exceeding a preset threshold as the target paragraph set includes:
[0165] Based on the conditional probability sorting result, select the paragraph with the maximum conditional probability as the first paragraph set;
[0166] Calculate the total word count of the first paragraph set and determine whether the total word count of the first paragraph set is greater than or equal to the preset threshold;
[0167] When the total word count of the first paragraph set is greater than or equal to the preset threshold, use the first paragraph set as the target paragraph set;
[0168] When the total number of words in the first paragraph set is less than the preset threshold, select the paragraph with the highest conditional probability from the remaining paragraphs, add it to the first paragraph set, and jump to the step of calculating the total number of words in the first paragraph set.
[0169] Figure 5 It is the second flowchart of the local knowledge base RAG method based on Bayesian inference provided by the embodiments of the present application. As Figure 5 shown, the method includes the following steps:
[0170] (1) The user inputs a question.
[0171] (2) Extract the target professional vocabulary set: According to the question input by the user, extract all the involved professional vocabulary, and determine whether these professional vocabulary are included in the professional vocabulary dataset. If included, form the target professional vocabulary set of this question.
[0172] (3) Determine whether the target professional vocabulary set is empty: If it is empty, it indicates that the question does not involve professional content and does not require local knowledge as the background, and directly enter step 23. If it is not empty, it means the question involves professional content, and enter step 4.
[0173] (4) Select a question professional vocabulary Aq: Select a professional vocabulary Aq from the target professional vocabulary set in sequence. If it is the first selection, select the first professional vocabulary. If it is not the first time, select the next professional vocabulary.
[0174] (5) Obtain P(Aq): Calculate the semantic probability P(Aq) corresponding to the current professional vocabulary Aq.
[0175] (6) Determine whether to traverse the target professional vocabulary set: Determine whether the semantic probabilities of all the words in the target professional vocabulary set have been obtained. If the traversal is not completed, return to step 4 to select the next professional vocabulary. If the traversal is completed, enter step 7.
[0176] (7) Calculate P(A1,A2,…,Aq): Assume that each professional vocabulary in the target professional vocabulary set is independent, and calculate the joint semantic probability P(A1,A2,…,Aq) of the target professional vocabulary set.
[0177] (8) Initialize an empty set Glist: Create an empty set Glist to store the paragraph number, the conditional probability of this paragraph under the current question condition, and the paragraph word count.
[0178] (9) Select a paragraph Bj: Select a paragraph Bj from the paragraph set in sequence as the current processing object. If it is the first selection, select the first paragraph. If it is not the first time, select the next paragraph. After completion, enter step 10.
[0179] (10) Obtain the current paragraph information: Extract the paragraph number, the semantic probability P(Bj) of this paragraph, and the number of words in the paragraph from the current paragraph Bj.
[0180] (11) Obtain the conditional probability of the target professional vocabulary set under the current paragraph: Calculate the conditional probability P(A1 / Bj), P(A2 / Bj), …, P(Aq / Bj) of each word in the target professional vocabulary set under the current paragraph.
[0181] (12) Calculate P(Bj / A1, A2, …, Aq): Based on the current paragraph Bj, calculate its comprehensive conditional probability P(Bj / A1, A2, …, Aq) under the condition of the target professional vocabulary set.
[0182] (13) Store the paragraph information in Glist: Generate an element, including the paragraph number, the conditional probability P(Bj / A1, A2, …, Aq), and the number of words in the paragraph, and add this element to the Glist set.
[0183] (14) Determine whether all paragraphs have been traversed: Determine whether the conditional probability calculation for all paragraphs in the paragraph set has been completed and stored in Glist: If not, return to step 9 to select the next paragraph; if so, proceed to step 15.
[0184] (15) Sort Glist: Sort Glist in descending order according to the conditional probability P(Bj / A1, A2, …, Aq), ensuring that the conditional probability of the first element is the largest and the conditional probability of the last element is the smallest.
[0185] (16) Initialize Count_num and i: Initialize two variables: the background text statistical variable Count_num and the Glist set sorting index variable i, both with an initial value of 0.
[0186] (17) Temp_num = the number of characters of Glist(i): Set Temp_num to the number of words in the i-th element of Glist.
[0187] (18) Accumulate the number of words in the paragraph: Add the value of Temp_num to Count_num, i.e., Count_num = Count_num + Temp_num, for counting the total number of words in the currently selected paragraph.
[0188] (19) Determine whether Count_num > 2000: Check whether Count_num exceeds 2000: If it exceeds 2000, it means that the number of characters in the background information is sufficient and no more paragraphs need to be added, and proceed to step 21; if it does not exceed 2000, it means that the background information is still insufficient, and proceed to step 20.
[0189] (20) Update the index variable i: Increment the sorting index variable i of the Glist set by 1, and continue to process the next paragraph with a later sort order.
[0190] (21) Obtain the selected paragraph numbers: According to the sorting index variable i of Glist, obtain the paragraph numbers from 0 to i - 1, which represent the selected paragraph range.
[0191] (22) Generate the final input text: Generate different final input texts according to the branch paths of each step: If it jumps directly from step 3, the final input text is the problem entered by the user itself; if it goes through step 21, the final input text is the text content corresponding to the selected paragraph plus the problem entered by the user.
[0192] (23) Submit to the large language model: Submit the finally generated input text to the large language model (such as Deepseek) and wait for the model to output an answer.
[0193] In this embodiment, by selecting the paragraph with the maximum conditional probability based on the conditional probability sorting result and calculating its total number of words, it is determined whether the preset threshold is met. If the total number of words does not reach the threshold, the remaining paragraphs are continuously selected and added to the paragraph set until the word count requirement is met. It realizes extracting background information related to the target problem from the local knowledge base as the target retrieval information, improves the relevance and accuracy of the retrieval results, reduces redundant information, improves the retrieval efficiency, and meets the requirements of the professional field for high-precision information retrieval.
[0194] In the embodiment of the present application, the RAG method for the local knowledge base based on Bayesian inference may have an execution entity as the RAG device for the local knowledge base based on Bayesian inference. In the embodiment of the present application, taking the RAG device for the local knowledge base based on Bayesian inference to execute the RAG method for the local knowledge base based on Bayesian inference as an example, the RAG device for the local knowledge base based on Bayesian inference provided by the embodiment of the present application is described.
[0195] The embodiment of the present application also provides a RAG device for the local knowledge base based on Bayesian inference, as Figure 6 shown. The RAG device for the local knowledge base based on Bayesian inference includes: an acquisition module 610, a first processing module 620, a second processing module 630, a third processing module 640, a fourth processing module 650, a fifth processing module 660, a sixth processing module 670, and a generation module 680.
[0196] The acquisition module 610 is used to acquire all knowledge data of the local knowledge base, perform paragraph classification and encoding on the knowledge data to obtain a paragraph set, and perform professional vocabulary extraction and encoding on the knowledge data to obtain a professional vocabulary data set;
[0197] The first processing module 620 is configured to calculate the semantic probability of each paragraph in the paragraph set;
[0198] The second processing module 630 is configured to obtain the target professional vocabulary set of the target problem, and calculate the semantic probability of each vocabulary in the target professional vocabulary set in the paragraph set;
[0199] The third processing module 640 is configured to obtain the semantic probability of all the vocabulary in the target professional vocabulary set in the paragraph set based on the semantic probability of each vocabulary in the target professional vocabulary set in the paragraph set;
[0200] The fourth processing module 650 is configured to calculate the occurrence frequency of each professional vocabulary in the target professional vocabulary set in each paragraph in the paragraph set, and calculate the conditional probability of each vocabulary in the target professional vocabulary set in each paragraph based on the occurrence frequency;
[0201] The fifth processing module 660 is configured to calculate and sort the conditional probability of the target professional vocabulary set in each paragraph by using the Bayesian formula based on the conditional probability of each vocabulary in the target professional vocabulary set in each paragraph, the semantic probability of all the vocabulary in the target professional vocabulary set in the paragraph set, and the semantic probability of each paragraph, so as to obtain a conditional probability sorting result;
[0202] The sixth processing module 670 is configured to select paragraphs with a total number of words not exceeding a preset threshold as the target paragraph set based on the conditional probability sorting result, and the number of paragraphs in the target paragraph set is at least 1;
[0203] The generation module 680 is configured to generate target retrieval information based on the target paragraph set and the target problem.
[0204] According to the local knowledge base RAG method based on Bayesian inference provided by the embodiments of the present application, by obtaining all the knowledge data in the local knowledge base, classifying paragraphs, extracting and encoding professional vocabulary, combining the semantic probability of each paragraph and the conditional probability of the target professional vocabulary, calculating and sorting the conditional probability of the target professional vocabulary set by using the Bayesian formula, background information related to the target problem is extracted from the local knowledge base as the target retrieval information, the relevance and accuracy of the retrieval result are improved, redundant information is reduced, the retrieval efficiency is improved, and the demand for high-precision information retrieval in the professional field is met.
[0205] The local knowledge base RAG device based on Bayesian inference provided by the embodiments of the present application can implement Figures 1 to 5 Each process implemented by the embodiment of the local knowledge base RAG method based on Bayesian inference is not described herein again to avoid repetition.
[0206] In some embodiments, such as Figure 7As shown in the figure, an embodiment of the present application further provides an electronic device 700, including a processor 701, a memory 702, and a computer program stored on the memory 702 and executable on the processor 701. When the program is executed by the processor 701, it implements each process of the above-mentioned embodiment of the local knowledge base RAG method based on Bayesian inference and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0207] It should be noted that the electronic device in the embodiment of the present application includes the above-mentioned mobile electronic device and non-mobile electronic device.
[0208] An embodiment of the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiment of the local knowledge base RAG method based on Bayesian inference and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0209] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.
[0210] An embodiment of the present application further provides a computer program product, including a computer program, which implements the above-mentioned local knowledge base RAG method based on Bayesian inference when executed by a processor.
[0211] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.
[0212] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement each process of the above-mentioned embodiment of the local knowledge base RAG method based on Bayesian inference and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0213] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a device-level chip, device chip, chip device, or on-chip device chip, etc.
[0214] It should be noted that in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising such element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0215] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the local knowledge base RAG method based on Bayesian inference in various embodiments of the present application.
[0216] In the description of the present application, "the first feature", "the second feature" may include one or more of such features.
[0217] In the description of the present application, "a plurality of" means two or more.
[0218] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit and scope protected by the claims of the present application, can also make many forms, all of which fall within the protection scope of the present application.
[0219] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.
[0220] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.
Claims
1. A local knowledge base RAG method based on Bayesian reasoning, characterized in that: The method comprises: Acquire all knowledge data of the local knowledge base, classify and encode the knowledge data into paragraphs to obtain a paragraph set, extract and encode professional vocabulary from the knowledge data to obtain a professional vocabulary data set; Calculating the semantic probability of each paragraph in the paragraph set; Obtain a target professional vocabulary set for a target question, and calculate the semantic probability of each word in the target professional vocabulary set in the paragraph set; Based on the semantic probability of each word in the target professional vocabulary set in the paragraph set, the semantic probability of all words in the target professional vocabulary set in the paragraph set is obtained; Calculate the frequency of occurrence of each professional word in the target professional vocabulary set in each paragraph in the paragraph set, and calculate the conditional probability of each word in the target professional vocabulary set in each paragraph based on the frequency of occurrence; Based on the conditional probability of each word in the target professional vocabulary set in each paragraph, the semantic probability of all words in the target professional vocabulary set in the paragraph set and the semantic probability of each paragraph, the conditional probability of the target professional vocabulary set in each paragraph is calculated and sorted using the Bayesian formula to obtain a conditional probability sorting result; Based on the conditional probability sorting result, selecting paragraphs whose total word count does not exceed a preset threshold as a target paragraph set, wherein the number of paragraphs in the target paragraph set is at least 1; Based on the target paragraph set and the target question, target retrieval information is generated.
2. The local knowledge base RAG method based on Bayesian reasoning according to claim 1 is characterized in that: The calculating the semantic probability of each paragraph in the paragraph set includes: Using the BERT model to generate a word vector for each professional vocabulary in the professional vocabulary dataset and a semantic vector for each paragraph in the paragraph set; Calculate the average word vectors of all professional vocabulary to obtain the average word vector of the professional vocabulary data set; Calculate the similarity between the semantic vector of each paragraph in the paragraph set and the average word vector; Based on the similarity, the semantic probability of each paragraph in the paragraph set is obtained.
3. The local knowledge base RAG method based on Bayesian reasoning according to claim 2 is characterized in that: The similarity between the semantic vector of each paragraph in the paragraph set and the average word vector is calculated by the cosine similarity formula. The calculation formula of the cosine similarity is as follows: Among them, VA is the average word vector, VB j is the semantic vector of the jth paragraph, SimB j is the similarity between the semantic vector of the jth paragraph and the average word vector.
4. The local knowledge base RAG method based on Bayesian reasoning according to claim 2 is characterized in that: The obtaining, based on the similarity, the semantic probability of each paragraph in the paragraph set includes: The similarity is normalized to obtain the semantic probability of each paragraph in the paragraph set. The normalized calculation formula is as follows: Among them, SimB j is the similarity between the semantic vector of the jth paragraph and the average word vector, P(B j ) is the semantic probability of the jth paragraph.
5. The local knowledge base RAG method based on Bayesian reasoning according to claim 1 is characterized in that: The paragraph classification and encoding of the knowledge data to obtain a paragraph set includes: Convert the knowledge data into a first format to obtain first knowledge text data; Preprocessing the first knowledge text data to obtain second knowledge text data; Based on content logic and word count threshold, the second knowledge text data is subjected to paragraph classification and encoding to obtain a paragraph set.
6. The local knowledge base RAG method based on Bayesian reasoning according to claim 5 is characterized in that: The calculation formula of the conditional probability of the target professional vocabulary set in each paragraph is as follows: Among them, P(B j / (A1,A2...A n )) is the conditional probability of the target professional vocabulary set in the jth paragraph, P((A1,A2...A n ) / B j ) is the semantic probability of all words in the target professional vocabulary set in the jth paragraph, P(B j ) is the semantic probability of the jth paragraph in the paragraph set, P(A1,A2...A n ) is the semantic probability of all words in the target professional vocabulary set in the paragraph set.
7. The local knowledge base RAG method based on Bayesian reasoning according to claim 1 is characterized in that: The step of selecting paragraphs whose total word count does not exceed a preset threshold as a target paragraph set based on the conditional probability sorting result includes: Based on the conditional probability sorting result, selecting the paragraph with the largest conditional probability as the first paragraph set; Calculating the total number of words in the first paragraph set, and determining whether the total number of words in the first paragraph set is greater than or equal to a preset threshold; When the total number of words in the first paragraph set is greater than or equal to a preset threshold, taking the first paragraph set as a target paragraph set; When the total number of words in the first paragraph set is less than a preset threshold, a paragraph with the largest conditional probability is selected from the remaining paragraphs and added to the first paragraph set, and the process jumps to the step of calculating the total number of words in the first paragraph set.
8. A local knowledge base RAG device based on Bayesian reasoning, implemented by the local knowledge base RAG method based on Bayesian reasoning according to any one of claims 1 to 7, characterized in that: The device comprises: An acquisition module is used to acquire all knowledge data of a local knowledge base, classify and encode the knowledge data into paragraphs to obtain a paragraph set, and extract and encode professional vocabulary from the knowledge data to obtain a professional vocabulary data set; A first processing module, used for calculating the semantic probability of each paragraph in the paragraph set; The second processing module is used to obtain a target professional vocabulary set for the target question and calculate the semantic probability of each word in the target professional vocabulary set in the paragraph set; A third processing module is used to obtain the semantic probabilities of all the words in the target professional vocabulary set in the paragraph set based on the semantic probability of each word in the target professional vocabulary set in the paragraph set; A fourth processing module is used to calculate the frequency of occurrence of each professional vocabulary in the target professional vocabulary set in each paragraph in the paragraph set, and based on the frequency of occurrence, calculate the conditional probability of each vocabulary in the target professional vocabulary set in each paragraph; A fifth processing module is used to calculate the conditional probability of the target professional vocabulary set in each paragraph and sort it based on the conditional probability of each word in the target professional vocabulary set in each paragraph, the semantic probability of all words in the target professional vocabulary set in the paragraph set and the semantic probability of each paragraph, so as to obtain a conditional probability sorting result; A sixth processing module, configured to select paragraphs whose total word count does not exceed a preset threshold as a target paragraph set based on the conditional probability sorting result, wherein the number of paragraphs in the target paragraph set is at least 1; A generating module is used to generate target retrieval information based on the target paragraph set and the target question.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the local knowledge base RAG method based on Bayesian reasoning is implemented as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the local knowledge base RAG method based on Bayesian reasoning is implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Visualization concept detector and method for constructing semantic field
CN102982072A
Method for remotely supervising the noise reduction of retrieved data
CN109063032A
Domain speech recognition method and system based on RAG
CN119296516A
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search
US20240386015A1
Service keyword extraction method, apparatus, and device, and storage medium
WO2022134575A1
Cited By
Text matching method and device based on probability distribution, equipment and storage medium
CN121029964A
Text matching method and device based on probability distribution, equipment and storage medium
CN121029964B