Hybrid enhanced indexing method and system based on vector retrieval and BM25 algorithm

By combining vector search and BM25 algorithm, using large language models to optimize word segmentation, the problem of insufficient accuracy and recall of RAG system when dealing with complex queries and fresh Internet vocabulary is solved, achieving more efficient and accurate information retrieval effects.

CN119961376APending Publication Date: 2025-05-09XIAMEN MEIYABAIKE INFORMATION SECURITY RES INST CO LTD
View PDF 0 Cites 21 Cited by

Patent Information

Application Number
CN202411928098.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing RAG systems have problems with low accuracy and insufficient recall when dealing with complex queries and fresh Internet vocabulary, especially in highly concise vocabulary queries in government affairs and emergency areas.

Method used

A hybrid enhanced indexing method based on vector search and BM25 algorithm is adopted, dense vector representation and keyword extraction are performed through the combination of vector library and vocabulary list, and word segmentation is optimized using a large language model, keyword search is combined with BM25 algorithm, and answers are finally generated through reordering and filtering.

Benefits of technology

It improves the accuracy and comprehensiveness of information retrieval, improves the ability to process complex queries and new vocabulary, and enhances the query effect in government affairs and emergency areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961376A_ABST
    Figure CN119961376A_ABST
Patent Text Reader

Abstract

The invention discloses a hybrid enhanced index and system based on vector retrieval and a BM2 algorithm, and the method comprises the steps: uploading a document by a user, analyzing the document by the system to obtain text content, segmenting the text content, converting the segmented text content into dense vectors, storing the dense vectors in a vector library, extracting keywords of segmented texts by using a large model, processing the keywords, and inserting the keywords into a word list. Word segmentation is carried out on the text based on the constructed word list and a word segmentation device; in response to the received user query, vectorizing the user query, and calculating the similarity with each vector in the vector library to obtain a preliminary query result; performing word segmentation on user query, performing keyword retrieval by using a BM25 algorithm according to a word list and a word segmentation device to obtain query results, and filtering the similarity of the two query results by the reordering model according to a set threshold value; selecting a result with the highest similarity according to large model parameter limitation and an upper limit set by business requirements; and splicing the document content corresponding to the result and the user query into a cue word, inputting the cue word into the large model, and analyzing a semantic relationship and a logic structure in the cue word to generate an answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data retrieval, and in particular to a hybrid enhanced indexing method and system based on vector retrieval and BM25 algorithm. Background Art

[0002] The Retrieval Enhancement (RAG) system retrieves relevant information from an external knowledge base and then lets the big model summarize and generalize based on this information, thereby enhancing the contextual knowledge and generation effect of the big model in processing queries. This technology aims to solve the problems that big models may have in terms of accuracy, knowledge update speed, and answer transparency.

[0003] The development of RAG technology has achieved remarkable results. RAG enables large models to obtain necessary contextual information without updating model parameters, thereby alleviating the illusion problem of large models. In addition, with the continuous maturity of large language model technology and the deepening of industry applications, the application of RAG systems is becoming more and more extensive. RAG technology also faces some challenges in practical applications. On the one hand, the accuracy of the RAG model is currently about 70%, which is far lower than expected in many application scenarios. This may be due to the model's lack of sensitivity to specific types of data or its insufficient processing capabilities for complex queries. On the other hand, the RAG system is prone to errors when processing queries involving fresh Internet vocabulary, which shows that the model is insufficient in understanding and processing specific and detailed information.

[0004] In the fields of government affairs and emergency response, users are usually accustomed to using highly concise vocabulary to ask questions. This method often has incomplete semantic information, and relying solely on vectorized retrieval may result in a low recall rate. Or, related domain keywords are lost in semantic retrieval. For example, when a user enters a piece of background information and asks about how to apply for an ID card or residence permit, the difference between "ID card" and "residence permit" may be obscured by the previous background information, resulting in unsatisfactory recall results. In the industry, the speed at which users create new words far exceeds the training speed of the embedding model. Summary of the invention

[0005] In order to solve the above technical problems existing in the prior art, the present invention proposes a hybrid enhanced indexing method and system based on vector retrieval and BM25 algorithm to solve the above technical problems.

[0006] According to a first aspect of the present invention, a hybrid enhanced indexing method based on vector retrieval and BM25 algorithm is proposed, comprising:

[0007] S1: The user uploads a document, the system parses the file to obtain the text content, segments the text content, converts it into dense vectors and stores it in the vector library, uses the large model to extract keywords from the segmented text, inserts the keywords into the vocabulary after processing, and segments the text based on the constructed vocabulary and word segmenter;

[0008] S2: In response to receiving a user query, the user query is vectorized, and the similarity with each vector in the vector library is calculated to obtain the first batch of preliminary query results; at the same time, the user query is segmented, and the keyword search is performed using the BM25 algorithm based on the vocabulary and the word segmenter to obtain the second batch of query results, and the similarity of the results of the vector matching retrieval and the BM25 keyword retrieval by the re-ranking model is filtered according to the set threshold; according to the upper limit set by the large model parameter restrictions and business needs, the result with the highest similarity is selected from the filtered results; the document content corresponding to the result and the user query are spliced ​​into prompt words, which are input into the large model, and the large model parses the semantic relationship and logical structure in the prompt words to generate answers.

[0009] In some specific embodiments, the text content is segmented in S1, converted into dense vectors and stored in a vector library, specifically including: mapping the text to a low-dimensional vector space through a pre-trained model in the Transformer architecture, and using a large-scale corpus for pre-training when training the vector representation model so that the model can learn a common semantic representation pattern; fine-tuning training is performed on the knowledge base text in the corresponding field to adapt to the domain-specific semantic features; in the vectorization process, the text is subjected to lexical, syntactic and semantic analysis, key features are extracted, and the key features are combined into a final dense vector representation.

[0010] In some specific embodiments, in the process of extracting keywords from segmented text in S1, the segmented text is first preprocessed, and the preprocessing includes removing interference information such as stop words and punctuation marks; the large model extracts keywords based on its pre-trained language model and domain knowledge by combining multiple algorithms and techniques, including using statistical methods to calculate the frequency of words and inverse document frequency indicators to determine the importance of words; at the same time, the attention mechanism in the deep learning model is used to focus on the parts of the text that are highly relevant to the topic and extract key words therein; for the extraction of domain nouns and new nouns, the large model refers to the domain glossary, industry knowledge base, and the latest language trends, and identifies domain-specific and novel keywords by comparing and analyzing the words in the text with known domain words and emerging vocabulary patterns.

[0011] In some specific embodiments, in the process of extracting keywords from segmented text in S1, when extracting names, prompt words in the following format are used to guide the large model to operate:

[0012] Role setting: Make it clear that the role of the big model in the task is a natural language processing expert;

[0013] Background: It is necessary to conduct in-depth analysis of the text to identify possible name entities;

[0014] Expert description Profile: Specializes in using algorithms and pattern recognition techniques to identify names in text

[0015] Skills: Capable of text analysis, entity recognition, pattern matching, and algorithm design, and able to extract names from text efficiently and accurately

[0016] Goals: Accurately identify and extract all names from the text, ensuring that no possible name is missed;

[0017] Constraints: impose constraints on the name extraction process. The extracted names must have textual basis and the extraction results must meet the pre-set format specifications.

[0018] OutputFormat: The output format of names is clearly defined, which facilitates the system to accurately identify and process the extraction results.

[0019] According to the hybrid enhanced indexing method based on vector retrieval and BM25 algorithm of claim 4, the workflow of name extraction includes:

[0020] Read and analyze the given text. The large model reads the text provided by the user and performs a comprehensive and in-depth analysis on it.

[0021] Identify possible name entities in the text. Based on the analysis of the text, use algorithms and pattern recognition technology to accurately locate possible name entities in the text.

[0022] Output the recognized names in the specified format. Output the recognized names in the pre-set format, with each name separated by "##".

[0023] Confirm that all possible names have been extracted. After outputting the names, the large model will check the entire text again to ensure that no possible names are missed;

[0024] Output <|COMPLETE|> The extraction process is complete.

[0025] In some specific embodiments, the keyword search using the BM25 algorithm based on the word list and the word segmenter in S2 to obtain the second batch of query results specifically includes:

[0026] Keyword preprocessing and analysis: After receiving a user query, the user query is first segmented. The word segmenter constructed in the knowledge base construction phase is used to accurately segment the user query into separate vocabulary units based on the vocabulary information in the vocabulary. In the word segmentation process, the language habits and grammatical structure characteristics of the user query are fully considered. Through context analysis and part-of-speech tagging, ambiguity is eliminated to ensure that the word segmentation results accurately reflect the true semantic intent of the user query.

[0027] Keyword statistics and weight calculation: for each keyword after word segmentation, calculate the frequency of occurrence TF of each keyword in each document, and determine the inverse document frequency IDF of the keyword in the entire knowledge base;

[0028] Document relevance score calculation: For each knowledge base document, the relevance score with the user query is calculated by comprehensively considering the keyword frequency, inverse document frequency and weight factors. Among them, Q represents the user query, d represents the document, and q i is the i-th keyword in the query, f(q i ,d) is q i The word frequency in document d, IDF(q i ) is q i The inverse document frequency of k1 and b are adjustment parameters, |d| is the length of document d, and avgdl is the average length of documents in the knowledge base;

[0029] Document sorting and result screening: sort the documents in the knowledge base according to the calculated relevance score of each document, and select a part of the sorted documents as the results of BM25 keyword search according to pre-set screening rules or thresholds.

[0030] According to a second aspect of the present invention, a computer-readable storage medium is provided, on which one or more computer programs are stored. When the one or more computer programs are executed by a computer processor, the above method is implemented.

[0031] According to a third aspect of the present invention, a hybrid enhanced indexing system based on vector retrieval and BM25 algorithm is proposed, comprising:

[0032] The knowledge base construction unit is configured for users to upload documents. The system parses the files to obtain text content, segments the text content, converts it into dense vectors and stores it in the vector library. The large model is used to extract keywords from the segmented text. After processing, the keywords are inserted into the vocabulary, and the text is segmented based on the constructed vocabulary and the word segmenter.

[0033] The search and comparison result generation unit is configured to respond to receiving a user query, vectorize the user query, calculate the similarity with each vector in the vector library to obtain the first batch of preliminary query results; at the same time, segment the user query, use the BM25 algorithm to perform keyword search based on the vocabulary and the word segmenter to obtain the second batch of query results, re-rank the model to the vector matching retrieval and the BM25 keyword retrieval results similarity, and filter the results according to the set threshold; according to the upper limit set by the large model parameter limit and business needs, select the result with the highest similarity from the screening results; splice the document content corresponding to the result with the user query into a prompt word, input the large model, and the large model parses the semantic relationship and logical structure in the prompt word to generate an answer.

[0034] In some specific embodiments, the knowledge base construction unit segments the text content, converts it into dense vectors and stores it in the vector library, specifically including: mapping the text to a low-dimensional vector space through a pre-trained model in the Transformer architecture, and using a large-scale corpus for pre-training when training the vector representation model so that the model learns a common semantic representation pattern; fine-tuning training is performed on the knowledge base text in the corresponding field to adapt to the domain-specific semantic features; in the vectorization process, the text is subjected to lexical, syntactic and semantic analysis to extract key features, and the key features are combined into a final dense vector representation.

[0035] In some specific embodiments, in the process of extracting keywords from segmented text in the knowledge base construction unit, the segmented text is first preprocessed, and the preprocessing includes removing interference information such as stop words and punctuation marks; the large model uses a combination of multiple algorithms and technologies to extract keywords based on its pre-trained language model and domain knowledge, including using statistical methods to calculate the frequency of words and inverse document frequency indicators to determine the importance of words; at the same time, the attention mechanism in the deep learning model is used to focus on the parts of the text that are highly relevant to the topic and extract key words therein; for the extraction of domain nouns and new nouns, the large model refers to the domain glossary, industry knowledge base, and the latest language trends, and identifies domain-specific and novel keywords by comparing and analyzing the words in the text with known domain words and emerging vocabulary patterns.

[0036] In some specific embodiments, in the process of extracting keywords from segmented texts in the knowledge base construction unit, when extracting names, prompt words in the following format are used to guide the large model to operate:

[0037] Role setting: Make it clear that the role of the big model in the task is a natural language processing expert;

[0038] Background: It is necessary to conduct in-depth analysis of the text to identify possible name entities;

[0039] Expert description Profile: Specializes in using algorithms and pattern recognition techniques to identify names in text

[0040] Skills: Capable of text analysis, entity recognition, pattern matching, and algorithm design, and able to extract names from text efficiently and accurately

[0041] Goals: Accurately identify and extract all names from the text, ensuring that no possible name is missed;

[0042] Constraints: impose constraints on the name extraction process. The extracted names must have textual basis and the extraction results must meet the pre-set format specifications.

[0043] OutputFormat: The output format of names is clearly defined, which facilitates the system to accurately identify and process the extracted results.

[0044] The workflow of name extraction includes:

[0045] Read and analyze the given text. The large model reads the text provided by the user and performs a comprehensive and in-depth analysis on it.

[0046] Identify possible name entities in the text. Based on the analysis of the text, use algorithms and pattern recognition technology to accurately locate possible name entities in the text.

[0047] Output the recognized names in the specified format. Output the recognized names in the pre-set format, with each name separated by "##".

[0048] Confirm that all possible names have been extracted. After outputting the names, the large model will check the entire text again to ensure that no possible names are missed;

[0049] Output <|COMPLETE|> The extraction process is complete.

[0050] In some specific embodiments, the search comparison and result generation unit uses the BM25 algorithm to perform keyword search based on the word list and the word segmenter to obtain the second batch of query results, which specifically includes:

[0051] Keyword preprocessing and analysis: After receiving a user query, the user query is first segmented. The word segmenter constructed in the knowledge base construction phase is used to accurately segment the user query into separate vocabulary units based on the vocabulary information in the vocabulary. In the word segmentation process, the language habits and grammatical structure characteristics of the user query are fully considered. Through context analysis and part-of-speech tagging, ambiguity is eliminated to ensure that the word segmentation results accurately reflect the true semantic intent of the user query.

[0052] Keyword statistics and weight calculation: for each keyword after word segmentation, calculate the frequency of occurrence TF of each keyword in each document, and determine the inverse document frequency IDF of the keyword in the entire knowledge base;

[0053] Document relevance score calculation: For each knowledge base document, the relevance score with the user query is calculated by comprehensively considering the keyword frequency, inverse document frequency and weight factors. Among them, Q represents the user query, d represents the document, and q i is the i-th keyword in the query, f(q i ,d) is q i The word frequency in document d, IDF(q i ) is q i The inverse document frequency of k1 and b are adjustment parameters, |d| is the length of document d, and avgdl is the average length of documents in the knowledge base;

[0054] Document sorting and result screening: sort the documents in the knowledge base according to the calculated relevance score of each document, and select a part of the sorted documents as the results of BM25 keyword search according to pre-set screening rules or thresholds.

[0055] The present invention proposes a hybrid enhanced indexing method and system based on vector retrieval and BM25 algorithm. On the solid foundation of native RAG dense vector retrieval, it cleverly integrates the excellent vocabulary extraction ability of large language models, and innovatively introduces the BM25 keyword retrieval mechanism to jointly construct an efficient and comprehensive hybrid retrieval strategy. This application aims to make up for the limitations of a single retrieval method and improve the accuracy and comprehensiveness of information retrieval. On the other hand, after the retrieval process is started, the system first uses the BM25 algorithm for preliminary keyword matching. The core of this process lies in accurate word segmentation. Traditional word segmentation technology is often limited to a fixed vocabulary, and it is difficult to flexibly respond to emerging new words and professional terms, which affects the timeliness and accuracy of the retrieval. This application introduces a large language model to optimize the word segmentation link. With its deep language understanding and context-awareness capabilities, the large model can accurately identify domain-specific vocabulary and emerging nouns in the text, significantly improving the quality and adaptability of word segmentation. By combining the advantages of RAG dense vector retrieval, vocabulary extraction of large language models, and BM25 keyword retrieval technology, supplemented by optimized word segmentation and re-ranking mechanisms, the hybrid retrieval solution of this application not only improves retrieval efficiency and accuracy, but also demonstrates strong flexibility and the ability to capture new information. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and together with the description are used to explain the principles of the present invention. Other embodiments and many expected advantages of the embodiments will be readily appreciated as they become better understood by reference to the following detailed description. Other features, objects and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments made with reference to the following drawings:

[0057] Figure 1 is a flow chart of a hybrid enhanced indexing method based on vector retrieval and BM25 algorithm according to an embodiment of the present application;

[0058] Figure 2 It is a flowchart of a hybrid enhanced indexing method based on vector retrieval and BM25 algorithm in a specific embodiment of the present application;

[0059] Figure 3 This is an architecture diagram of a hybrid enhanced indexing system based on vector retrieval and BM25 algorithm according to an embodiment of the present application;

[0060] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application. DETAILED DESCRIPTION

[0061] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0062] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0063] Figure 1 FIG. 4 is a flow chart of a hybrid enhanced indexing method based on vector retrieval and BM25 algorithm according to an embodiment of the present application. Figure 1 As shown, the method comprises the following steps:

[0064] S101: The user uploads a document, the system parses the file to obtain the text content, segments the text content, converts it into dense vectors and stores it in the vector library, uses a large model to extract keywords from the segmented text, inserts the keywords into the vocabulary after processing, and segments the text based on the constructed vocabulary and word segmenter.

[0065] In a specific embodiment, the text is mapped to a low-dimensional vector space through a pre-trained model in a Transformer architecture. When training the vector representation model, a large-scale corpus is used for pre-training so that the model can learn a general semantic representation pattern. Fine-tuning training is performed on the knowledge base text in the corresponding field to adapt to the domain-specific semantic features. In the vectorization process, the text is subjected to lexical, syntactic and semantic analysis to extract key features, and the key features are combined into a final dense vector representation.

[0066] In a specific embodiment, in the process of extracting keywords from segmented text, the segmented text is first preprocessed, and the preprocessing includes removing interference information such as stop words and punctuation marks; the large model uses a combination of multiple algorithms and technologies to extract keywords based on its pre-trained language model and domain knowledge, including using statistical methods to calculate the frequency of words and inverse document frequency indicators to determine the importance of words; at the same time, the attention mechanism in the deep learning model is used to focus on the parts of the text that are highly relevant to the topic and extract key words therein; for the extraction of domain nouns and new nouns, the large model refers to the domain glossary, industry knowledge base and the latest language trends, and identifies domain-specific and novel keywords by comparing and analyzing the words in the text with known domain words and emerging vocabulary patterns.

[0067] In a specific embodiment, when extracting names, prompt words in the following format are used to guide the large model to operate:

[0068] Role setting: Make it clear that the role of the big model in the task is a natural language processing expert;

[0069] Background: It is necessary to conduct in-depth analysis of the text to identify possible name entities;

[0070] Expert description Profile: Specializes in using algorithms and pattern recognition techniques to identify names in text

[0071] Skills: Capable of text analysis, entity recognition, pattern matching, and algorithm design, and able to extract names from text efficiently and accurately

[0072] Goals: Accurately identify and extract all names from the text, ensuring that no possible name is missed;

[0073] Constraints: impose constraints on the name extraction process. The extracted names must have textual basis and the extraction results must meet the pre-set format specifications.

[0074] OutputFormat: The output format of names is clearly defined, which facilitates the system to accurately identify and process the extraction results.

[0075] According to the hybrid enhanced indexing method based on vector retrieval and BM25 algorithm of claim 4, the workflow of name extraction includes:

[0076] Read and analyze the given text. The large model reads the text provided by the user and performs a comprehensive and in-depth analysis on it.

[0077] Identify possible name entities in the text. Based on the analysis of the text, use algorithms and pattern recognition technology to accurately locate possible name entities in the text.

[0078] Output the recognized names in the specified format. Output the recognized names in the pre-set format, with each name separated by "##".

[0079] Confirm that all possible names have been extracted. After outputting the names, the large model will check the entire text again to ensure that no possible names are missed;

[0080] Output <|COMPLETE|> The extraction process is complete.

[0081] S102: In response to receiving a user query, the user query is vectorized, and the similarity with each vector in the vector library is calculated to obtain the first batch of preliminary query results; at the same time, the user query is segmented, and the keyword search is performed using the BM25 algorithm based on the vocabulary and the word segmenter to obtain the second batch of query results, and the similarity of the results of the vector matching search and the BM25 keyword search are re-ranked by the model, and the results are filtered according to the set threshold; according to the upper limit set by the large model parameter restrictions and business needs, the result with the highest similarity is selected from the filtered results; the document content corresponding to the result and the user query are spliced ​​into prompt words, which are input into the large model, and the large model parses the semantic relationship and logical structure in the prompt words to generate answers.

[0082] In a specific embodiment, the second batch of query results are obtained by performing keyword search using the BM25 algorithm based on the word list and the word segmenter, specifically including:

[0083] Keyword preprocessing and analysis: After receiving a user query, the user query is first segmented. The word segmenter constructed in the knowledge base construction phase is used to accurately segment the user query into separate vocabulary units based on the vocabulary information in the vocabulary. In the word segmentation process, the language habits and grammatical structure characteristics of the user query are fully considered. Through context analysis and part-of-speech tagging, ambiguity is eliminated to ensure that the word segmentation results accurately reflect the true semantic intent of the user query.

[0084] Keyword statistics and weight calculation: for each keyword after word segmentation, calculate the frequency of occurrence TF of each keyword in each document, and determine the inverse document frequency IDF of the keyword in the entire knowledge base;

[0085] Document relevance score calculation: For each knowledge base document, the relevance score with the user query is calculated by comprehensively considering the keyword frequency, inverse document frequency and weight factors. Among them, Q represents the user query, d represents the document, and q i is the i-th keyword in the query, f(q i ,d) is q i The word frequency in document d, IDF(q i ) is q i The inverse document frequency of k1 and b are adjustment parameters, |d| is the length of document d, and avgdl is the average length of documents in the knowledge base;

[0086] Document sorting and result screening: sort the documents in the knowledge base according to the calculated relevance score of each document, and select a part of the sorted documents as the results of BM25 keyword search according to pre-set screening rules or thresholds.

[0087] Figure 2 A flowchart of a hybrid enhanced indexing method based on vector retrieval and BM25 algorithm in a specific embodiment of the present application is shown. Figure 2 As shown in the figure, the specific implementation of this method can be divided into two parts: knowledge base construction and retrieval comparison:

[0088] 1. Knowledge base construction: The specific steps are as follows:

[0089] 1.1 The user uploads a document and obtains the text content of the document after parsing the file.

[0090] 1.2 The text content is segmented, vectorized, formed into dense vectors, and synchronized to the vector library.

[0091] 1.3 Using the natural language processing capabilities of the large model itself, extract keywords from the segmented text with high concurrency, including domain nouns, new nouns, names of people, etc., and insert them into the word list of the segmentation.

[0092] In a specific example, the prompt words for name extraction can be referenced as follows:

[0093] -Role: Natural Language Processing Expert

[0094] -Background: Users need to automatically identify and extract names from text, which usually requires in-depth analysis of the text to identify possible name entities.

[0095] -Profile: You are an expert in natural language processing, using algorithms and pattern recognition techniques to identify names in text.

[0096] -Skills: You have the ability of text analysis, entity recognition, pattern matching and algorithm design to extract names from text efficiently and accurately.

[0097] -Goals: Accurately identify and extract all names from the text, ensuring that no possible name is missed.

[0098] -Constrains: The extracted names must be explicitly mentioned in the text, and the extracted results must conform to the specified format.

[0099] -OutputFormat: Use ## to separate each name, and output <|COMPLETE|> after extraction.

[0100] -Workflow:

[0101] 1. Read and analyze the given text.

[0102] 2. Identify possible name entities in the text.

[0103] 3. Output the recognized name in the specified format.

[0104] 4. Confirm that all possible names have been extracted.

[0105] 5. The output <|COMPLETE|> indicates the end of the extraction process

[0106] 1.4 builds the corresponding vocabulary and word segmenter to segment the text and use Elasticsearch index.

[0107] 2. Retrieve and compare the generated results. The specific steps are as follows:

[0108] 2.1 The user query is vectorized, and vector matching retrieval is performed in the vector library constructed in 1.2 to obtain a batch of query results.

[0109] 2.2 User query is segmented, and a batch of query results are obtained based on the vocabulary and analyzer constructed in 1.4 and BM25 keyword search.

[0110] 2.3 uses the re-ranking model to unify the similarity of the query results of 2.1 and 2.2, and obtains a batch of query results by setting threshold filtering.

[0111] 2.4 Set the query result upper limit n according to the large model parameter restrictions and business restrictions, and select the n query results with the highest similarity.

[0112] 2.5 The filtered documents and user queries are combined into the final prompt words, which are input into the large model to generate the answer.

[0113] The above method of the present application summarizes the enhanced indexing of vector retrieval, key word extraction and hybrid with BM25 algorithm through practice. Users can choose the corresponding steps according to their own business to complete the enhanced indexing, so that the large model can more fully and reliably utilize the knowledge base information. By combining the advantages of RAG dense vector retrieval, vocabulary extraction of large language models and BM25 keyword retrieval technology, and supplemented by optimized word segmentation and reordering mechanisms, the hybrid retrieval scheme of the present application not only improves retrieval efficiency and accuracy, but also demonstrates strong flexibility and keen ability to capture new information. The indexing scheme based on hybrid enhancement of vector retrieval and BM25 algorithm proposed in this application has shown stable and broad application prospects in domestic government agencies, Internet companies and other types of systems. In view of the rapid change of Internet information and the continuous emergence of new hot words, it has become an irreversible industry trend to build an intelligent vocabulary update mechanism and integrate it accurately into the business system.

[0114] Figure 3 FIG. 4 shows an architecture diagram of a hybrid enhanced indexing system based on vector retrieval and BM25 algorithm according to an embodiment of the present application. Figure 3 As shown, the system includes a knowledge base construction unit 301 and a search and comparison result generation unit 302. Among them, the knowledge base construction unit 301 is configured for users to upload documents, the system performs file parsing to obtain text content, segments the text content, converts it into dense vectors and stores it in the vector library, uses the big model to extract keywords of the segmented text, inserts the keywords into the word list after processing, and segments the text based on the constructed word list and word segmenter; the search and comparison result generation unit 302 is configured to respond to receiving user queries, vectorize the user queries, calculate the similarity with each vector in the vector library to obtain the first batch of preliminary query results; at the same time, segment the user queries, perform keyword retrieval using the BM25 algorithm based on the word list and word segmenter to obtain the second batch of query results, re-rank the model to the vector matching retrieval and BM25 keyword retrieval results similarity, and filter the results according to the set threshold; according to the big model parameter restrictions and the upper limit set by business needs, select the result with the highest similarity from the screening results; splice the document content corresponding to the result and the user query into prompt words, input the big model, and the big model parses the semantic relationship and logical structure in the prompt words to generate answers. .

[0115] Reference below Figure 4 , which shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing an embodiment of the present application. Figure 4The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0116] like Figure 4 As shown, the computer system includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage part 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the system 400 are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0117] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed, so that a computer program read therefrom is installed into the storage section 408 as needed.

[0118] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, - but not limited to - a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wireline, optical cable, RF, etc., or any suitable combination of the foregoing.

[0119] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0120] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0121] The modules involved in the embodiments of the present application may be implemented by software or by hardware.

[0122] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: when a user uploads a document, the system performs file parsing to obtain text content, segments the text content, converts it into dense vectors and stores it in a vector library, extracts keywords from the segmented text using a large model, inserts the keywords into a vocabulary after processing, and segments the text based on the constructed vocabulary and word segmenter; in response to receiving a user query, vectorizes the user query, calculates the similarity with each vector in the vector library to obtain the first batch of preliminary query results; at the same time, segment the user query, performs keyword search using the BM25 algorithm based on the vocabulary and word segmenter to obtain the second batch of query results, re-ranks the model to calculate the similarity between the results of vector matching search and BM25 keyword search, and filters the results based on a set threshold; selects the result with the highest similarity from the screening results based on the large model parameter restrictions and the upper limit set by business needs; splices the document content corresponding to the result and the user query into prompt words, inputs the large model, and the large model parses the semantic relationship and logical structure in the prompt words to generate answers.

[0123] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.

Claims

1. A hybrid enhanced indexing method based on vector retrieval and BM25 algorithm, characterized in that: include: S1: The user uploads a document, the system parses the file to obtain the text content, segments the text content, converts it into dense vectors and stores it in the vector library, uses the large model to extract keywords from the segmented text, inserts the keywords into the vocabulary after processing, and segments the text based on the constructed vocabulary and the word segmenter; S2: In response to receiving a user query, the user query is vectorized, and the similarity with each vector in the vector library is calculated to obtain a first batch of preliminary query results; at the same time, the user query is segmented, and a keyword search is performed using the BM25 algorithm based on the vocabulary and the word segmenter to obtain a second batch of query results, and the similarity of the results of the vector matching retrieval and the BM25 keyword retrieval by the re-ranking model is filtered according to a set threshold; according to the upper limit set by the large model parameter limit and business needs, the result with the highest similarity is selected from the filtered results; the document content corresponding to the result and the user query are spliced ​​into a prompt word, which is input into the large model, and the large model parses the semantic relationship and logical structure in the prompt word to generate an answer.

2. The hybrid enhanced indexing method based on vector retrieval and BM25 algorithm according to claim 1 is characterized in that: The text content is segmented in S1, converted into dense vectors and stored in a vector library, specifically including: mapping the text to a low-dimensional vector space through a pre-trained model in the Transformer architecture, and using a large-scale corpus for pre-training when training the vector representation model so that the model can learn a general semantic representation pattern; fine-tuning training is performed on the knowledge base text in the corresponding field to adapt to the domain-specific semantic features; in the vectorization process, the text is subjected to lexical, syntactic and semantic analysis, key features are extracted, and the key features are combined into a final dense vector representation.

3. The hybrid enhanced indexing method based on vector retrieval and BM25 algorithm according to claim 1 is characterized in that: In the process of extracting keywords from segmented text in S1, the segmented text is first preprocessed, and the preprocessing includes removing interference information such as stop words and punctuation marks; the large model extracts keywords based on its pre-trained language model and domain knowledge by combining multiple algorithms and technologies, including using statistical methods to calculate the frequency of words and inverse document frequency indicators to determine the importance of words; at the same time, the attention mechanism in the deep learning model is used to focus on the parts of the text that are highly relevant to the topic and extract key words therein; for the extraction of domain nouns and new nouns, the large model refers to the domain glossary, industry knowledge base and the latest language trends, and identifies keywords with domain specificity and novelty by comparing and analyzing the words in the text with known domain words and emerging vocabulary patterns.

4. The hybrid enhanced indexing method based on vector retrieval and BM25 algorithm according to claim 3 is characterized in that: In the process of extracting keywords from segmented text in S1, when extracting names, the following format of prompt words is used to guide the large model to operate: Role setting: Make it clear that the role played by the large model in the task is a natural language processing expert; Background: It is necessary to conduct in-depth analysis of the text to identify possible name entities; Expert Description Profile: Expertise in using algorithms and pattern recognition techniques to identify names in text; Skills: Have the ability to analyze text, identify entities, match patterns, and design algorithms, and be able to extract names from text efficiently and accurately; Goals: Accurately identify and extract all names from the text, ensuring that no possible name is missed; Constraints: impose constraints on the name extraction process. The extracted names must have textual basis and the extraction results must meet the pre-set format specifications. OutputFormat: The output format of names is clearly defined, which facilitates the system to accurately identify and process the extraction results.

5. The hybrid enhanced indexing method based on vector retrieval and BM25 algorithm according to claim 4 is characterized in that: The workflow of name extraction includes: Read and analyze a given text. The large model reads the text provided by the user and performs a comprehensive and in-depth analysis on it; Identify possible name entities in the text, and based on the analysis of the text, use algorithms and pattern recognition technology to accurately locate possible name entities in the text; Output the recognized names in the specified format, and output the recognized names in the pre-set format, with each name separated by "##"; Confirming that all possible names have been extracted, after outputting the names, the large model rechecks the entire text to ensure that no possible names are missed; Output <|COMPLETE|> The extraction process is complete.

6. The hybrid enhanced indexing method based on vector retrieval and BM25 algorithm according to claim 1 is characterized in that: In S2, the keyword search is performed using the BM25 algorithm based on the word list and the word segmenter to obtain the second batch of query results, which specifically includes: Keyword preprocessing and analysis: After receiving a user query, the user query is first segmented. The word segmenter constructed in the knowledge base construction phase is used to accurately segment the user query into separate vocabulary units based on the vocabulary information in the vocabulary. In the word segmentation process, the language habits and grammatical structure characteristics of the user query are fully considered. Through the technical means of context analysis and part-of-speech tagging, ambiguity is eliminated to ensure that the word segmentation result accurately reflects the true semantic intention of the user query; Keyword statistics and weight calculation: for each keyword after word segmentation, calculate the frequency of occurrence TF of each keyword in each document, and determine the inverse document frequency IDF of the keyword in the entire knowledge base; Document relevance score calculation: For each knowledge base document, the relevance score with the user query is calculated by comprehensively considering the keyword frequency, inverse document frequency and weight factors. Among them, Q represents the user query, d represents the document, and q i is the i-th keyword in the query, f(q i ,d) is q i The word frequency in document d, IDF(q i ) is q i The inverse document frequency of k1 and b are adjustment parameters, |d| is the length of document d, and avgdl is the average length of documents in the knowledge base; Document sorting and result screening: sort the documents in the knowledge base according to the calculated relevance score of each document, and select a part of the sorted documents as the results of BM25 keyword search according to pre-set screening rules or thresholds.

7. A computer-readable storage medium having one or more computer programs stored thereon, characterized in that: When the one or more computer programs are executed by a computer processor, the method according to any one of claims 1 to 6 is implemented.

8. A hybrid enhanced indexing system based on vector retrieval and BM25 algorithm, characterized in that: include: A knowledge base construction unit is configured for users to upload documents, the system performs file parsing to obtain text content, segments the text content, converts it into dense vectors and stores it in a vector library, extracts keywords from the segmented text using a large model, inserts the keywords into a vocabulary after processing, and segments the text based on the constructed vocabulary and a word segmenter; The search and comparison result generation unit is configured to respond to receiving a user query, vectorize the user query, calculate the similarity with each vector in the vector library to obtain the first batch of preliminary query results; at the same time, segment the user query, use the BM25 algorithm to perform keyword search based on the vocabulary and the word segmenter to obtain the second batch of query results, re-rank the model to the vector matching retrieval and the BM25 keyword retrieval results similarity, and filter the results according to the set threshold; according to the upper limit set by the large model parameter limit and business needs, select the result with the highest similarity from the screening results; splice the document content corresponding to the result with the user query into a prompt word, input the large model, and the large model parses the semantic relationship and logical structure in the prompt word to generate an answer.

9. The hybrid enhanced indexing system based on vector retrieval and BM25 algorithm according to claim 8, characterized in that: The text content is segmented in the knowledge base construction unit, converted into dense vectors and stored in the vector library, specifically including: mapping the text to a low-dimensional vector space through a pre-trained model in the Transformer architecture, and using a large-scale corpus for pre-training when training the vector representation model so that the model learns a general semantic representation pattern; fine-tuning training is performed on the knowledge base text in the corresponding field to adapt to the field-specific semantic features; in the vectorization process, the text is subjected to lexical, syntactic and semantic analysis, key features are extracted, and the key features are combined into a final dense vector representation.

10. The hybrid enhanced indexing system based on vector retrieval and BM25 algorithm according to claim 8, characterized in that: In the process of extracting keywords from segmented text in the knowledge base construction unit, the segmented text is first preprocessed, and the preprocessing includes removing interference information such as stop words and punctuation marks; the large model extracts keywords based on its pre-trained language model and domain knowledge by combining multiple algorithms and technologies, including using statistical methods to calculate the frequency of words and inverse document frequency indicators to determine the importance of words; at the same time, the attention mechanism in the deep learning model is used to focus on the parts of the text that are highly relevant to the topic and extract key words therein; for the extraction of domain nouns and new nouns, the large model refers to the domain glossary, industry knowledge base and the latest language trends, and identifies domain-specific and novel keywords by comparing and analyzing the words in the text with known domain words and emerging vocabulary patterns.

11. The hybrid enhanced indexing system based on vector retrieval and BM25 algorithm according to claim 8, characterized in that: In the process of extracting keywords from segmented texts in the knowledge base construction unit, when extracting names, prompt words in the following format are used to guide the large model to operate: Role setting: Make it clear that the role played by the large model in the task is a natural language processing expert; Background: It is necessary to conduct in-depth analysis of the text to identify possible name entities; Expert Description Profile: Expertise in using algorithms and pattern recognition techniques to identify names in text; Skills: Have the ability to analyze text, identify entities, match patterns, and design algorithms, and be able to extract names from text efficiently and accurately; Goals: Accurately identify and extract all names from the text, ensuring that no possible name is missed; Constraints: impose constraints on the name extraction process. The extracted names must have textual basis and the extraction results must meet the pre-set format specifications. OutputFormat: The output format of names is clearly defined, which facilitates the system to accurately identify and process the extraction results. The workflow of name extraction includes: Read and analyze a given text. The large model reads the text provided by the user and performs a comprehensive and in-depth analysis on it; Identify possible name entities in the text, and based on the analysis of the text, use algorithms and pattern recognition technology to accurately locate possible name entities in the text; Output the recognized names in the specified format, and output the recognized names in the pre-set format, with each name separated by "##"; Confirming that all possible names have been extracted, after outputting the names, the large model rechecks the entire text to ensure that no possible names are missed; Output <|COMPLETE|> The extraction process is complete.

12. The hybrid enhanced indexing system based on vector retrieval and BM25 algorithm according to claim 8, characterized in that: The search comparison and result generation unit uses the BM25 algorithm to perform keyword search based on the word list and the word segmenter to obtain the second batch of query results, which specifically includes: Keyword preprocessing and analysis: After receiving a user query, the user query is first segmented. The word segmenter constructed in the knowledge base construction phase is used to accurately segment the user query into separate vocabulary units based on the vocabulary information in the vocabulary. In the word segmentation process, the language habits and grammatical structure characteristics of the user query are fully considered. Through the technical means of context analysis and part-of-speech tagging, ambiguity is eliminated to ensure that the word segmentation result accurately reflects the true semantic intention of the user query; Keyword statistics and weight calculation: for each keyword after word segmentation, calculate the frequency of occurrence TF of each keyword in each document, and determine the inverse document frequency IDF of the keyword in the entire knowledge base; Document relevance score calculation: For each knowledge base document, the relevance score with the user query is calculated by comprehensively considering the keyword frequency, inverse document frequency and weight factors. Among them, Q represents the user query, d represents the document, and q i is the i-th keyword in the query, f(q i ,d) is q i The word frequency in document d, IDF(q i ) is q i The inverse document frequency of k1 and b are adjustment parameters, |d| is the length of document d, and avgdl is the average length of documents in the knowledge base; Document sorting and result screening: sort the documents in the knowledge base according to the calculated relevance score of each document, and select a part of the sorted documents as the results of BM25 keyword search according to pre-set screening rules or thresholds.

Citation Information

Cited By

  • Heterogeneous A / B data source run-through strengthening method, system, equipment and medium

    CN120353835A

  • Literature semantic search method and system based on elastic search

    CN120429311A

  • Online content recommendation method and system based on semantic discovery

    CN120492737A

  • An online content recommendation method and system based on semantic discovery

    CN120492737B

  • Intelligent file centralized management method and device in information system, equipment and medium

    CN120705114A