Literature semantic search method and system based on elastic search
By building a hybrid search architecture, combining Apache Lucene and Milvus databases, semantic vectorization and error correction processing are used using the CoSENT model and the Spacy model, the problem of low efficiency in semantic understanding and high-dimensional vector data retrieval is solved, and efficient and accurate literature retrieval is achieved.
Patent Information
- Application Number
- CN202510946180.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Traditional Elasticsearch has shortcomings in semantic understanding and multilingual processing, and cannot recognize synonyms or synonyms, and cannot capture context information, resulting in inaccurate query matching and low retrieval efficiency of high-dimensional vector data.
A hybrid search architecture is built, combined with Apache Lucene's BM25 algorithm and Milvus vector database, the text2vec-base-multilingual model trained by the CoSENT method is used for semantic vectorization, combined with the Spacy model for error correction and noun phrase extraction, and the approximate nearest neighbor search and BM25 algorithm are used for keyword matching, and the Sequence Matcher algorithm is used for deduplication and weight allocation, realizing deep semantic understanding and efficient retrieval.
It improves the accuracy and efficiency of literature search, can identify synonyms and synonyms, understand context information, and improves user experience and retrieval performance.
Smart Images

Figure CN120429311A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data retrieval, and in particular to a document semantic search method and system based on elastic search. Background Art
[0002] In terms of traditional search engine technology, Apache Lucene, a representative open-source full-text search library, implements efficient text retrieval through technologies such as inverted indexing, word segmentation and text analysis, and relevance scoring. Based on Lucene technology, a variety of distributed search engine solutions have emerged. Among them, Solr, an early distributed search engine, is based on Lucene and supports distributed indexing and querying, achieving horizontal scalability and high availability through sharding and replication. However, its configuration and management are relatively complex, and its real-time performance is poor.
[0003] Elasticsearch (ES), a next-generation distributed search engine, inherits Lucene's core technologies and features significant innovations and improvements in distributed architecture optimization, real-time performance, ease of use, and ecosystem integration. Built on Apache Lucene, ES is a distributed, real-time search and analytics engine primarily used for full-text search, structured data retrieval, log analysis, and data visualization. Its core technologies include inverted indexing, distributed storage, word segmentation, and relevance scoring algorithms (such as TF-IDF and BM25). ES's core features include a distributed architecture, real-time performance, efficient full-text search, and versatility, enabling large-scale data processing and horizontal scalability. Compared to traditional technologies, ES features built-in distributed storage and retrieval capabilities, supports near-real-time indexing and search, provides an easy-to-use RESTful API, and integrates with tools such as Logstash and Kibana to form a complete log analysis and data visualization solution (ELK Stack). These advantages have made ES one of the most popular search engines, suitable for large-scale data retrieval and analysis.
[0004] However, with the rapid development of artificial intelligence and natural language processing technologies, traditional keyword-matching search techniques have gradually exposed their limitations. Modern semantic retrieval technologies, based on natural language processing (NLP) techniques such as word embeddings (Word2Vec, GloVe) and sentence embeddings (BERT, Sentence-BERT), can capture synonyms, near-synonyms, and contextual information, enabling deep semantic understanding.
[0005] Specifically, as a keyword-based search engine, Elasticsearch, while excelling in full-text search and distributed processing, suffers from significant shortcomings in semantic retrieval due to its core design. ES is built on Apache Lucene and uses an inverted index and the BM25 (or TF-IDF) algorithm for keyword retrieval. This mechanism enables it to efficiently process literal matches, but it lacks semantic understanding. For example, ES cannot recognize synonyms or near-synonyms, nor can it capture contextual information, resulting in the query "AI technology" not matching documents containing "artificial intelligence." Furthermore, ES has limited capabilities for multilingual and complex semantic processing. While its word segmenter and text analyzer support multiple languages, they still rely on rules or statistical methods and are unable to process deep semantic information.
[0006] Furthermore, ES's core design lacks an integrated semantic model and cannot directly support vector-based semantic search. The high computational complexity of semantic search conflicts with ES's design goals of high performance and low latency, resulting in limited performance in scenarios requiring deep semantic understanding, such as intelligent question answering and semantic search. This is an inherent flaw of ES as a traditional search engine.
[0007] Therefore, how to combine the efficient keyword retrieval capabilities of traditional Elasticsearch with modern semantic retrieval technology to build a hybrid retrieval system that has both high performance and supports semantic understanding has become an important direction for the current development of search engine technology and a technical problem that needs to be solved urgently. Summary of the Invention
[0008] In response to the lack of semantic understanding in Elasticsearch's keyword-based retrieval, this application provides a document semantic search method and system based on Elasticsearch. By constructing a hybrid retrieval architecture that includes a traditional keyword retrieval module and a semantic vector retrieval module, it combines Apache Lucene's BM25 algorithm with the semantic retrieval capabilities of the Milvus vector database. It can achieve deep semantic understanding and synonym recognition capabilities while maintaining the efficient keyword matching performance of traditional search engines, significantly improving the accuracy of document retrieval and user experience.
[0009] One aspect of the present application provides a document semantic search method based on elastic search, comprising: constructing a hybrid retrieval architecture including a first retrieval module and a second retrieval module, wherein the first retrieval module adopts Elasticsearch based on Apache Lucene, processes structured text data through an inverted index and a BM25 algorithm, and the second retrieval module adopts a Milvus vector database, and processes vector data using a vector index structure; preprocessing the input query text data to generate a noun phrase list; vectorizing the noun phrase list using a text2vec-base-multilingual model trained based on a CoSENT method to obtain a semantic vector; using the semantic vector as a query vector, performing an approximate nearest neighbor search in the second retrieval module, obtaining extended word data, a similarity score, and a retrieval ranking position as first candidate data; inputting the query text data into the first retrieval module, performing a keyword matching retrieval through the BM25 algorithm, calculating a statistical score based on word frequency and inverse document frequency, and a retrieval ranking position as second candidate data; fusing the first candidate data and the second candidate data, and aggregating the fused candidate data according to logical operators in the query text data to obtain third candidate data; and using Sequence The Matcher algorithm calculates the string similarity between the extended words in the third candidate data, sets a similarity threshold based on the longest common subsequence length and performs deduplication processing to obtain the fourth candidate data; the fourth candidate data is weighted based on position and similarity score, and the distinction of the extended words is enhanced by expanding the score range to obtain the final extended word recommendation list data.
[0010] Further, a list of noun phrases is generated, including: counting the number of characters N in the query text data, and when N is greater than a threshold When the query text data is corrected, pycorrector based on the statistical language model is used to perform text correction on the query text data and output the corrected query text data; otherwise, the original query text data is used as the corrected query text data; The value range is 3 to 10.
[0011] Determine whether the length of the query text data after error correction is greater than the threshold , The value range is 2 to 5. If yes, the Transformer-based Spacy en_core_web_trf model is used to perform lexical and syntactic analysis on the corrected query text data to generate document object data containing part-of-speech tags and grammatical structure; if no, the corrected query text data is used as the document object data; based on the document object data, the doc.noun_chunks method built into the Spacy model is used to extract the noun phrase set; the set data structure is used to deduplicate the noun phrase set to obtain the final noun phrase list.
[0012] Among them, pycorrector: a Chinese text automatic error correction tool based on a statistical language model. By analyzing the frequency of occurrence and contextual relationships of words in a large-scale corpus, combined with the edit distance algorithm and n-gram model, it can automatically identify and correct spelling errors, vocabulary errors, and grammatical errors in user input. In this solution, it is used to preprocess query texts with a character count exceeding the threshold N0 to improve the accuracy of the input text.
[0013] The Spacy en_core_web_trf model is a pre-trained model for English natural language processing based on the Transformer architecture. It integrates deep learning technologies such as BERT and has powerful lexical analysis, syntactic analysis, named entity recognition, and semantic understanding capabilities. In this solution, it is used to perform deep language analysis on corrected query text whose length exceeds the threshold C0, generating document object data containing part-of-speech tags, dependencies, and grammatical structure.
[0014] Spacy model: An industrial-grade open source natural language processing framework that provides a complete text processing pipeline, including word segmentation, part-of-speech tagging, named entity recognition, dependency parsing, and other functions. In this solution, it serves as the core language analysis engine, responsible for processing document object data and providing noun phrase extraction capabilities.
[0015] doc.noun_chunks method: The built-in noun phrase extraction method of the Spacy model is based on dependency parsing and grammatical rules. It can automatically identify and extract noun phrase chunks in the text. These noun phrases usually contain core semantic information. In this solution, it is used to extract a set of noun phrases that carry the main semantic content from the document object data.
[0016] The set data structure is a Python data type that is implemented based on a hash table. It features unique elements and fast lookups, and can perform insertion and lookup operations with an average time complexity of O(1). In this solution, it is used to automatically deduplicate the extracted noun phrase set, eliminating duplicate noun phrases and improving subsequent processing efficiency.
[0017] Noun phrase list: The final output after preprocessing, language analysis, extraction, and deduplication. It contains the core semantic elements of the query text. This list-based collection of noun phrases serves as input data for subsequent vectorization processing in this solution, providing refined semantic units for semantic retrieval.
[0018] In particular, spelling errors in user input lead to retrieval failures. In addition, it is difficult to extract core semantic information in long text queries, and repeated noun phrases affect retrieval efficiency. Therefore, this application introduces a multi-level preprocessing mechanism of pycorrector automatic error correction and Spacy deep language analysis.
[0019] Further, obtaining a semantic vector includes: using the CoSENT method and a ranking loss function L of cosine similarity to train a text2vec-base-multilingual model; using the trained text2vec-base-multilingual model to semantically encode the list of noun phrases and convert each noun phrase into a high-dimensional semantic vector;
[0020] Sorting loss function L expression:
[0021] ; where i, j, k, l represent the sample index in the sample pair, represents the sentence embedding vector of the corresponding sample, represents the set of all positive sample pairs, Represents the set of all negative sample pairs; is a boundary parameter used to control the interval between positive and negative sample pairs; .
[0022] The text2vec-base-multilingual model is a Transformer-based, sentence-level semantic encoding model. Built on Transformer architectures like BERT, it employs a multi-head self-attention mechanism and a feedforward neural network to process all word positions in a sequence in parallel, capturing long-range semantic dependencies and generating dense vector representations of fixed dimensions. Pre-trained on a large-scale multilingual corpus, it supports semantic understanding in multiple languages, including Chinese and English, and possesses cross-language semantic alignment capabilities, mapping semantically similar texts in different languages to similar locations in the vector space.
[0023] In particular, traditional word vector models fail to capture sentence-level semantics. This application employs the CoSENT (Cosine Sentence) contrastive learning method. By constructing a large number of positive and negative sample pairs for supervised learning and using cosine similarity as a semantic similarity metric, the text2vec-base-multilingual model is trained to learn sentence-level semantic representations. Unlike traditional word-level vector models like Word2Vec and GloVe, the CoSENT method can holistically understand the semantics of sentences or phrases, capturing the complex semantic relationships and contextual dependencies between words. The resulting high-dimensional semantic vectors contain the complete semantic information of noun phrases, rather than simple word combinations.
[0024] On the other hand, the loss function L in this application maximizes the cosine similarity of positive pairs (semantically similar sentence pairs) while minimizing the cosine similarity of negative pairs (semantically dissimilar sentence pairs). The boundary parameter λ controls the distance between positive and negative samples. When the similarity of the positive pair exceeds the similarity of the negative pair plus the boundary λ, the loss is zero. Otherwise, gradient optimization is performed. This sorting mechanism forces the model to learn more precise semantic boundaries, significantly improving the ability to distinguish between positive and negative samples.
[0025] Furthermore, the text2vec-base-multilingual model, trained on large-scale multilingual corpora, possesses cross-lingual semantic understanding capabilities and can handle mixed Chinese and English query scenarios. Leveraging the Transformer architecture's multi-head self-attention mechanism, the model processes all positions in a sequence in parallel, capturing long-range dependencies and generating semantic vectors with good semantic continuity and comparability.
[0026] Further, obtaining the first candidate data includes: inputting the semantic vector as a query vector into the Milvus vector database;
[0027] An approximate nearest neighbor search based on cosine similarity is performed in the Milvus vector database to calculate the similarity score between the query vector and the word vector stored in the Milvus vector database; the top 1 candidate expansion word data with the highest similarity score is obtained, and the candidate expansion word data includes the expansion word text, the similarity score, and the sorting position index in the search results; the top 1 candidate expansion word data is threshold filtered to determine whether the similarity score of the first-ranked expansion word is lower than a preset threshold. If so, it is determined that there is no valid synonym expansion word and an empty first candidate data is returned; if not, the candidate expansion word data that meets the threshold requirement is retained; the filtered candidate expansion word data, the similarity score, and the sorting position index are combined into a triplet as the first candidate data.
[0028] In particular, traditional databases cannot efficiently handle similarity calculations for high-dimensional vectors, and semantic retrieval response times for tens of millions of documents are excessively long. This application, on the one hand, employs the Milvus vector database and replaces traditional B+ tree indexes with Milvus-specific vector index structures (such as IVF, HNSW, and ANNOY). These vector indexes are optimized for the geometric characteristics of high-dimensional spaces and can reduce linear search complexity from O(n) to O(log n) or even lower. Milvus also employs hardware optimization technologies such as SIMD instruction set acceleration and GPU parallel computing to significantly improve the execution efficiency of vector operations, addressing the performance bottlenecks of traditional relational databases when processing high-dimensional vectors.
[0029] On the other hand, this application uses an approximate nearest neighbor (ANN) search strategy instead of an exact brute force search. Its technical principle is to divide the high-dimensional vector space into multiple subspaces by building a multi-layer graph structure or cluster index. During the search, only part of the candidate set needs to be traversed instead of the entire data, thus significantly improving performance at the cost of a small loss of accuracy. Combined with the efficient calculation formula of cosine similarity ,Through vector dot product and modulus length pre-computation optimization, a millisecond-level similarity calculation response is achieved.
[0030] In addition, the present application performs early filtering of search results by presetting a similarity threshold. When the similarity score of the first-ranked expansion word is lower than the threshold, it is directly determined that there is no valid synonym expansion word and an empty result is returned, avoiding subsequent meaningless calculations and processing.
[0031] Furthermore, obtaining the second candidate data includes: inputting the query text data into the first retrieval module of Elasticsearch of Apache Lucene; mapping each word in the query text data to a list of documents containing the corresponding word using an inverted index structure, and establishing a corresponding relationship between the word and the document; and calculating the similarity score between the query text data Q and the document D based on the BM25 algorithm. ; Get the top two matching document results with the highest BM25 algorithm scores. The matching document results include the document identifier, similarity score, and position index in the search ranking; use the top two matching document results as the second candidate data.
[0032] Apache Lucene, an open-source full-text search engine library, serves as a core foundational technology in the field of information retrieval, providing comprehensive text indexing and search capabilities. It utilizes an inverted index data structure for rapid text retrieval, supports complex query syntax and Boolean logic combinations, offers a variety of text analyzers for word segmentation and text processing, and features efficient disk storage and memory management. In this solution, it serves as the underlying engine for the first retrieval module, responsible for indexing structured text data and performing keyword-matching retrieval.
[0033] Elasticsearch: A distributed, real-time search and analytics engine built on Apache Lucene. Its distributed architecture supports horizontal scalability and high availability, provides a RESTful API for easy system integration, supports real-time indexing and near-real-time search, and offers powerful aggregate analysis and data visualization capabilities, along with built-in load balancing and fault recovery mechanisms. As the first retrieval module in this solution, it encapsulates Lucene's complexity and provides distributed search capabilities. It is responsible for receiving query text data, performing inverted indexing operations, running the BM25 algorithm, and returning ranked matching documents.
[0034] BM25 algorithm: Best Matching 25 algorithm, a document ranking algorithm based on a probabilistic information retrieval model, calculates the relevance score between a document and a query through a nonlinear combination of term frequency (TF) and inverse document frequency (IDF).
[0035] Furthermore, the similarity score : ; Where Q represents the query text data, which contains 1 to n query terms ;D indicates literature; Representing terms Term frequency TF in document D; Indicates the number of words in document D; Represents the average length of all documents in the document collection; and b represent parameters; The value range of is 1.2 to 2; the value range of b is 0.5 to 0.85; Representing terms The inverse document frequency IDF: , where N represents the total number of documents in the document collection; Indicates that it contains a term The number of documents;
[0036] In particular, on the one hand, for the traditional BM25 algorithm parameters The problem that fixed a and b cannot adapt to different document set characteristics is solved by introducing a parameter range constraint mechanism. The parameter is limited to the range of 1.2 to 2, and the b parameter is limited to the range of 0.5 to 0.85, realizing the dynamic adjustment capability of the parameters. The parameter controls the word frequency saturation point, the smaller the A value of 1.2 is suitable for short document collections, larger values A value of 2.0 is suitable for long document collections. The b parameter controls the strength of document length normalization: a smaller b value (0.5) reduces the effect of length, while a larger b value (0.85) strengthens length normalization. By dynamically selecting the optimal parameter combination based on statistical characteristics such as the average length of the document collection and vocabulary distribution, the algorithm can adapt to the characteristics of literature collections of different domains and sizes.
[0037] On the other hand, for the traditional inverse document frequency calculation In extreme cases, numerical anomalies may occur. This application uses Laplace smoothing technology to improve the IDF calculation formula to By adding 1 to both the numerator and the denominator, the smoothing operation effectively avoids This prevents the zero division error that occurs when This smoothing process eliminates the loss of weights caused by IDF values of 0, ensuring that all terms receive reasonable inverse document frequency weights. This smoothing process is particularly suitable for calculating the weights of rare and high-frequency words, improving the numerical stability and robustness of the algorithm.
[0038] Furthermore, obtaining the third candidate data includes: obtaining logical operators in the query text data, the logical operators including a first operator representing and, a second operator representing or, and a third operator representing not; fusing the matching results in the first candidate data and the second candidate data, extracting the expansion words therein, and forming a candidate vocabulary set; identifying repeated expansion words in the candidate vocabulary set, and aggregating the repeated expansion words according to the logical operators: when the logical operator is the first operator and, performing an accumulation operation on the similarity scores of the repeated expansion words to strengthen the intersection result; when the logical operator is the second operator or, performing an average calculation on the similarity scores of the repeated expansion words to balance the union result; when the logical operator is the third operator not, deleting the corresponding expansion words from the candidate vocabulary set; retaining the original similarity scores of the non-repeated expansion words in the first candidate data and the second candidate data, and merging them with the repeated expansion words after aggregation processing; calculating the sorting position index of the merged candidate data, and generating a data set containing expansion words, similarity scores and sorting position indexes as the third candidate data.
[0039] In particular, to address the problem that traditional retrieval systems cannot effectively handle complex logical combination queries, this application decomposes complex Boolean queries into executable logical operation sequences by identifying the three logical operators and, or, and not in the query text.
[0040] To address the issue of a single aggregation strategy for duplicate result scores, this application designs differentiated aggregation algorithms based on the semantic characteristics of different logical operators. When the logical operator is "and," an accumulation strategy is adopted. Based on the "logical AND" requirement for multiple conditions to be met simultaneously, the weight of the intersection result is strengthened through score accumulation, reflecting the confidence of multiple matches. When the logical operator is "or," an averaging strategy is adopted. Based on the "logical OR" inclusiveness, the score is averaged to avoid excessive bias towards a single high-scoring result and ensure the balance of the union result. When the logical operator is "not," a deletion strategy is adopted to directly remove excluded words from the candidate set, achieving precise negative filtering.
[0041] Further, obtaining the fourth candidate data includes: obtaining the extended word text in the third candidate data to form an extended word list to be processed; and using the Sequence Matcher algorithm to calculate the string similarity ratio between any two words in the extended word list: ,in, Indicates the length of the longest common subsequence between string a and string b; and Represent the lengths of string a and string b respectively; determine whether the string similarity ratio is greater than a preset threshold, and if so, perform deduplication processing; combine the expanded words after deduplication processing and the corresponding similarity scores and sorting position indexes to obtain the fourth candidate data.
[0042] In particular, to address the problem of a large number of highly similar, repeated, expanded words in search results, this application employs the Sequence Matcher algorithm to perform precise similarity calculations based on the longest common subsequence (LCS). Using a dynamic programming algorithm, the longest common subsequence length (LCS) between two strings is calculated. Compared to simple edit distance or character overlap calculations, the LCS algorithm can identify consecutive matching segments and common parts within strings that maintain order, more accurately reflecting the structural similarity between words. Furthermore, the LCS algorithm does not require completely continuous character matching; it allows for skips and gaps, enabling it to identify semantically related word pairs that differ slightly in form, such as "machine learning" and "machine-learning," "AI" and "artificial intelligence," and "deep learning" and "deep neural network learning." By calculating the character-level LCS, the algorithm can discover the inherent connections between words and accurately identify similarities even with variations in form, such as hyphens, spaces, and word order adjustments.
[0043] In addition, this application uses the character-level analysis capabilities of the Sequence Matcher algorithm to achieve intelligent filtering of cross-language vocabulary. By analyzing the character composition and structural patterns of the string, the algorithm can identify the characteristic differences between different languages, such as the distribution patterns of Chinese characters, English letters, and digital symbols. When words with large character set differences appear in the expanded word list, their LCS length is usually short and the similarity ratio value is low. Filtering through a preset threshold can automatically exclude irrelevant small language interference words, maintaining the language consistency and quality stability of the results.
[0044] Furthermore, the final expansion word recommendation list data is obtained, including: obtaining the similarity score of the expansion word in the fourth candidate data and sort position index ; Based on similarity score and sort position index , calculate the weight of the expanded word : , where ɑ and β are coefficients, is the similarity score of the i-th extended word, is the sort position index of the i-th expansion word; determines the highest weight Is it less than the preset threshold? If so, clear the fourth candidate data and return an empty expansion word recommendation list; if not, execute the next step; use linear transformation to correct the weight and obtain the corrected weight : ;in, and Respectively represent the maximum and minimum values of the weight of the expanded word sequence after descending; and Respectively represent the upper and lower limits of the target score range; the target score range represents the preset standardized score range; according to the revised weight Sort the expansion word list in descending order to obtain the final expansion word recommendation list.
[0045] In particular, when the highest weight of all expansion words When both are less than the preset threshold, the system determines that the retrieval quality does not meet the standard, directly clears the candidate data and returns an empty recommendation list, avoiding pushing low-relevance expansion words to users and preventing the system from making incorrect recommendations when the semantic matching degree is low.
[0046] In addition, the discrimination of expansion words refers to the degree of differentiation and distinguishability of the similarity scores between different expansion words in the recommendation list. After multi-source retrieval fusion, the weight distribution of candidate expansion words often presents a centralized feature, that is, the score differences of most expansion words are small, resulting in a lack of obvious priority distinction in the sorting results. In the hybrid retrieval fusion process, after the candidate data from the semantic retrieval module and the keyword retrieval module are weighted, the score distribution often becomes too concentrated. For example, the weights of most expansion words are concentrated in the narrow range of 0.6-0.8. Through linear transformation Remapping the original weights to a wider target range effectively widens the score gap between expansion terms of varying quality, creating a clear numerical distinction between high-quality expansion terms and common expansion terms. In literature retrieval scenarios, enhancing expansion term discrimination helps researchers quickly identify the most relevant research topics and keywords, avoiding ineffective screening among a large number of expansion terms with similarity scores.
[0047] Another aspect of the present application is to provide a document semantic search system based on elastic search, comprising: a first retrieval module, which processes structured text data through an inverted index based on Apache Lucene and the BM25 algorithm; a second retrieval module, which uses the Milvus vector database and a vector index structure to process vector data; a preprocessing module, which preprocesses the input query text data, performs error correction processing by counting the number of characters, and uses Spacy to calculate the search results. The en_core_web_trf model performs lexical and syntactic analysis, extracts noun phrases and performs deduplication processing to generate a noun phrase list; the vectorization module uses the text2vec-base-multilingual model trained based on the CoSENT method to vectorize the noun phrase list and convert each noun phrase into a high-dimensional semantic vector; the semantic retrieval module uses the semantic vector as the query vector and performs approximate nearest neighbor search in the second retrieval module. The candidate expansion word data, similarity score and retrieval ranking position are obtained through cosine similarity calculation, and threshold filtering is performed to obtain the first candidate data; the keyword retrieval module inputs the query text data into the first retrieval module, performs keyword matching retrieval through the BM25 algorithm, uses the inverted index structure to establish the correspondence between words and documents, calculates the statistical score based on word frequency and inverse document frequency and the retrieval ranking position as the second candidate data; the data fusion module fuses the first candidate data with the second candidate data, aggregates the fused candidate data according to the logical operators in the query text data, and obtains the third candidate data; the deduplication processing module uses Sequence The Matcher algorithm calculates the string similarity between the extended words in the third candidate data, sets a similarity threshold based on the longest common subsequence length, and performs deduplication processing to obtain the fourth candidate data; the weight allocation module assigns weights to the fourth candidate data based on position and similarity scores, expands the score range through linear transformation to enhance the differentiation of the extended words, and sorts them in descending order to obtain the final extended word recommendation list data.
[0048] Compared with the existing technology, the advantages of this application are:
[0049] Traditional ES cannot understand the deep semantics of query statements based on keyword matching. A single retrieval method cannot simultaneously meet the requirements of precise matching and semantic understanding. The ES inverted index is inefficient when processing high-dimensional vector data. This application separates traditional statistical-based keyword matching from deep learning-based semantic vector retrieval by constructing a dual-module hybrid retrieval architecture. The text2vec-base-multilingual model trained with the CoSENT method is used to convert query text into high-dimensional semantic vectors. The Milvus vector database is used for approximate nearest neighbor search, achieving a technological leap from literal matching to semantic understanding, enabling the system to identify synonyms and antonyms and understand contextual semantic relationships.
[0050] In addition, this application adopts a parallel dual-channel retrieval strategy, simultaneously running Apache Lucene's BM25 algorithm for precise keyword matching and Milvus's vector similarity calculation for semantic retrieval, and then aggregates the two retrieval results according to logical operators through an intelligent fusion algorithm, achieving the complementary advantages of precise matching and semantic understanding, ensuring the high precision of traditional search and achieving a high recall rate of semantic retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The present application will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbers represent the same structures, wherein:
[0052] Figure 1 is an exemplary flow chart of a document semantic search method based on elastic search according to some embodiments of the present application;
[0053] Figure 2 is an exemplary flow chart of ES retrieval according to some embodiments of the present application;
[0054] Figure 3 This is an exemplary flowchart of Milvus semantic retrieval according to some embodiments of the present application. DETAILED DESCRIPTION
[0055] The method and system provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0056] Example 1
[0057] like Figure 1As shown, a hybrid retrieval architecture including a first retrieval module and a second retrieval module is constructed. The first retrieval module adopts Elasticsearch based on Apache Lucene and processes structured text data through inverted index and BM25 algorithm. The second retrieval module adopts Milvus vector database and processes vector data using vector index structure. The input query text data is preprocessed to generate a list of noun phrases. The noun phrase list is vectorized using the text2vec-base-multilingual model trained based on the CoSENT method to obtain a semantic vector. The semantic vector is used as the query vector and an approximate nearest neighbor search is performed in the second retrieval module to obtain the expanded word data, similarity score and retrieval ranking position as the first candidate data. The query text data is input into the first retrieval module and a keyword matching retrieval is performed through the BM25 algorithm to calculate the statistical score based on word frequency and inverse document frequency, as well as the retrieval ranking position as the second candidate data. The first candidate data and the second candidate data are fused, and the fused candidate data are aggregated according to the logical operators in the query text data to obtain the third candidate data. Sequence The Matcher algorithm calculates the string similarity between the extended words in the third candidate data, sets a similarity threshold based on the longest common subsequence length and performs deduplication processing to obtain the fourth candidate data; the fourth candidate data is weighted based on position and similarity score, and the distinction of the extended words is enhanced by expanding the score range to obtain the final extended word recommendation list data.
[0058] like Figure 2 As shown in the figure, this solution, based on Elasticsearch (ES), Milvus builds an efficient and scalable literature retrieval system designed to achieve rapid retrieval, precise matching, and intelligent sorting of large-scale literature data. By combining ES's distributed architecture, full-text search capabilities, and modern natural language processing (NLP) technology, the system not only supports traditional keyword search but also introduces semantic search and personalized recommendation features to meet users' diverse literature search needs.
[0059] ES is built on Apache Lucene and uses an inverted index and the BM25 algorithm for efficient full-text search. The inverted index maps each word in a document to a list of documents containing that word, allowing for rapid location of relevant documents. The BM25 algorithm calculates the relevance score between the query keyword and the document by counting the term frequency (TF), inverse document frequency (IDF), and document length. The formula is: ,Formula description: D: document. Q: query, consisting of 1~n queries. , term The term frequency (TF) in document D. , the length of document D (number of words). , the average length of all documents in the document collection. and b: adjustable parameters, , b=0.75. , term The inverse document frequency (IDF) is calculated as: , N: the total number of documents in the document collection. , containing terms The number of documents.
[0060] like Figure 3 As shown in the figure, for semantic retrieval, we chose Milvus to achieve accurate semantic search. The entire process includes the following core steps: input processing and language error correction, semantic vector search, result merging and filtering, and weight allocation and output.
[0061] Automatic error correction based on word length: Use pycorrector to correct text, and enable the correction function when the number of input characters exceeds 5. It can intelligently identify and correct spelling errors in the text, helping users to find potential problems in a timely manner. Error detection: Identify spelling errors in the text through dictionaries or language models. By comparing with the standard vocabulary, EnSpellCorrector can quickly find words that are not in the dictionary. Processing long texts: Use the Spacy model to conduct in-depth analysis of the input text, extract key noun phrases, and decompose long texts into more understandable semantic core units, thereby providing efficient and accurate input for subsequent vectorization, semantic expansion and retrieval. For texts that have been automatically corrected, if its length is greater than 2, Spacy's en_core_web_trf model is applied for word segmentation, and noun phrases are extracted through doc.noun_chunks. The set structure is then used to remove duplicates, and finally a phrase list is returned for subsequent vectorization and semantic expansion.
[0062] This project uses a text2vec-based multilingual model trained using the CoSENT method. This model can convert tens of millions of article keywords into vectors and store them in a vector database to support efficient semantic retrieval and semantic expansion word recommendations. Cosine similarity is used for vector comparison.
[0063] CoSENT calculates the cosine similarity of the embedding of sentence pairs: ;in is the embedding of two sentences.
[0064] CoSENT uses a ranking loss function based on cosine similarity. The loss function formula is: , , all positive sample pairs, The core of CoSENT is a new loss function that optimizes the cos value.
[0065] Merge: Logical operators are used to process query results. When the first logical operator is "and", the similarity scores of repeated expansions are accumulated to strengthen the intersection result; when the second logical operator is "or", the similarity scores of repeated expansions are averaged to balance the union result; when the third logical operator is "not", the corresponding expansion is deleted from the candidate vocabulary set;
[0066] Filtering: Use Sequence Matcher to calculate string similarity to filter out highly similar words to ensure the uniqueness and accuracy of the results. The similarity formula is: , where LCS is the longest common subsequence; a and b are both strings.
[0067] Combining the traditional Elasticsearch (ES) module with the semantic search module ultimately achieves efficient and accurate document retrieval. Built on Apache Lucene, the traditional ES module utilizes an inverted index and the BM25 algorithm for keyword search, enabling rapid location of documents matching the query keywords. However, its limitation lies in its inability to understand semantics, resulting in insufficient processing of synonyms, near-synonyms, and contextual information.
[0068] To address the shortcomings of existing technologies in handling complex logical relationships, this project developed a semantic expansion word recommendation system based on semantic vector retrieval and weight assignment. This system integrates the Milvus vector database, text2vec tools, and natural language processing (NLP) technology to improve the ability to generate highly relevant expansion words under complex logical relationships. The following are some of the system's core features:
[0069] Merge the results of ES searches and Milvus vector database searches and dynamically adjust the result merging strategy: Based on user-entered logical conditions such as "and," "or," and "not," the system flexibly adjusts the result merging method to improve the relevance of the expanded keywords. This approach overcomes the problem that traditional keyword expansion methods cannot effectively handle complex logical relationships.
[0070] Using the Sequence Matcher algorithm, error correction, and input optimization: By implementing spelling correction, phrase recognition, and language restriction, the system can significantly improve the quality of input content and generate semantically complete and accurate word candidates. This approach particularly enhances the error tolerance of cross-language searches (for example, between Chinese and English).
[0071] Efficient vector retrieval and filtering mechanism: Utilize Milvus's vector similarity search function, combined with specific thresholds for filtering and removing duplicate words, to ensure the relevance and purity of the final results.
[0072] Intelligent weight assignment strategy: obtaining similarity scores of extended words in candidate data and sort position index ; Based on similarity score and sort position index , calculate the weight of the expanded word : ; where ɑ and β are coefficients, is the similarity score of the i-th extended word, is the sort position index of the i-th expansion word; determines the highest weight Is it less than the preset threshold? If so, clear the fourth candidate data and return an empty expansion word recommendation list; if not, execute the next step; use linear transformation to correct the weight and obtain the corrected weight : ;in, and Respectively represent the maximum and minimum values of the weight of the expanded word sequence after descending; and Respectively represent the upper and lower limits of the target score range; the target score range represents the preset standardized score range; according to the revised weight Sort the expansion word list in descending order to obtain the final expansion word recommendation list.
[0073] This solution significantly improves the accuracy of semantic expansion and user experience in various search scenarios, particularly for searching and recommending a wide range of resources, including articles, books, journals, and preprints. This innovation allows users to more easily find the information they need, while also bringing new ideas and technical approaches to the field of information retrieval.
[0074] The system first applies spelling correction technology to correct input errors, then uses the en_core_web_trf model for word segmentation and identifies key concepts by extracting noun phrases from doc.noun_chunks. Next, it uses text2vec to vectorize the text and leverages the Milvus vector database for efficient related word searches, ultimately recommending expansion terms.
[0075] The system specifically supports logical groupings such as "and," "or," and "not." Through merging strategies and similarity deduplication (using the Sequence Matcher method), it effectively improves the relevance of expanded terms. Furthermore, the system implements language restrictions, a weighting mechanism, and duplicate word filtering. These measures collectively enhance retrieval efficiency in long text processing and complex query scenarios, significantly improving the user experience.
[0076] Overall, this intelligent semantic search system not only enhances the quality and diversity of expanded terms, but also ensures the accuracy and efficiency of the retrieval process, making it ideal for searching articles, books, journals, and other types of content. Through this innovative approach, users can more easily find the information they need and enjoy a more seamless search experience.
[0077] Example 2
[0078] A document semantic search system based on elastic search includes: a first retrieval module, which processes structured text data through an inverted index based on Apache Lucene and the BM25 algorithm; a second retrieval module, which uses the Milvus vector database and a vector index structure to process vector data; a preprocessing module, which preprocesses the input query text data, performs error correction processing by counting the number of characters, and uses Spacy The en_core_web_trf model performs lexical and syntactic analysis, extracts noun phrases and performs deduplication processing to generate a noun phrase list; the vectorization module uses the text2vec-base-multilingual model trained based on the CoSENT method to vectorize the noun phrase list and convert each noun phrase into a high-dimensional semantic vector; the semantic retrieval module uses the semantic vector as the query vector and performs approximate nearest neighbor search in the second retrieval module. The candidate expansion word data, similarity score and retrieval ranking position are obtained through cosine similarity calculation, and threshold filtering is performed to obtain the first candidate data; the keyword retrieval module inputs the query text data into the first retrieval module, performs keyword matching retrieval through the BM25 algorithm, uses the inverted index structure to establish the correspondence between words and documents, calculates the statistical score based on word frequency and inverse document frequency and the retrieval ranking position as the second candidate data; the data fusion module fuses the first candidate data with the second candidate data, aggregates the fused candidate data according to the logical operators in the query text data, and obtains the third candidate data; the deduplication processing module uses Sequence The Matcher algorithm calculates the string similarity between the extended words in the third candidate data, sets a similarity threshold based on the longest common subsequence length, and performs deduplication processing to obtain the fourth candidate data; the weight allocation module assigns weights to the fourth candidate data based on position and similarity scores, expands the score range through linear transformation to enhance the differentiation of the extended words, and sorts them in descending order to obtain the final extended word recommendation list data.
[0079] The invention of the present application and its implementation methods are described schematically above. This description is not restrictive. Without departing from the spirit or basic features of the present application, the present application can be implemented in other specific forms. What is shown in the accompanying drawings is only one of the implementation methods of the invention of the present application, and the actual structure is not limited to this. Therefore, if a person of ordinary skill in the art is inspired by it, without departing from the purpose of the invention, a structural method and embodiment similar to the technical solution are designed without creativity, which should all fall within the scope of protection of the present application. In addition, the word "including" does not exclude other elements or steps, and the word "one" before an element does not exclude the inclusion of "multiple" elements. Words such as first and second are used to indicate names and do not indicate any specific order.
Claims
1. A document semantic search method based on elastic search, characterized in that: include: A hybrid retrieval architecture was constructed, consisting of a first retrieval module and a second retrieval module. The first retrieval module used Elasticsearch based on Apache Lucene to process structured text data using an inverted index and the BM25 algorithm. The second retrieval module used the Milvus vector database to process vector data using a vector index structure. Preprocess the input query text data to generate a noun phrase list; Use the text2vec-base-multilingual model trained based on the CoSENT method to vectorize the noun phrase list to obtain semantic vectors; Using the semantic vector as the query vector, an approximate nearest neighbor search is performed in the second retrieval module to obtain the expanded word data, similarity score, and retrieval ranking position as the first candidate data; Input the query text data into the first retrieval module, perform keyword matching retrieval using the BM25 algorithm, calculate the statistical score based on word frequency and inverse document frequency, and the retrieval ranking position as the second candidate data; Fusing the first candidate data and the second candidate data, and performing aggregation processing on the fused candidate data according to the logical operator in the query text data to obtain third candidate data; The Sequence Matcher algorithm is used to calculate the string similarity between the extended words in the third candidate data, and a similarity threshold is set based on the longest common subsequence length and duplicate removal is performed to obtain the fourth candidate data; The fourth candidate data is weighted based on the position and similarity score, and the discrimination of the extended word is enhanced by expanding the score range to obtain the final extended word recommendation list data.
2. The document semantic search method based on elastic search according to claim 1, characterized in that: Generate a list of noun phrases, including: Count the number of characters N in the query text data. When N is greater than the threshold When the query text data is corrected, pycorrector based on the statistical language model is used to perform text correction on the query text data and output the corrected query text data; otherwise, the original query text data is used as the corrected query text data; Determine whether the length of the query text data after error correction is greater than the threshold If yes, the Spacy en_core_web_trf model based on the Transformer architecture is used to perform lexical and syntactic analysis on the corrected query text data to generate document object data containing part-of-speech tags and grammatical structure; if no, the corrected query text data is used as the document object data; According to the document object data, the noun phrase set is extracted through the built-in doc.noun_chunks method of the Spacy model; The set data structure is used to remove duplicates from the noun phrase set to obtain the final noun phrase list.
3. The document semantic search method based on elastic search according to claim 1, characterized in that: Get the semantic vector, including: The text2vec-base-multilingual model is trained using the CoSENT method with the cosine similarity ranking loss function L. Using the trained text2vec-base-multilingual model, we semantically encode the list of noun phrases and convert each noun phrase into a high-dimensional semantic vector. Sorting loss function L expression: ; Among them, i, j, k, l represent the sample index in the sample pair, represents the sentence embedding vector of the corresponding sample, represents the set of all positive sample pairs, Represents the set of all negative sample pairs; is a boundary parameter used to control the interval between positive and negative sample pairs; in, .
4. The document semantic search method based on elastic search according to claim 2, characterized in that: Obtain the first candidate data, including: Input the semantic vector into the Milvus vector database as the query vector; Perform an approximate nearest neighbor search based on cosine similarity in the Milvus vector database and calculate the similarity score between the query vector and the word vectors stored in the Milvus vector database; Obtain the top 1 candidate expansion word data with the highest similarity score, wherein the candidate expansion word data includes the expansion word text, similarity score, and ranking position index in the search results; Perform threshold filtering on the top 1 candidate expansion word data to determine whether the similarity score of the top expansion word is lower than the preset threshold. If so, determine that there is no valid synonym expansion word and return the empty first candidate data; if not, retain the candidate expansion word data that meets the threshold requirement; The filtered candidate expansion word data, similarity score and sorting position index are combined into a triplet as the first candidate data.
5. The document semantic search method based on elastic search according to claim 2, characterized in that: Obtain the second candidate data, including: Input query text data into the first retrieval module based on Apache Lucene; Use the inverted index structure to map each word in the query text data to a list of documents containing the corresponding word, and establish a corresponding relationship between the word and the document; Calculate the similarity score between the query text data Q and the document D based on the BM25 algorithm ; Get the top two matching document results with the highest BM25 algorithm scores. The matching document results include the document identifier, similarity score, and position index in the search ranking. The top 2 matching document results are used as the second candidate data.
6. The document semantic search method based on elastic search according to claim 5, characterized in that: Similarity score : ; Where Q represents the query text data, which contains 1 to n query terms. ;D indicates literature; Representing terms Term frequency TF in document D; Indicates the number of words in document D; Represents the average length of all documents in the document collection; and b represent parameters; The value range of is 1.2 to 2; the value range of b is 0.5 to 0.85; Representing terms The inverse document frequency IDF: , where N represents the total number of documents in the document collection; Indicates that it contains a term The number of documents.
7. The document semantic search method based on elastic search according to claim 5, characterized in that: The third candidate data is obtained, including: Acquire logical operators in the query text data, where the logical operators include a first operator representing and, a second operator representing or, and a third operator representing not; Merging the matching results of the first candidate data and the second candidate data, extracting the expanded words therein, and forming a candidate vocabulary set; Identify repeated expansions in the candidate vocabulary set and aggregate them according to logical operators: When the logical operator is the first operator and, the similarity scores of repeated expanded words are accumulated to strengthen the intersection result; When the logical operator is the second operator or, the similarity scores of repeated expanded words are averaged to balance the union results; When the logical operator is the third operator not, the corresponding expansion word is deleted from the candidate vocabulary set; The original similarity scores of the non-repeated expansion words in the first candidate data and the second candidate data are retained, and the non-repeated expansion words are merged with the repeated expansion words after the aggregation process; The sorting position index of the merged candidate data is calculated to generate a data set including the expanded word, the similarity score and the sorting position index as the third candidate data.
8. The document semantic search method based on elastic search according to claim 2, characterized in that: The fourth candidate data is obtained, including: Obtaining the expanded word text in the third candidate data to form an expanded word list to be processed; The Sequence Matcher algorithm is used to calculate the string similarity ratio between any two words in the expanded word list: ;in, Indicates the length of the longest common subsequence between string a and string b; and Respectively represent the length of string a and string b; Determine whether the string similarity ratio is greater than a preset threshold. If so, perform deduplication processing; The expanded words after deduplication processing, the corresponding similarity scores, and the sorting position indexes are combined to obtain the fourth candidate data.
9. The document semantic search method based on elastic search according to any one of claims 2 to 8, characterized in that: Get the final expanded word recommendation list data, including: Get the similarity score of the expanded word in the fourth candidate data and sort position index ; Based on similarity score and sort position index , calculate the weight of the expanded word : ; where ɑ and β are coefficients, is the similarity score of the i-th extended word, is the sort position index of the i-th expansion word; Determine the highest weight Is it less than a preset threshold? If so, clear the fourth candidate data and return an empty expansion word recommendation list; if not, proceed to the next step; Use linear transformation to correct the weight and get the corrected weight : ;in, and Respectively represent the maximum and minimum values of the weight of the expanded word sequence after descending order; and They represent the upper and lower limits of the target score interval respectively; the target score interval represents the preset standardized score range; According to the revised weight Sort the expansion word list in descending order to obtain the final expansion word recommendation list.
10. A document semantic search system based on elastic search, characterized in that: include: The first retrieval module processes structured text data using an inverted index based on Apache Lucene and the BM25 algorithm; The second retrieval module uses the Milvus vector database and processes vector data using a vector index structure; The preprocessing module preprocesses the input query text data, performs error correction by counting the number of characters, uses the Spacy en_core_web_trf model to perform lexical and syntactic analysis, extracts noun phrases and removes duplicates, and generates a noun phrase list; The vectorization module uses the text2vec-base-multilingual model trained based on the CoSENT method to vectorize the noun phrase list and convert each noun phrase into a high-dimensional semantic vector; The semantic retrieval module uses the semantic vector as the query vector and performs an approximate nearest neighbor search in the second retrieval module. It obtains candidate expansion word data, similarity scores, and retrieval ranking positions through cosine similarity calculation, and performs threshold filtering to obtain the first candidate data. The keyword search module inputs the query text data into the first search module, performs keyword matching search using the BM25 algorithm, establishes the correspondence between words and documents using the inverted index structure, calculates the statistical score based on word frequency and inverse document frequency, and the search ranking position as the second candidate data; A data fusion module, which fuses the first candidate data and the second candidate data, and aggregates the fused candidate data according to the logical operator in the query text data to obtain the third candidate data; A deduplication processing module uses a Sequence Matcher algorithm to calculate the string similarity between the extended words in the third candidate data, sets a similarity threshold based on the longest common subsequence length, and performs deduplication processing to obtain the fourth candidate data; The weight assignment module assigns weights to the fourth candidate data based on position and similarity scores, expands the score interval through linear transformation to enhance the discrimination of the extended word, and sorts the data in descending order to obtain the final extended word recommendation list data.
Citation Information
Patent Citations
Short text query expansion and indexing method based on word vector
CN104765769A
Pseudo-correlation feedback model information retrieval method and system based on semantic similarity
CN109829104A
Pseudo-correlation feedback model information retrieval method and system based on BERT
CN110442777A
ConceptNet-based information retrieval query expansion method
CN114840639A
Multi-feature image retrieval method based on inverted index fusion
CN116737973A
Cited By
Data mining system and method for mail data
CN121070985A
A data mining system and method for mail data
CN121070985B
Multi-prediction channel fusion HSCODE complementation calibration method
CN121278661A
Method for constructing original text reference library based on improved reinforcement learning
CN121326978A
Method, system and equipment for evaluating conformity of text and question and medium
CN121579696A