Hybrid search method and device, equipment and storage medium
By introducing hybrid search methods into search technology, combining keyword extraction and vectorized indexing, the shortcomings of the existing technology in multi-vector query and dynamic environments are solved, and a more accurate, efficient and flexible search effect is achieved.
Patent Information
- Application Number
- CN202411933677.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-02
AI Technical Summary
Existing search technologies are limited by their dependence on static datasets when dealing with complex multivector queries, lack flexibility, and are difficult to achieve the required accuracy and efficiency in a dynamic environment.
The hybrid search method is adopted to perform traditional search by extracting keywords in the query statement, and vectorized indexing is used to combine and sort the results to improve the accuracy and efficiency of the search.
It realizes a more accurate, efficient and flexible search technology. By combining vectorized indexes and keyword searches, the accuracy of search results and the scalability of the system are improved.
Smart Images

Figure CN119917614A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information retrieval technology, and in particular to a hybrid search method, device, equipment and storage medium. Background Art
[0002] Traditional search technologies can be divided into two categories: one is database-based search, whether it is a relational database or a NoSQL database, they generally support the like type search function. This search method directly checks whether a specific field contains the specified keyword. The other is a full-text search engine such as ElasticSearch, which uses an inverted index mechanism to pre-establish a mapping relationship between keywords and specific documents. When performing a search, the corresponding document list is returned by matching keywords. The core of these two search methods is based on keyword matching. Once the keyword is not selected properly, the corresponding product or document cannot be found. A significant drawback of this strategy is that if a document does not contain a specific keyword, even if there are synonyms with similar meanings in the document, they cannot be retrieved, thus limiting the breadth and depth of the search. At the same time, current information retrieval systems, such as Milvus, are committed to providing strong support for the management of large-scale vector data. However, these systems are limited by their reliance on static data sets and lack of necessary flexibility when processing complex multi-vector queries. Traditional algorithms and libraries often rely too much on main memory storage and cannot effectively distribute data among multiple machines, which undoubtedly limits their scalability. This limitation makes it difficult for them to adapt to real-world scenarios where data is constantly changing. As a result, existing solutions have difficulty achieving the required accuracy and efficiency in dynamic environments. Summary of the invention
[0003] The embodiments of the present invention provide a hybrid search method, apparatus, device and storage medium, aiming to realize a more accurate, efficient and flexible search technology.
[0004] In a first aspect, an embodiment of the present invention provides a hybrid search method, the method comprising:
[0005] In response to receiving a query statement, extracting keywords of the query statement;
[0006] Search the target document database according to the keyword to obtain the first search result, and use the vectorized data search algorithm to perform vectorized indexing according to the query statement and the target vector database to obtain the second search result, where the target vector database is a vector database obtained by performing knowledge vectorization based on the target document database;
[0007] The first search result and the second search result are merged and sorted to obtain a search result set.
[0008] In a second aspect, an embodiment of the present invention provides a hybrid search device, the hybrid search device comprising:
[0009] A keyword extraction unit, configured to extract keywords of the query statement in response to receiving the query statement;
[0010] A dual-path recall unit is used to search the target document database according to keywords to obtain a first search result, and to use a vectorized data search algorithm to perform vectorized indexing according to the query statement and the target vector database to obtain a second search result, where the target vector database is a vector database obtained by performing knowledge vectorization based on the target document database;
[0011] The merging and sorting unit is used to merge and sort the first search result and the second search result to obtain a search result set.
[0012] In a third aspect, an embodiment of the present invention further provides a hybrid search device, including a processor and a memory, wherein the memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of any hybrid search method provided by the embodiment of the present invention.
[0013] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to enable the electronic device to execute the steps of any hybrid search method provided by the embodiment of the present invention.
[0014] The beneficial effects of the present invention are:
[0015] The hybrid search method of the present invention extracts the keywords of the query statement when receiving the query statement, and searches the target document database according to the keywords to obtain the first search result, and at the same time uses the vectorized search algorithm to perform vectorized indexing according to the query statement and the target vector database to obtain the second search result, and then merges and sorts the first search result and the second search result to obtain a search result set. Therefore, the present invention utilizes the high efficiency of vectorized indexing and the flexibility of keyword search, and sorts and merges the results of vectorized indexing and the results of keyword search, so that the obtained search result set is more accurate. Therefore, the hybrid search method of the present invention realizes a more accurate, efficient and flexible search technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 is a flow chart of an embodiment of a hybrid search method provided in an embodiment of the present invention;
[0018] Figure 2 is a schematic diagram of a search result interface of an embodiment of a hybrid search method provided in an embodiment of the present invention;
[0019] Figure 3 is a flowchart of a specific embodiment of the hybrid search method provided in an embodiment of the present invention;
[0020] Figure 4 is a schematic diagram of a search result interface of a specific embodiment of the hybrid search method provided in an embodiment of the present invention;
[0021] Figure 5 is a schematic diagram of the structure of a hybrid search device provided in an embodiment of the present invention;
[0022] Figure 6 It is a schematic diagram of the structure of a hybrid search device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention. At the same time, in the description of the embodiments of the present invention, the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.
[0024] Embodiments of the present invention provide a hybrid search method, apparatus, device and storage medium.
[0025] Specifically, this embodiment will be described from the perspective of a hybrid search device, which can be integrated into a hybrid search device. The hybrid search device can be an image sensor chip, an image sensor, a camera, or an electronic device such as a mobile phone or a computer. That is, the hybrid search method of the embodiment of the present invention can be executed by a hybrid search device.
[0026] The following is a detailed description in conjunction with the accompanying drawings. In this embodiment, the execution subject is a hybrid search device as an example. It should be noted that the description order of the following embodiments is not intended to limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in an order different from that shown in the accompanying drawings.
[0027] Traditional search technologies can be divided into two categories: one is database-based search, whether it is a relational database or a NoSQL database, they generally support the like type search function. This search method directly checks whether a specific field contains the specified keyword. The other is a full-text search engine such as ElasticSearch, which uses an inverted index mechanism to pre-establish a mapping relationship between keywords and specific documents. When performing a search, the corresponding document list is returned by matching keywords. The core of these two search methods is based on keyword matching. Once the keyword is not selected properly, the corresponding product or document cannot be found. A significant drawback of this strategy is that if a document does not contain a specific keyword, even if there are synonyms with similar meanings in the document, they cannot be retrieved, thus limiting the breadth and depth of the search. At the same time, current information retrieval systems, such as Milvus, are committed to providing strong support for the management of large-scale vector data. However, these systems are limited by their reliance on static data sets and lack of necessary flexibility when processing complex multi-vector queries. Traditional algorithms and libraries often rely too much on main memory storage and cannot effectively distribute data among multiple machines, which undoubtedly limits their scalability. This limitation makes it difficult for them to adapt to real-world scenarios where data is constantly changing. As a result, existing solutions have difficulty achieving the required accuracy and efficiency in dynamic environments.
[0028] In order to solve the above problems, the present invention discloses a hybrid search method, which integrates advanced language models, hybrid indexing technology and multi-vector query processing mechanism, greatly improving the accuracy and scalability of retrieval; by cleverly combining vector embedding technology with traditional indexing methods, the hybrid search method can efficiently cope with the management of large-scale data sets and become an efficient tool for processing complex search tasks; in addition, the hybrid search method also incorporates a cache mechanism and an optimized search algorithm, effectively shortening the response time and further improving system performance. These innovative features make the hybrid search method of the embodiment of the present invention significantly different from the traditional search method, and provide a comprehensive and efficient solution for the document retrieval field.
[0029] Please refer to Figure 1 The specific process of the hybrid search method can be as follows: Step S101 to Step S103, wherein:
[0030] Step S101: In response to receiving a query statement, extract keywords of the query statement.
[0031] The query statement is a statement input by the user for searching for a certain document or a certain type of document that the user wants. For example, the query statement may be "Please give me a benchmark project case in the industrial industry." When receiving the query statement input by the user, the embodiment of the present invention extracts the keywords of the query statement.
[0032] Specifically, keywords in the query statement may be extracted by methods based on statistics, rules, machine learning, deep learning, and vocabulary resources.
[0033] The statistical method for extracting keywords from the query may include:
[0034] 1) Split the query into individual words. Understandably, for languages like Chinese that do not have clear space delimiters, a specialized word segmentation tool is required.
[0035] 2) Calculate the frequency of each word in the query statement or the entire corpus.
[0036] 3) For each word, calculate its term frequency (TF) in the query and its inverse document frequency (IDF) in the corpus. Words with high TF-IDF are often considered keywords.
[0037] 4) Based on TF-IDF or other statistical indicators, select the words with the highest scores as keywords.
[0038] The rule-based method for extracting keywords from the query statement may include:
[0039] 1) Establish a set of rules in advance. For example, nouns, verbs, adjectives and other parts of speech often contain a high amount of information, or have specific phrase structures.
[0040] 2) Use part-of-speech tagging tools to tag the query statements.
[0041] 3) According to the pre-established rules, extract the words that meet the rules from the annotated sentences.
[0042] 4) Filter out the most important words as keywords according to the rules.
[0043] The machine learning-based method for extracting keywords from the query may include:
[0044] 1) Collect a large amount of query statements and their corresponding keyword annotation data.
[0045] 2) Extract features from the query, such as word frequency, part of speech, location information, etc.
[0046] 3) Train a keyword extraction model using supervised learning algorithms such as support vector machines (SVM), decision trees, random forests, or neural networks.
[0047] 4) Apply the trained model to new query statements to extract keywords.
[0048] The deep learning-based method for extracting keywords from the query may include:
[0049] 1) Collect a large amount of query statements and their corresponding keyword annotation data.
[0050] 2) Design a deep learning model, such as a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory network (LSTM), or a Transformer.
[0051] 3) Use a large amount of query data to train the model so that it can learn the distribution characteristics of keywords in the query.
[0052] 4) Input the newly received query statement into the trained model, and the model outputs keywords.
[0053] Optionally, in some embodiments, step S101 may include:
[0054] Based on intent recognition, keywords are extracted from the query statement to obtain keywords.
[0055] Specifically, use machine learning or deep learning technology to train an intent classifier based on a labeled intent dataset; extract features from the query statement after word segmentation, such as the bag-of-words model, TF-IDF, word embedding, etc.; input the features of the query statement into the intent classifier to identify the user's query intent; then, based on the identified query intent, select an appropriate keyword extraction strategy, such as TF-IDF, TextRank, LDA, etc., and calculate the weight for each word in the query statement, and then select the most relevant words as keywords based on the weight.
[0056] Furthermore, in some embodiments, the extracted keywords may be post-processed, such as synonym replacement, context adjustment, and result verification.
[0057] Synonym replacement: replace synonyms in extracted keywords with standard forms to reduce ambiguity.
[0058] Contextual adjustment: Adjust the keyword list based on the query intent and context to ensure that the keywords accurately reflect the query intent.
[0059] Result verification: Verify the accuracy of keyword extraction by interacting with users or using additional validation datasets.
[0060] Finally, the post-processed keywords are integrated and the final keywords are output for subsequent processing.
[0061] Step S102, searching the target document database according to the keyword to obtain a first search result, and using a vectorized data search algorithm to perform vectorized indexing according to the query statement and the target vector database to obtain a second search result.
[0062] The target vector database is a vector database obtained by performing knowledge vectorization based on the target document database.
[0063] The vectorized data search algorithm can be HNSW (Hierarchical Navigable Small World) or Faiss (Facebook AI Similarity Search).
[0064] It is understandable that both HNSW and Faiss are algorithms and data structures for efficient similarity search, and they perform well in dealing with large-scale vector search problems.
[0065] Among them, HNSW is an algorithm for approximate nearest neighbor search (ANN). It is a hierarchical graph structure that combines the characteristics of small-world networks and NNG (Navigable Small World) to accelerate the similarity search between vectors. HNSW organizes data points by building a multi-layer graph structure. Each layer is an NNG, and the layers are connected by "edges", so that the search can be carried out between different layers, thereby speeding up the search. In HNSW, the search starts from the top layer and gradually goes deeper to the lower layers until the nearest neighbor is found. This hierarchical search mechanism can significantly reduce the length of the search path. HNSW is suitable for large-scale data sets because it can process tens of thousands of vectors without significantly reducing the search efficiency. In addition, HNSW also supports multiple distance metrics, which is suitable for different application scenarios.
[0066] In the actual application of HNSW, by adjusting the two key parameters M and EF_CONSTRUCTION, a fine balance can be achieved between performance and accuracy. Generally speaking, the larger the M and the larger the EF_CONSTRUCTION, the longer the index construction time, the higher the accuracy, and the longer the search latency.
[0067] Faiss is an open source library developed by Facebook AI Research for efficient similarity search and dense vector clustering. It provides a variety of algorithms and data structures and supports CPU and GPU acceleration. The embodiment of the present invention provides a variety of commonly used index types and allows performance to be optimized by fine-tuning parameters. At the same time, the embodiment of the present invention can adaptively adjust some parameters according to the change in data volume, and only retain those parameters that are closely related to the static characteristics of the data to ensure the stability and efficiency of the index.
[0068] Specifically, the embodiment of the present invention uses a two-way recall method to retrieve the query statement input by the user, combining vectorized indexing and keyword search. The vectorized indexing process includes vectorizing and similarity calculation of the query statement and the target document database, and the keyword search process includes full-text retrieval based on the keywords extracted in step S101, or fuzzy matching using the like operator to find documents containing specific keywords. In the embodiment of the present invention, the two searches of vectorized indexing and keyword search are performed in parallel.
[0069] Further, in some embodiments, step S102 searches the target document database according to the keyword to obtain the first search result, which may include:
[0070] Based on the keywords, use MySQL query statements to search the target document database and obtain the first search result.
[0071] It should be noted that MySQL is a widely used open source relational database management system that uses SQL (Structured Query Language) for data query and management.
[0072] Furthermore, in some embodiments, before step S102 uses a vectorized data search algorithm to perform vectorized indexing according to the query statement and the target vector database to obtain the second search result, the hybrid search method may further include:
[0073] Perform conditional filtering on the target vector database.
[0074] It is understandable that in the vectorized indexing of the hybrid search method of the embodiment of the present invention, conditional filtering of the target vector database also plays a vital role. Conditional filtering before vectorized indexing is an important preprocessing step, the purpose of which is to reduce the search space and improve search efficiency and accuracy. That is, the embodiment of the present invention excludes documents that do not meet the query conditions by conditionally filtering the target vector database, which not only improves the efficiency of the query, but also ensures the accuracy of the final result.
[0075] For example, by setting a non-empty string filter condition, abnormal documents that may be matched can be effectively eliminated before vectorized indexing, thereby purifying the search results.
[0076] Furthermore, in some embodiments, conditionally filtering the target vector database may include:
[0077] 1) Predetermine the search target and define a series of filter conditions based on the target. It is understandable that these filter conditions can be based on the metadata of the document (such as date, author, category, etc.), or based on keywords or specific attributes of the document content (such as the document's language, size, tags, etc.).
[0078] 2) Extracting relevant metadata from the target vector database to be searched. It is understandable that these metadata are usually stored in the header information of the database or document, which is convenient for quick access and screening.
[0079] 3) Filter the target vector database according to the defined filtering conditions. Specifically, the target vector database can be filtered based on the filtering conditions by using simple matching, range query, logical combination, etc. Among them, simple matching directly excludes documents that do not meet simple keyword matching; range query excludes documents that are not within a specific range; logical combination combines multiple filtering conditions and uses AND / OR and other logics for further precise screening.
[0080] Furthermore, in some embodiments, performing knowledge vectorization based on the target document database to obtain a target vector database may include:
[0081] 1) Obtain document knowledge in the target document database;
[0082] 2) Loading structured data and unstructured data in the document knowledge respectively to obtain a first text block and a second text block;
[0083] 3) Vectorizing the first text block and the second text block, and obtaining a target vector database based on the vectorization results.
[0084] S103: Merge and sort the first search result and the second search result to obtain a search result set.
[0085] It can be understood that the embodiment of the present invention combines the high efficiency of vectorized indexing and the flexibility of keyword searching by merging and sorting the first search result and the second search result, so that the search result set has the advantages of being more accurate, efficient, and flexible than the results obtained by existing search methods.
[0086] Specifically, the embodiment of the present invention can merge and sort the first search result and the second search result by merging and re-ranking, priority-based merging, score-weighted merging, etc. to obtain a search result set.
[0087] Among them, the post-merge re-sorting merges the document lists of the first search result and the second search result into a large list, and removes duplicate documents during the merging process. At this time, if the scoring criteria of the two search results are inconsistent, it is necessary to unify the score calculation method. It is understandable that a comprehensive score can be recalculated based on the scores of the documents in the two results. According to the recalculated comprehensive score, a sorting algorithm (such as quick sort, merge sort, etc.) is used to sort the merged document list to obtain a search result set. Furthermore, the sorted list can be trimmed according to the number of results that need to be returned, and only the top N results are retained to obtain a search result set.
[0088] Priority merging first determines the priority of the first search result and the second search result. For example, the document of the first search result may be considered more relevant. In order of priority, the documents of the two search results are merged (the document in the result with a higher priority is ranked first), and duplicate documents are removed during the merging process to obtain a search result set. Furthermore, a quantity threshold can be set. If the merged results exceed the quantity threshold, the results are trimmed according to the priority to obtain a search result set.
[0089] Score weighted merging first assigns weights to the scores of the first and second search results. The weights can be set based on the reliability or relevance of the search results. Then, a weighted score is calculated for each document based on the weights, and the weighted scores are used to sort the first and second search results when merging them (duplicate documents are also removed during the merging process). Furthermore, the sorted list can also be pruned to retain the top N results.
[0090] It is understandable that the hybrid search method of the embodiment of the present invention may also use other methods to merge and sort the first search result and the second search result, and is not limited to the above three methods of re-sorting, priority merging, and score weighted merging.
[0091] Optionally, in some embodiments, merging and sorting the first search result and the second search result to obtain a search result set may include:
[0092] 1) Calculate the relevance scores of the first search result and the second search result;
[0093] 2) confirming that the relevance score is greater than a preset threshold, merging and sorting the first search result and the second search result by weighted average and sorting fusion to obtain a search result set;
[0094] 3) Confirm that the relevance score is less than a preset threshold, and use a linear weighted sum or Rank-based Result Fusion (RRF) to merge and sort the first search result and the second search result to obtain a search result set.
[0095] Specifically, confirming that the correlation score is greater than a preset threshold value indicates that the first search result and the second search result do not conflict. At this time, the embodiment of the present invention uses weighted average and sorting fusion to merge and sort the first search result and the second search result to obtain a search result set.
[0096] More specifically, the formula for weighted average is:
[0097]
[0098] Where x is the value of each node and w is the weight corresponding to x.
[0099] The formula for sorting fusion is:
[0100]
[0101] Where D is the document set, R is a set of rankings as permutations of 1…|D|, and k usually defaults to 60.
[0102] Confirming that the correlation score is less than the preset threshold indicates that the first search result and the second search result conflict. At this time, the embodiment of the present invention uses a linear weighted sum or a fusion sorting based on the inverse of the results to merge and sort the first search result and the second search result to obtain a search result set.
[0103] It is understandable that due to the significant differences in the relevance scores obtained by weighted scoring for different search methods, a single weight setting is difficult to meet the needs of diverse scenarios, especially when the weight of full-text retrieval is increased. Therefore, at this time, a method RRF can be used to achieve effective fusion and sorting of results without relying on relevance scoring. After applying the RRF strategy, the presentation of search results no longer depends on the relevance score, but is comprehensively integrated based on the ranking of documents in multi-way recall, thereby improving the accuracy of the results and the rationality of the sorting. Alternatively, the first search result and the second search result can be merged and sorted by a linear weighted sum to obtain a search result set. Specifically, the formula for the linear weighted sum is:
[0104] Y=w 1 X 1 +w2 X 2 +…w n X m +ε
[0105] Where Y is the overall evaluation score, X_1…X_n are the indicators, w_1…w_n are the corresponding weights, and ε is the error term.
[0106] Furthermore, in some embodiments, the hybrid search method may further include:
[0107] A document list is output according to a search result set and displayed on a preset search result interface. The search result interface has a first control and a second control. The first control is used to select a corresponding document in the document list after receiving a first operation, and the second control is used to download the document selected by the first operation after receiving the second operation.
[0108] The embodiment of the present invention does not limit the specific shapes and forms of the first control and the second control.
[0109] Optionally, the first control is a button, and the first control in the search results interface can select the corresponding document in the document list after receiving the first operation (such as a single-click operation on the first control); or, the first control is a slider, and the first control in the search results interface can select the corresponding document in the document list after receiving the first operation (such as moving or dragging the first control); or, the first control is a text box, and the user can select the corresponding document in the document list by entering content in the search results interface. In addition, the embodiment of the present invention does not limit the display position of the first control, and the first control can be displayed at any position in the search results interface. In the search results interface, a guide mark is used to remind the user of the document corresponding to the first control, that is, the document that will be selected after the first operation is performed on the first control. To make the display effect more intuitive, the first control can be set to the right of the corresponding document, such as Figure 2 shown.
[0110] Optionally, the second control is a button, and the second control in the search results interface can download the document selected by the first operation after receiving a second operation (such as a single-click operation on the second control); or, the second control is a slider, and the second control in the search results interface can download the document selected by the first operation after receiving a second operation (such as a move or drag operation on the second control); or, the second control is a text box, and the user can download the document selected by the first operation by entering content in the search results interface. In addition, the embodiment of the present invention does not limit the display position of the first control, and the first control can be displayed at any position on the search results interface. To make the display effect more intuitive, text such as "Download" can be set in the second control. Figure 2 shown.
[0111] Figure 3 The specific steps of the hybrid search method of an embodiment of the present invention are shown. In a specific embodiment of the present invention (5G capability cube intelligent solution search), it includes:
[0112] The system is an intelligent assistant tailored for users in the 5G customized network industry. It has a built-in rich 5G customized network knowledge base, and provides convenient browsing, downloading and other operation functions. Behind the system, a complete 5G customized network document library and vector database have been built. Users only need to enter the requirements of the required documents in the system, such as "industrial industry private network solutions", and the system's large model will be able to recognize user intent and extract core keywords: "industrial industry" and "private network solutions". Subsequently, the system builds a keyword index through SQL statements, and filters the search conditions while querying the document library. Since no additional filtering conditions are required here, the system directly uses the HNSW and Faiss algorithms to establish a vectorized index for the vector database for retrieval. Next, the retrieved document set is fused and sorted through weighted averaging and RRF technology, and finally three selected documents are presented to the user, such as Figure 4 As shown in Figure 2. The system uses this process to efficiently provide users with a list of document search results, from which users can select and download the documents they think are most suitable. This approach not only optimizes the user experience, but also provides a solid practical foundation for the practicality and effectiveness of the system.
[0113] To summarize, the embodiment of the present invention extracts the keywords of the query statement when receiving the query statement, and searches the target document database according to the keywords to obtain the first search result, and at the same time uses the vectorized search algorithm to perform vectorized indexing according to the query statement and the target vector database to obtain the second search result, and then merges and sorts the first search result and the second search result to obtain a search result set. Therefore, the present invention utilizes the high efficiency of the vectorized index and the flexibility of the keyword search, and sorts and merges the results of the vectorized index and the results of the keyword search, so that the obtained search result set is more accurate. Therefore, the hybrid search method of the present invention realizes a more accurate, efficient and flexible search technology.
[0114] Specifically, the embodiment of the present invention aims to propose a knowledge base document search method based on a hybrid search combining metadata and vector data. The core technical point is to establish a hybrid search method that combines the advantages of multiple indexing technologies, including FAISS for distributed indexing and HNSW for hierarchical search optimization. This method can seamlessly manage large-scale data sets across multiple machines. At the same time, the functions of filtering index conditions and merging and sorting results make it different from traditional systems, providing a comprehensive solution for document retrieval, and providing a new method for knowledge base document search, making query results more accurate and providing users with better experience and services.
[0115] This embodiment also provides a hybrid search device, which can be integrated into a hybrid search device. For example, the pixel array of an image sensor chip includes a photosensitive area and an optical dark area, such as Figure 5 As shown, the hybrid search device may include:
[0116] The keyword extraction unit 501 is configured to extract keywords of a query statement in response to receiving the query statement.
[0117] The dual-path recall unit 502 is used to search the target document database according to keywords to obtain a first search result, and use a vectorized data search algorithm to perform vectorized indexing based on the query statement and the target vector database to obtain a second search result. The target vector database is a vector database obtained by performing knowledge vectorization based on the target document database.
[0118] The merging and sorting unit 503 is used to merge and sort the first search result and the second search result to obtain a search result set.
[0119] like Figure 6 As shown, Figure 6 A schematic diagram of the structure of a hybrid search device provided for an embodiment of the present invention. The hybrid search device 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. The processor 1101 is electrically connected to the memory 1102. Those skilled in the art will appreciate that the hybrid search device structure shown in the figure does not constitute a limitation on the hybrid search device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0120] The processor 1101 is the control center of the hybrid search device 1100. It uses various interfaces and lines to connect various parts of the entire hybrid search device 1100, and executes various functions and processes data of the hybrid search device 1100 by running or loading software programs and / or units stored in the memory 1102, and calling data stored in the memory 1102, so as to monitor the hybrid search device 1100 as a whole. The processor 1101 can be a processor CPU, a graphics processor GPU, a network processor (Network Processor, NP), etc., and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention.
[0121] In an embodiment of the present invention, the processor 1101 in the hybrid search device 1100 will load the instructions corresponding to the processes of one or more applications into the memory 1102 according to the following steps, and the processor 1101 will run the applications stored in the memory 1102 to implement various functions. Please refer to the previous embodiments and will not repeat them here.
[0122] Optional, such as Figure 6 As shown, the hybrid search device 1100 further includes: a touch screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. The processor 1101 is electrically connected to the touch screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107, respectively. Those skilled in the art can understand that Figure 6 The hybrid search device structure shown in the figure does not constitute a limitation on the hybrid search device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0123] The touch display screen 1103 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user and various graphical user interfaces of the hybrid search device, and these graphical user interfaces can be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD, Liquid Crystal Display), an organic light-emitting diode (OLED, Organic Light-EmittingDiode) and the like. The touch panel can be used to collect the user's touch operation on or near it (such as the user uses any suitable object or attachment such as a finger, a stylus, etc. on the touch panel or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 1101, and can receive the command sent by the processor 1101 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 1101 to determine the type of touch event, and then the processor 1101 provides corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present invention, the touch panel and the display panel can be integrated into the touch display screen 1103 to realize the input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 1103 can also be used as a part of the input unit 1106 to realize the input function.
[0124] The radio frequency circuit 1104 may be used to send and receive radio frequency signals, so as to establish wireless communication with a network device or other hybrid search device through wireless communication, and to send and receive signals between the network device or other hybrid search device.
[0125] The audio circuit 1105 can be used to provide an audio interface between the user and the hybrid search device through a speaker and a microphone. The audio circuit 1105 can transmit the electrical signal converted from the received audio data to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and converted into audio data, and then the audio data is output to the processor 1101 for processing, and then sent to another hybrid search device through the radio frequency circuit 1104, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earplug jack to provide communication between an external headset and the hybrid search device.
[0126] The input unit 1106 may be used to receive input numbers, character information or user feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0127] The power supply 1107 is used to supply power to various components of the hybrid search device 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management system, so as to manage charging, discharging, and power consumption through the power management system. The power supply 1107 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0128] although Figure 6 Not shown in the figure, the hybrid search device 1100 may also include a camera, a sensor, a wireless fidelity unit, a Bluetooth unit, etc., which will not be described in detail here.
[0129] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0130] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0131] To this end, an embodiment of the present invention provides a computer-readable storage medium, in which a plurality of computer programs are stored, and the computer program can be loaded by a processor to execute any hybrid search method provided in an embodiment of the present invention. The computer program can execute the steps of the aforementioned hybrid search method, which can be referred to in the previous embodiment and will not be repeated here.
[0132] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0133] Since the computer program stored in the computer-readable storage medium can execute any hybrid search method provided in the embodiments of the present invention, the beneficial effects that can be achieved by any hybrid search method provided in the embodiments of the present invention can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0134] In the above hybrid search device, computer-readable storage medium, hybrid search equipment, and computer program product embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process and beneficial effects of the hybrid search device, computer-readable storage medium, computer program product, hybrid search device, and its corresponding units described above can refer to the description of the hybrid search method in the above embodiment, and will not be repeated here.
[0135] The above is a detailed introduction to a hybrid search method, a hybrid search device, a hybrid search equipment, a computer-readable storage medium and a computer program product provided in an embodiment of the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A hybrid search method, characterized in that: The method comprises: In response to receiving a query statement, extracting keywords of the query statement; Searching a target document database according to the keyword to obtain a first search result, and using a vectorized data search algorithm to perform vectorized indexing according to the query statement and a target vector database to obtain a second search result, wherein the target vector database is a vector database obtained by performing knowledge vectorization based on the target document database; The first search result and the second search result are combined and sorted to obtain a search result set.
2. The hybrid search method according to claim 1, characterized in that: Before the vectorized data search algorithm is used to perform vectorized indexing according to the query statement and the target vector database to obtain the second search result, the method further includes: Conditional filtering is performed on the target vector database.
3. The hybrid search method according to claim 1, characterized in that: The step of searching the target document database according to the keyword to obtain the first search result includes: According to the keyword, the target document database is searched using a MySQL query statement to obtain the first search result.
4. The hybrid search method according to claim 1, characterized in that: The extracting keywords of the query statement includes: Keyword extraction is performed on the query statement based on intent recognition to obtain the keywords.
5. The hybrid search method according to claim 1, characterized in that: The merging and sorting the first search result and the second search result to obtain a search result set includes: Calculating relevance scores of the first search result and the second search result; Confirming that the relevance score is greater than a preset threshold, merging and sorting the first search result and the second search result by weighted average and sorting fusion to obtain the search result set; It is confirmed that the relevance score is less than the preset threshold, and the first search result and the second search result are merged and sorted by using a linear weighted sum or a fusion sorting based on the inverse of the results to obtain the search result set.
6. The hybrid search method according to claim 1, characterized in that: The performing knowledge vectorization based on the target document database to obtain the target vector database comprises: Acquiring document knowledge in the target document database; Loading structured data and unstructured data in the document knowledge respectively to obtain a first text block and a second text block; The first text block and the second text block are vectorized, and the target vector database is obtained based on the vectorization result.
7. The hybrid search method according to any one of claims 1 to 6, characterized in that: The method further comprises: A document list is output according to the search result set and the document list is displayed on a preset search result interface, wherein the search result interface has a first control and a second control, wherein the first control is used to select a corresponding document in the document list after receiving a first operation, and the second control is used to download the document selected by the first operation after receiving a second operation.
8. A hybrid search device, characterized in that: The hybrid search device comprises: A keyword extraction unit, configured to extract keywords of the query statement in response to receiving the query statement; A dual-path recall unit, configured to search a target document database according to the keyword to obtain a first search result, and to use a vectorized data search algorithm to perform vectorized indexing according to the query statement and a target vector database to obtain a second search result, wherein the target vector database is a vector database obtained by performing knowledge vectorization based on the target document database; The merging and sorting unit is used to merge and sort the first search result and the second search result to obtain a search result set.
9. A hybrid search device, characterized in that: It comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the hybrid search method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises a computer program. When the computer program is run on an electronic device, the computer program is used to enable the electronic device to execute the steps of the hybrid search method according to any one of claims 1 to 7.
Citation Information
Cited By
Data retrieval method, device and system
CN120316119A
Mixed information retrieval generation method, system, equipment and medium
CN120763229A
Data retrieval method, system and equipment based on hybrid storage architecture and medium
CN121434272A
Generation method and device of target reply information, computer equipment, readable storage medium and program product
CN121579763A
Text retrieval method and device, equipment, storage medium and program product
CN121658637A