Document retrieval method and device, electronic equipment, storage medium and product
By combining multi-path retrieval with a large language model, the problem of inaccurate retrieval results in existing technologies is solved, efficient and convenient natural language retrieval in complex service systems is achieved, and the accuracy and comprehensiveness of retrieval results are improved.
Patent Information
- Application Number
- CN202410384095.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-09-30
AI Technical Summary
In the existing technology, the results of searches based on keywords extracted from the search text are not accurate enough, cannot adapt to the diversity and professionalism of complex service systems, and cannot effectively process natural language input.
It uses multiple search paths to process the search text entered by the user through preset search terms, uses a large language model to identify natural language intent, and combines document retrieval, vector retrieval and internal interface retrieval to achieve multi-path parallel retrieval and result fusion.
It improves the accuracy and comprehensiveness of search results, provides an efficient and convenient human-computer interaction method, and adapts to the diversity and professional needs of complex service systems.
Smart Images

Figure CN120723883A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a document retrieval method, device, electronic device, storage medium and product. Background Art
[0002] With the rapid development of information technology, service systems have evolved into highly complex and feature-rich network structures, consisting of multi-layered data architectures, numerous subsystems, and cross-platform integrated modules. The sheer volume of documents in service systems requires a specific method for retrieval.
[0003] In the related art, the search text entered by the user account is obtained, keywords are extracted from the search text, and then based on the extracted keywords, the search engine's inverted index is queried to see whether the keyword is included. If the keyword is included, the document corresponding to the keyword is returned to the user account.
[0004] However, the related technology only performs retrieval based on keywords extracted from the retrieval text, and the retrieval results are not accurate enough. Summary of the Invention
[0005] The present invention provides a document retrieval method, device, electronic device, storage medium, and product that can use multiple search paths to perform searches, resulting in more accurate search results. The technical solution is as follows:
[0006] In a first aspect, a document retrieval method is provided, the method comprising:
[0007] In response to a document search request including target search text sent by a user account, processing the target search text to obtain a search term in a preset format, wherein the preset format is a data format recognizable by a preset search interface, and the preset search interface is an interface capable of performing searches using multiple search paths;
[0008] Based on the search term and the preset search interface, multiple search paths are used to perform searches, and document search results corresponding to each search path are obtained;
[0009] The document retrieval results corresponding to the multiple retrieval paths are sent to the user account.
[0010] In a second aspect, a document retrieval device is provided, the device comprising:
[0011] a processing module configured to respond to a document search request including target search text sent by a user account, process the target search text, and obtain a search term in a preset format, wherein the preset format is a data format recognizable by a preset search interface, and the preset search interface is an interface capable of performing searches using multiple search paths;
[0012] A retrieval module, configured to perform retrieval using multiple retrieval paths based on the retrieval term and the preset retrieval interface, and obtain document retrieval results corresponding to each retrieval path;
[0013] The sending module is used to send the document retrieval results corresponding to the multiple search paths to the user account.
[0014] In a third aspect, an electronic device is provided, comprising a processor and a memory; the memory stores at least one program code; the at least one program code is used to be called and executed by the processor to implement the document retrieval method as described in the first aspect.
[0015] In a fourth aspect, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, and when the at least one computer program is executed by a processor, the document retrieval method as described in the first aspect can be implemented.
[0016] In a fifth aspect, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, it can implement the document retrieval method as described in the first aspect.
[0017] The beneficial effects of the technical solution provided by the embodiments of the present application are:
[0018] The present application provides a preset search interface capable of performing searches using multiple search paths. Upon receiving a document search request from a user account containing target search text, the target search text is processed to obtain a search term in a preset format. This preset format is a data format recognizable by the preset search interface. Therefore, the preset search interface can perform searches based on the search term using multiple search paths, thereby obtaining document search results corresponding to each search path. Compared to searching based on keywords extracted from the search text, this interface provides a richer range of search methods and more accurate search results. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 is a schematic diagram of an implementation environment involved in a document retrieval method provided in an embodiment of the present application;
[0021] Figure 2This is a framework diagram of a document retrieval system provided in an embodiment of the present application;
[0022] Figure 3 This is a schematic diagram of a method for constructing an index in a document retrieval system provided in an embodiment of the present application;
[0023] Figure 4 This is a flow chart of a document retrieval method provided in an embodiment of the present application;
[0024] Figure 5 This is a schematic diagram of a document retrieval process provided by an embodiment of the present application;
[0025] Figure 6 This is a schematic diagram of the structure of a document retrieval device provided in an embodiment of the present application;
[0026] Figure 7 A structural block diagram of an electronic device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0027] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0028] It should be understood that the terms "each," "plurality," and "any" used in the embodiments of this application include two or more, "each" refers to each of the corresponding plurality, and "any" refers to any one of the corresponding plurality. For example, if a plurality of words includes 10 words, "each" refers to each of the 10 words, and "any" refers to any one of the 10 words.
[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0030] Before executing the embodiments of the present application, the terms involved in the embodiments of the present application are first explained.
[0031] Large Language Models (LLMs) are a class of complex algorithmic models built using deep learning techniques that can process, understand, and generate natural language text. They are typically based on the Transformer architecture, a self-attention mechanism that captures long-range dependencies in input text sequences.
[0032] A retrieval system is an information system designed to help users find and retrieve specific information from large, complex data sets. The primary purpose of a retrieval system is to provide an efficient method for users to locate desired data or documents based on various search criteria. Retrieval systems are widely used in areas such as internet search engines, library catalogs, corporate databases, electronic archives, and professional knowledge bases.
[0033] In computer science and natural language processing, vectorization refers to the process of converting data structures into vector form. This process is particularly important in machine learning and data analysis, as most algorithmic models are based on mathematical operations, which are often performed in vector space. Simply put, vectorization is the process of representing data using numerical vectors.
[0034] A search engine is an online information retrieval system designed to help users find content on the internet by entering a query keyword or phrase. Search engines crawl, index, and rank web pages to provide fast, relevant search results. These results are typically displayed as a list of links, each pointing to a web page with a title and a brief description or summary.
[0035] An inverted index is a database indexing technique widely used in full-text search engines. In an inverted index, the index is built around content keywords, allowing for quick location of documents containing specific keywords. This approach contrasts with traditional forward indexing, which organizes data by document, recording the attributes and content of each document.
[0036] Natural Language Processing (NLP) is an interdisciplinary field in computer science, artificial intelligence, and linguistics. Its goal is to develop algorithms and technologies that enable computers to understand, interpret, and generate content in human language. Natural language processing involves the analysis and processing of everything from basic language units (such as words, phrases, and sentences) to complex text and voice conversations.
[0037] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0038] With the research and advancement of artificial intelligence technology, it has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robotics, smart healthcare, and smart customer service. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The solutions provided in the embodiments of this application involve artificial intelligence technologies such as natural language processing, which are specifically explained through the subsequent embodiments. Natural language processing is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language, that is, the language people use in daily life, and is closely related to the study of linguistics. Natural language processing technologies generally include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.
[0039] With the rapid development of information technology, service systems have evolved into highly complex and feature-rich network structures. These systems often consist of multi-layered data architectures, numerous subsystems, and cross-platform integrated modules to meet the diverse needs of service systems for data, entities, and functions. Taking the log system as an example, the log system contains a huge amount of data, entities, and functions, requiring a considerable amount of time and effort to use. Furthermore, it requires certain expertise and experience, and accessing each function requires a specific operation path, which significantly reduces the user experience. At the same time, with the emergence of large language models, human-computer interaction methods are also changing rapidly. Therefore, finding a more convenient and efficient text retrieval method that can support both humans and large language models has become a major concern for those skilled in the art.
[0040] The embodiment of the present application provides a document retrieval method, in which the retrieval text input by the method can be specific data within the service system, or can be category data within the service system, or can be a semantically complete natural language or semantic fragment. Regardless of which of the above the retrieval text is input, the corresponding information can be intelligently retrieved. These information sources are related to multiple aspects of the service system, and have the characteristics of diversity, complexity, and expert experience. Compared with the retrieval method based on keywords extracted from the retrieval text, the information retrieved is richer and more comprehensive. Moreover, the embodiment of the present application does not need to spend a lot of time and cost to train a large language model. It only fine-tunes the customized output of the large language model to the specified interface, and then guides to different retrieval paths through the interface, thereby not only achieving an efficient, low-cost, and easy-to-understand retrieval method, but also providing a new human-computer interaction paradigm.
[0041] Please refer to Figure 1 , which shows the implementation environment involved in the document retrieval method provided in the embodiment of the present application, see Figure 1 The implementation environment includes: a terminal 101 and a server 102. The terminal 101 and the server 102 can communicate directly or indirectly through a network 103, and the network 103 can be a wired network or a wireless network.
[0042] The terminal 101 may be, but is not limited to, a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc. The terminal 101 may have an application installed thereon for searching for documents. Based on the installed application, the terminal 101 may send a document search request carrying target search text to the server 102, so that the server 102 may search for documents that meet the user's requirements and then return the searched documents to the user.
[0043] The server 102 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms. The server 102 may receive a document search request carrying a target search text from the terminal 101, process the target search text to obtain a search term in a preset format, and then, based on the search term and a preset search interface, use multiple search paths to perform a search, obtain document search results corresponding to each search path, and then send the document search results corresponding to the multiple search paths to the user account.
[0044] Figure 2 The framework diagram of the information fusion retrieval system provided by the embodiment of the present application is shown in FIG. Figure 2,The information fusion retrieval system consists of two main layers, namely the unified input layer and the ,multi-path retrieval fusion layer.
[0045] Among them, there are three main types of search texts input into the unified input layer: the first type is specific data that can be identified within the service system, such as entity data; the second type is category data that can be identified within the service system; and the third type is natural language. Among them, specific data is mainly used to retrieve document databases or documents with specific sources. When the input search text is specific data, the system will automatically input it into the information fusion search interface (i.e., the preset search interface described in the embodiment of the present application) to facilitate rapid retrieval of the required documents. Category data is mainly used to search for specific functional web pages in the system. When the input search text is category data, the system will perform further judgment and extraction to determine the type of the input search text, and then convert it into a search term in a preset format, and then input it into the information fusion search interface for retrieval. When the input search text is a natural language that the service system cannot recognize, the system will call the large language model, and then determine the user's intentions and needs through interaction with the user, and then generate a search term in a preset format based on the user's intentions and needs, and then input the search term into the information fusion retrieval interface for retrieval.
[0046] Among them, the multi-path retrieval fusion layer includes a multiple search path module and a result unified sorting module. The multiple search path module is used to provide multiple search paths. The system can quickly obtain search results by using multiple search paths to perform parallel searches on the search terms input into the information fusion retrieval interface. The multiple search paths include document retrieval paths, vector retrieval paths, and internal interface retrieval paths. Among them, the document retrieval path directly inputs the search terms, the vector retrieval path inputs the feature vector generated based on the search terms, and the internal interface retrieval path inputs the interface data generated based on the search terms. For the document retrieval path, the system can input the search terms directly into the document retrieval system. The document retrieval system queries the document database to see whether there are document keywords matching the search terms. If there are document keywords matching the search terms stored in the document database, the document corresponding to the matching document keywords will be returned. The document retrieval path makes full use of the system's retrieval capabilities and can quickly and accurately complete information search and matching. For the vector search path, the system generates a corresponding feature vector based on the search term. This feature vector is then input into the vector search system. The vector search system queries the vector database for feature vectors with a similarity greater than a preset threshold. If a feature vector with a similarity greater than a preset threshold is stored in the vector database, the document corresponding to the feature vector with similarity greater than the preset threshold is returned. This vector search path leverages the efficiency of vector computation, enabling the system to achieve higher accuracy and speed in information retrieval. For the internal interface search path, the system converts the search term into the data format of the internal data query interface and then calls the internal system interface to retrieve documents from the database corresponding to the corresponding service interface. Depending on the service scenario, multiple links calling the internal system interface may exist simultaneously. For documents returned by different interfaces, the overlap ratio must be evaluated and queried. This internal interface search path leverages the efficiency of the internal system interface, providing the system with better support for handling large-scale data and complex service scenarios. The unified result ranking module is used to uniformly rank and merge the results retrieved from multiple search paths, providing the user with the ranked and merged results. This process can be achieved through algorithmic models to ensure the comprehensiveness, accuracy and real-time nature of the search results. Through this series of operations, the information fusion retrieval system can provide users with efficient, accurate and convenient information query services.
[0047] Figure 3This paper introduces the method for constructing the retrieval system index involved in the information fusion retrieval system. The index comprehensively utilizes data from multiple sources, including documents based on entity generalization and abstraction (such as service pages, etc.), service static documents, and documents extracted from service data. These index data are of vital importance to the information fusion retrieval system and can effectively improve the efficiency and accuracy of the information fusion retrieval system. In addition, the information fusion retrieval system may also involve some unstructured data. In order to further optimize the capabilities of the retrieval system, multimodal vector embedding tools can be used to convert these unstructured data into vector form and store them in a vector database, thereby supporting more diverse retrieval methods and providing users with a better retrieval experience.
[0048] The embodiment of the present application combines deep learning and natural language processing technology to provide a new solution for complex human-computer interaction systems. The method introduces a large-scale language model to identify the intent of the input natural language and convert it into input data in a preset format, thereby achieving effective natural language retrieval. In addition, the embodiment of the present application also uses text vectorization tools to model and extract information in the service system to achieve an indexable structure. In terms of retrieval, the document retrieval system and inverted index technology can be used to quickly retrieve the abstracted information and interact with the functional module interface within the service system, thereby achieving unified sorting of multiple data. This method not only improves the freedom of input and human-computer interaction, but also improves the data management efficiency and query performance of the service system.
[0049] This application embodiment provides a document retrieval method, see Figure 4 , the method process provided in the embodiment of the present application includes:
[0050] 401. In response to a document search request including a target search text sent by a user account, the target search text is processed to obtain a search term in a preset format.
[0051] The target search text can be specific data that can be identified within the service system, or it can be category data that can be identified within the service system, or it can be natural language that cannot be identified by the service system. The preset format is a data format that can be identified by the preset search interface, and the preset format can include specified parameters, such as qurey, type, etc. The preset search interface is an interface that can use multiple search paths for retrieval, and the preset search interface can be Figure 2 The information fusion retrieval interface in , etc. The search term can represent the search intention of the user account.
[0052] In an embodiment of the present application, in response to a document search request containing a target search text sent by a user account, the target search text is processed to obtain a search term in a preset format. The following method can be used:
[0053] 4011. In response to a document retrieval request, identify text content attributes of a target retrieval text.
[0054] The text content attributes of the target search text include specific data identifiable within the service system, category data identifiable within the service system, and natural language. In response to a document search request, a content recognition model can be invoked to recognize the target search text, thereby identifying the text content attributes of the target search text.
[0055] 4012. According to the text content attributes of the target search text, different processing methods are adopted to process the target search text to obtain the search terms.
[0056] The text content attributes of the target search text include but are not limited to the following:
[0057] In the first case, when the text content attribute of the target search text is specific data or category data within the system, the target search text can be subjected to word segmentation and string matching processing to obtain the search terms.
[0058] In the second case, when the text content attribute of the target search text is natural language, the large language model can be called to recognize the target search text to obtain the search terms.
[0059] The large language model is called to recognize the target search text to obtain the search term, which may specifically include the following steps:
[0060] The first step is to call the large language model to identify the target retrieval text and obtain multiple candidate intent information.
[0061] For example, the content of the target retrieval text is "I want to query the host". By calling the large language model to identify the target retrieval text, multiple candidate intent information can be obtained, such as "Do you want to query the information of the host?", "Do you want to query the parameters of the motherboard?", etc., and then the identified multiple candidate intent information can be provided to the user account.
[0062] The second step is to obtain the candidate intent information selected by the user account in response to the user account's selection operation on any candidate intent information.
[0063] After receiving multiple candidate intent information identified by the large language model, the terminal displays the multiple candidate intent information. The user account can select any one candidate intent information from the multiple candidate intent information according to its actual query intention. When the user account's selection operation on any candidate intent information is detected, the terminal obtains the candidate intent information selected by the user account, and then sends the candidate intent information to the server. When the server receives the candidate intent information selected by the user account sent by the terminal, it obtains the candidate intent information.
[0064] The third step is to process the candidate intent information to obtain the search terms.
[0065] The server can invoke the large language model to process the candidate intent information selected by the user account and obtain the search term corresponding to the candidate intent information. For example, if the candidate intent information selected by the user account is "Do you want to query host information?", the server can invoke the large language model to process this candidate intent information and obtain the search term "host information."
[0066] The embodiments of the present application provide a new human-computer interaction paradigm. With the help of a large language model, by interacting with the user, the user's intentions and needs can be identified, thereby achieving effective retrieval of natural language and improving the freedom of input and human-computer interaction.
[0067] By processing the target search text in the document search request, a search term in a unified format can be obtained. The search term can be read by a preset search interface, so that a search can be performed using multiple search paths based on the search term.
[0068] 402. Based on the search terms and the preset search interface, multiple search paths are used to perform search, and document search results corresponding to each search path are obtained.
[0069] The search path is the link used in the search process. Different links correspond to different search methods, which can retrieve different information and enrich the search results. Specifically, based on the search terms and the preset search interface, multiple search paths are used to search and obtain the document search results corresponding to each search path. The following method can be used:
[0070] 4021. Based on a preset search interface, the search terms are processed to obtain search path feature data corresponding to different search paths.
[0071] Since the embodiment of the present application provides multiple search paths, the search path feature data used for searching using different search paths is different. The search path feature data is the data required for searching using the corresponding search path. Therefore, for multiple search paths, the server processes the search terms based on the preset search interface, including but not limited to the following situations:
[0072] In the first case, for the document retrieval path, the search term can be directly used as the retrieval path feature data corresponding to the document retrieval path based on a preset retrieval interface.
[0073] In the second case, for vector search paths, a target feature vector corresponding to the search term can be obtained based on a preset search interface, and the target feature vector can be used as the search path feature data corresponding to the vector search path. When obtaining the target feature vector corresponding to the search term, a model for extracting feature vectors from the search term can be pre-trained. By invoking this model to process the search term, the target feature vector corresponding to the search term can be obtained.
[0074] In the third scenario, for internal interface search paths, the search terms can be converted into interface data recognizable within the system based on a preset search interface, and the interface data can then be used as search path feature data corresponding to the internal interface search path. When converting the search terms into interface data recognizable within the system, a correspondence between the search terms and the interface data can be pre-set, and based on the set correspondence, the search terms can be converted into interface data recognizable within the system.
[0075] 4022. Perform a search based on the search path feature data corresponding to different search paths to obtain a document search result corresponding to each search path.
[0076] For the above three search paths, in this step, the server performs searches based on the search path feature data corresponding to different search paths, including but not limited to the following three situations:
[0077] In the first case, the search path is the document search path
[0078] In this case, based on the search term, a document database can be searched for document keywords that match the search term. The document database is used to store the correspondence between document keywords and documents using an inverted document index. If a document keyword matching the search term is found in the document database, the document corresponding to the matching document keyword is used as the document search result corresponding to the document search path.
[0079] The document retrieval path can be used to retrieve documents that include document keywords, meeting the user's demand for accurate document retrieval.
[0080] In the second case, the search path is a vector search path
[0081] For the second scenario, based on the target feature vector, a vector database can be searched for feature vectors whose similarity to the target feature vector exceeds a preset threshold. This vector database is used to store correspondences between feature vectors and documents. If the vector database contains a feature vector whose similarity to the target feature vector exceeds the preset threshold, the document corresponding to the searched feature vector is used as the document retrieval result for the vector search path.
[0082] Using the vector retrieval path can retrieve documents with high similarity to the target feature vector, meeting the user's needs for fast document query.
[0083] The third case, the search path is the internal interface search path
[0084] In this case, based on the interface data, you can call the corresponding service interface within the system to perform a search, obtaining document retrieval results corresponding to the internal interface search path. Typically, the number of service interfaces that can be called by this interface data can be one or more, depending on the actual search requirements. Different service interfaces can correspond to different service scenarios, allowing you to query data generated by different service scenarios within the service system.
[0085] Using the internal interface search path for searching can retrieve data from different service scenarios, making the search results more comprehensive and accurate.
[0086] In this embodiment, the service system is characterized by diversity, complexity, and expert experience. This modeling process transforms the service system into an indexable structure, enabling rapid retrieval of this abstracted information using text vectorization tools and a search system using an inverted index. Simultaneously, the extracted information is queried by contacting the functional module interfaces within the service system.
[0087] 403. Send document search results corresponding to the multiple search paths to the user account.
[0088] Document retrieval results from multiple search paths can be sent directly to the user account. Alternatively, the results from multiple search paths can be merged to generate a fused result, which can then be sent to the user account. When merging document retrieval results from multiple search paths, duplicates must be removed and sorted before being sent to the user account based on the sorted results. This unified sorting of results from multiple search paths achieves superior retrieval results.
[0089] Existing retrieval systems primarily rely on keywords extracted from search text and are unable to summarize and extract the search text. Furthermore, existing retrieval systems only utilize inverted document technology, lacking the ability to search data from other paths within the system, resulting in incomplete searches. Furthermore, existing retrieval systems are unable to adapt to the highly flexible nature of natural language input, resulting in a relatively limited retrieval process. Furthermore, existing retrieval systems are built around search engines, lacking a deep understanding of the data within service systems and unable to adapt to the professionalism and experience inherent in these systems.
[0090] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.
[0091] Figure 5 The overall flow chart of the document retrieval method provided by the embodiment of the present application is shown in FIG. Figure 5 , which may include the following steps:
[0092] Step 110: The user account inputs a search text, and sends a document search request based on the input search text.
[0093] Step 120: In response to the document retrieval request, the system determines whether the retrieved text is in natural language that cannot be extracted. If the retrieved text is identified as specific data or categorized data from an internal system, steps 140-141 are executed to perform feature extraction and data matching. If the text is identified as natural language, steps 130-131 are executed to invoke the large language model for processing.
[0094] Step 130-Step 131: Call the large language model to identify the user's intention information, and convert the user's intention information into system-recognizable feature data pairs (i.e., search terms), which can then be input into the information fusion retrieval interface.
[0095] Step 140-Step 141: The system uses word segmentation, string matching and other methods to extract type or service data to obtain feature data pairs (ie, search terms), which can then be input into the information fusion retrieval interface.
[0096] Step 150: The unified fusion query interface is responsible for processing the input feature data pairs, thereby performing multiple parallel query tasks, wherein steps 160 to 162, steps 170 to 172, and steps 180 to 182 are parallel processes.
[0097] Step 160: The feature data is directly input into the document retrieval system for query.
[0098] Step 161: Query the search engine's inverted index to see whether there are document keywords that match the feature data pair.
[0099] Step 162: Return the search results and the score of the query relevance.
[0100] Step 170: Convert the input feature data pairs into features to obtain vectorized representations of the feature data pairs.
[0101] Step 171: Input the feature vector obtained in step 170 into the vector retrieval system, and obtain the matched feature vector with the highest similarity by matching in the vector retrieval system.
[0102] Step 172: Return the search result corresponding to the feature vector with the highest similarity and the query similarity score.
[0103] Step 180: Convert the input feature data into interface data of the internal data query interface, and call the internal service interface for retrieval.
[0104] Step 181: Integrate the data returned by the service interface.
[0105] Step 182: Return the integrated search results and keyword overlap.
[0106] Step 190: Comprehensively evaluate and sort the search results returned by the above parallel tasks according to their respective scores, and finally return the integrated results.
[0107] Please refer to Figure 6 , which shows a schematic diagram of the structure of a document retrieval device provided in an embodiment of the present application. The device can be implemented by software, hardware, or a combination of both, and becomes all or part of an electronic device. The device includes:
[0108] Processing module 601, configured to respond to a document search request including target search text sent by a user account, process the target search text, and obtain a search term in a preset format, where the preset format is a data format recognizable by a preset search interface, which is an interface capable of performing searches using multiple search paths;
[0109] A search module 602 is configured to perform searches using multiple search paths based on search terms and a preset search interface, and obtain document search results corresponding to each search path;
[0110] The sending module 603 is used to send the document search results corresponding to the multiple search paths to the user account.
[0111] In another embodiment of the present application, the processing module 601 is used to identify the text content attributes of the target retrieval text in response to a document retrieval request; and according to the text content attributes of the target retrieval text, different processing methods are used to process the target retrieval text to obtain retrieval terms.
[0112] In another embodiment of the present application, the processing module 601 is used to perform word segmentation and string matching on the target search text to obtain search terms when the text content attribute of the target search text is specific data or category data within the system.
[0113] In another embodiment of the present application, the processing module 601 is used to call the large language model to recognize the target search text and obtain the search terms when the text content attribute of the target search text is natural language.
[0114] In another embodiment of the present application, the processing module 601 is used to call the large language model to identify the target search text and obtain multiple candidate intent information; in response to the user account's selection operation for any candidate intent information, obtain the candidate intent information selected by the user account; and process the candidate intent information to obtain the search term.
[0115] In another embodiment of the present application, the retrieval module 602 is used to process the search terms based on a preset search interface to obtain search path feature data corresponding to different search paths, and the search path feature data is the data required when searching using the corresponding search path; based on the search path feature data corresponding to different search paths, a search is performed to obtain document retrieval results corresponding to each search path.
[0116] In another embodiment of the present application, the retrieval module 602 is used to use the search term directly as the retrieval path feature data corresponding to the document retrieval path based on a preset retrieval interface; based on the search term, query the document keywords matching the search term from the document database, and the document database is used to store the correspondence between the document keywords and the documents in an inverted document manner; when the document database stores document keywords matching the search term, the document corresponding to the matching document keyword is used as the document retrieval result corresponding to the document retrieval path.
[0117] In another embodiment of the present application, the retrieval module 602 is used to obtain a target feature vector corresponding to a search term based on a preset retrieval interface, and use the target feature vector as retrieval path feature data corresponding to the vector retrieval path; based on the target feature vector, query a feature vector whose similarity with the target feature vector is higher than a preset threshold from a vector database, and the vector database is used to store the correspondence between the feature vector and the document; when a feature vector whose similarity with the target feature vector is higher than a preset threshold is stored in the vector database, the document corresponding to the queried feature vector is used as the document retrieval result corresponding to the vector retrieval path.
[0118] In another embodiment of the present application, the search module 602 is configured to convert the search term into interface data recognizable within the system based on a preset search interface, and use the interface data as search path feature data corresponding to the internal interface search path;
[0119] Based on the interface data, the corresponding service interface within the system is called for retrieval to obtain the document retrieval results corresponding to the internal interface retrieval path.
[0120] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0121] Figure 7 FIG. 7 is a block diagram of an electronic device 700 according to an exemplary embodiment of the present application. Generally, the electronic device 700 includes a processor 701 and a memory 702 .
[0122] The processor 701 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 701 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state; the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 701 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 may also include an artificial intelligence processor, which is used to process computing operations related to machine learning.
[0123] The memory 702 may include one or more computer-readable storage media, which may be non-transitory computer-readable storage media, such as CD-ROMs (Compact Disc Read-Only Memory), ROMs, RAMs (Random Access Memory), magnetic tapes, floppy disks, and optical data storage devices. The computer-readable storage media may store at least one computer program, which, when executed, can implement the document retrieval method.
[0124] Of course, the electronic device described above may also include other components, such as input / output interfaces and communication components. The input / output interface provides an interface between the processor and a peripheral interface module, which may be an output device, an input device, etc. The communication component is configured to facilitate wired or wireless communication between the electronic device and other devices.
[0125] Those skilled in the art will understand that Figure 7 The structure shown in the figure does not constitute a limitation on the electronic device 700, and the electronic device 700 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0126] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores at least one computer program, and when the at least one computer program is executed by a processor, the above-mentioned document retrieval method can be implemented.
[0127] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it can implement the above-mentioned document retrieval method.
[0128] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A document retrieval method, characterized in that: The method comprises: In response to a document search request including target search text sent by a user account, processing the target search text to obtain a search term in a preset format, wherein the preset format is a data format recognizable by a preset search interface, and the preset search interface is an interface capable of performing searches using multiple search paths; Based on the search terms and the preset search interface, multiple search paths are used to perform searches, and document search results corresponding to each search path are obtained; The document retrieval results corresponding to the multiple retrieval paths are sent to the user account.
2. The method according to claim 1, characterized in that The step of responding to a document search request containing a target search text sent by a user account and processing the target search text to obtain a search term in a preset format includes: In response to the document retrieval request, identifying text content attributes of the target retrieval text; According to the text content attributes of the target search text, different processing methods are adopted to process the target search text to obtain the search term.
3. The method according to claim 2, characterized in that The process of processing the target search text in different ways according to the text content attributes of the target search text to obtain the search term includes: When the text content attribute of the target search text is specific data or category data within the system, the target search text is subjected to word segmentation processing and string matching processing to obtain the search term.
4. The method according to claim 2, characterized in that The process of processing the target search text in different ways according to the text content attributes of the target search text to obtain the search term includes: When the text content attribute of the target search text is natural language, the large language model is called to recognize the target search text to obtain the search term.
5. The method according to claim 4, characterized in that The calling of the large language model to recognize the target search text to obtain the search term includes: Calling the large language model to identify the target search text to obtain multiple candidate intent information; In response to a selection operation of the user account on any candidate intent information, obtaining the candidate intent information selected by the user account; The candidate intent information is processed to obtain the search term.
6. The method according to claim 1, characterized in that The method of searching by using multiple search paths based on the search term and the preset search interface to obtain document search results corresponding to each search path includes: Based on the preset search interface, the search term is processed to obtain search path characteristic data corresponding to different search paths, wherein the search path characteristic data is the data required for searching using the corresponding search path; Search is performed based on the search path feature data corresponding to different search paths to obtain document search results corresponding to each search path.
7. The method according to claim 6, characterized in that The processing of the search terms based on the preset search interface to obtain search path feature data corresponding to different search paths includes: Based on the preset search interface, the search term is directly used as the search path feature data corresponding to the document search path; The searching based on the search path feature data corresponding to different search paths to obtain the document search results corresponding to each search path includes: Based on the search term, querying a document database for document keywords that match the search term, the document database being configured to store correspondences between document keywords and documents in an inverted document manner; When a document keyword matching the search term is stored in the document database, the document corresponding to the matching document keyword is used as the document search result corresponding to the document search path.
8. The method according to claim 6, characterized in that The processing of the search terms based on the preset search interface to obtain search path feature data corresponding to different search paths includes: Based on the preset search interface, obtaining a target feature vector corresponding to the search term, and using the target feature vector as search path feature data corresponding to the vector search path; The searching based on the search path feature data corresponding to different search paths to obtain the document search results corresponding to each search path includes: Based on the target feature vector, searching a vector database for a feature vector whose similarity to the target feature vector is higher than a preset threshold, wherein the vector database is used to store the correspondence between feature vectors and documents; When a feature vector having a similarity with the target feature vector higher than a preset threshold is stored in the vector database, a document corresponding to the retrieved feature vector is used as a document retrieval result corresponding to the vector retrieval path.
9. The method according to claim 6, characterized in that The processing of the search terms based on the preset search interface to obtain search path feature data corresponding to different search paths includes: Based on the preset search interface, the search term is converted into interface data that can be recognized within the system, and the interface data is used as search path feature data corresponding to the internal interface search path; The searching based on the search path feature data corresponding to different search paths to obtain the document search results corresponding to each search path includes: Based on the interface data, the corresponding service interface within the system is called to perform retrieval, and a document retrieval result corresponding to the internal interface retrieval path is obtained.
10. A document retrieval device, characterized in that: The device comprises: a processing module configured to respond to a document search request including target search text sent by a user account, process the target search text, and obtain a search term in a preset format, wherein the preset format is a data format recognizable by a preset search interface, and the preset search interface is an interface capable of performing searches using multiple search paths; A retrieval module, configured to perform retrieval using multiple retrieval paths based on the retrieval term and the preset retrieval interface, and obtain document retrieval results corresponding to each retrieval path; The sending module is used to send the document retrieval results corresponding to the multiple search paths to the user account.
11. An electronic device, characterized in that: It comprises a processor and a memory; the memory stores at least one program code; the at least one program code is used to be called and executed by the processor to implement the document retrieval method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, and when the at least one computer program is executed by a processor, it can implement the document retrieval method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the document retrieval method according to any one of claims 1 to 9 can be implemented.
Citation Information
Patent Citations
Document retrieval method and device, electronic equipment and storage medium
CN111625621A
Retrieval method, system and device and medium
CN112650878A
Document processing method and device, equipment and storage medium
CN114398882A
Question answering method, apparatus and device, and medium
CN117591639A