Method and device for information retrieval, electronic equipment and storage medium
By combining multiple search strategies and information types, using the BM25 algorithm and m3e embedding model, the problem of low information retrieval adaptation rate of large language models is solved, and higher user query intention matching and information timeliness are achieved.
Patent Information
- Application Number
- CN202510523283.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, when searching information, the adaptation rate of mixed search results of large language models is low and cannot effectively match the user's query intention.
A variety of search strategies are used to combine information types, and the information in the text database is obtained through the first search strategy and the second search strategy, and the target text is obtained based on the information type. Taking into account the semantic complexity and entity existence of the user's query information, the BM25 algorithm and the m3e embedding model are used to improve the search accuracy.
The adaptation rate between information retrieval results and user query intentions is improved, and the adaptability and flexibility of the retrieval system is enhanced, especially in campus scenarios, which improves information timeliness and accuracy.
Smart Images

Figure CN120429403A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information retrieval technology, for example, to a method and device for information retrieval, an electronic device, and a storage medium. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, large language models (LLMs) have achieved great success in natural language processing tasks.
[0003] The large language model is pre-trained with a large amount of text data to learn rich language patterns and knowledge information. Before generating an answer, it retrieves an external knowledge base, extracts information related to the user query from the knowledge base, and then inputs this information as context into the large language model to generate an answer.
[0004] In related technologies, when searching external knowledge bases, the system primarily uses a hybrid search method that combines vector search and keyword matching, superimposing the results from each method to obtain the final search result. This results in a low degree of matching the final search result with the user's query intent.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0006] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0007] The embodiments of the present disclosure provide a method and apparatus, an electronic device, and a storage medium for information retrieval, so as to enable the retrieved information to have a higher adaptability to the user's query intent.
[0008] In some embodiments, the method includes: receiving user query information; obtaining the information type of the user query information; searching the user query information in a preset text database according to a preset first search strategy to obtain a first search text; searching the user query information in a preset text database according to a preset second search strategy to obtain a second search text; obtaining a target text based on the information type, the first search text and the second search text.
[0009] In some embodiments, the device includes: a receiving module configured to receive user query information; an information type acquisition module configured to obtain the information type of the user query information; a first retrieval module configured to search the user query information in a preset text database according to a preset first retrieval strategy to obtain a first retrieval text; a second retrieval module configured to search the user query information in a preset text database according to a preset second retrieval strategy to obtain a second retrieval text; and a target text acquisition module configured to obtain a target text based on the information type, the first retrieval text and the second retrieval text.
[0010] In some embodiments, the electronic device includes: a processor and a memory storing program instructions, and the processor is configured to execute the above-mentioned method for information retrieval when running the program instructions.
[0011] In some embodiments, the storage medium stores program instructions, and the program instructions are executed by a processor to implement the above-mentioned method for information retrieval.
[0012] The methods and devices, electronic devices, and storage media for information retrieval provided by the embodiments of the present disclosure can achieve the following technical effects: They can utilize multiple retrieval strategies to search a preset text database for user query information, obtaining a first retrieval text and a second retrieval text, respectively. When obtaining a target text based on the first and second retrieval texts, the first and second retrieval texts are not simply superimposed, but rather the information type of the user's query information is taken into account, resulting in a higher degree of compatibility between the target text and the user's query intent.
[0013] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] One or more embodiments are exemplarily described by corresponding drawings. These exemplary descriptions and drawings do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation. In addition,
[0015] Figure 1 is a schematic diagram of a method for information retrieval provided by an embodiment of the present disclosure;
[0016] Figure 2 is a schematic diagram of a device for information retrieval provided by an embodiment of the present disclosure;
[0017] Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0018] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The accompanying drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.
[0019] In the description and claims of the embodiments of the present disclosure, as well as in the accompanying drawings, the terms "first," "second," and the like are used to distinguish similar items and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate to describe the embodiments of the present disclosure herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions.
[0020] Unless otherwise stated, the term "plurality" means two or more.
[0021] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.
[0022] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0023] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.
[0024] The embodiment of the present disclosure provides a method for information retrieval, the execution subject of which is an electronic device. Among them, the electronic device includes a smart phone, a tablet computer, a computer or a server, etc. The electronic device receives user query information and obtains the information type of the user query information. Then, the user query information is searched in a preset text database according to a first search strategy and a second search strategy, respectively, to obtain a first search text and a second search text. Finally, according to the information type of the user query information, a target text is determined from the first search text and the second search text. In this way, when obtaining the target text, the information type of the user query information is taken into consideration, so that the target text finally obtained has a higher adaptation rate to the user query intention.
[0025] Combine Figure 1 As shown, the embodiment of the present disclosure provides a method for information retrieval, including:
[0026] Step S101: The electronic device receives user query information. The electronic device is provided with a display interface, and the user inputs query information through the display interface. The electronic device receives the query information input by the user.
[0027] Step S102: The electronic device obtains the information type of the user query information.
[0028] In step S103 , the electronic device searches the preset text database for the user query information according to a preset first search strategy to obtain a first search text.
[0029] In step S104, the electronic device searches the preset text database for the user query information according to the preset second search strategy to obtain a second search text.
[0030] Step S105: The electronic device obtains a target text according to the information type, the first search text, and the second search text.
[0031] The information retrieval method provided by the embodiments of the present disclosure can utilize multiple retrieval strategies to search a preset text database for user query information, obtaining a first search text and a second search text, respectively. When obtaining a target text based on the first and second search texts, the method does not simply superimpose the first and second search texts, but also considers the information type of the user's query information, resulting in a higher degree of compatibility between the target text and the user's query intent.
[0032] Optionally, the electronic device obtains the information type of the user query information, including: obtaining the semantic complexity of the user query information. When the semantic complexity of the user query information is greater than or equal to a first set threshold, the information type of the user query information is determined to be a semantic query. When the semantic complexity of the user query information is less than the first set threshold, determine whether there is an entity in the user query information. When an entity exists in the user query information, determine that the information type of the user query information is an entity query. When the semantic complexity of the user query information is less than the first set threshold, and when there is no entity in the user query information, obtain the number of words in the user query information, and when the number of words is less than or equal to the set number threshold, determine that the information type of the user query information is a short query. When the number of words is greater than the set number threshold, determine that the information type of the user query information is a long query. In some embodiments, the set number threshold is 3 or 5, etc.
[0033] In this way, the semantic complexity of the user's query information determines the information type of the query information. Furthermore, when obtaining the final target text based on the information type of the user's query information, the user's semantic intent is taken into account. This results in a higher degree of adaptation of the obtained target text to the user's query intent.
[0034] Optionally, the electronic device obtains the information type of the user query information, including: obtaining the semantic complexity of the user query information. When the semantic complexity of the user query information is less than the second set threshold, the information type of the user query information is determined to be a simple query. When the semantic complexity of the user query information is greater than or equal to the second set threshold, and less than or equal to the first set threshold, the information type of the user query information is determined to be a moderate query. When the semantic complexity of the user query information is greater than the first set threshold, the information type of the user query information is determined to be a complex query. The second set threshold is less than the first set threshold. In some embodiments, the first set threshold is 0.6. The second set threshold is 0.3.
[0035] In this way, the semantic complexity of the user's query information determines the information type of the query information. Furthermore, when obtaining the final target text based on the information type of the user's query information, the user's semantic intent is taken into account. This results in a higher degree of adaptation of the obtained target text to the user's query intent.
[0036] Furthermore, obtaining the semantic complexity of the user query information includes: inputting the user query information into a preset embedding model to obtain a query semantic vector. Obtaining the semantic complexity of the user query information based on the query semantic vector. The preset embedding model is an m3e embedding model.
[0037] Furthermore, determining whether an entity exists in the user query information includes: using a NER (Named Entity Recognition) tool to detect whether an entity exists in the user query information.
[0038] Furthermore, obtaining the number of words in the user query information includes: performing word segmentation processing on the user query information and removing stop words to obtain the number of words.
[0039] Furthermore, obtaining the semantic complexity of the user query information according to the query semantic vector includes: obtaining the variance of the query semantic vector, and determining the variance of the query semantic vector as the semantic complexity of the user query information.
[0040] Furthermore, the variance of the query semantic vector is obtained, including:
[0041] By calculation Get the variance of the query semantic vector. 2 is the variance of the query semantic vector, p is the dimension of the query semantic vector, x i is the value of the i-th dimension of the query semantic vector. μ is the mean of each dimension of the query semantic vector.
[0042] Among them, by calculating Get the mean of each dimension of the query semantic vector.
[0043] Optionally, the preset text database is obtained by: collecting initial text data from at least one data source in real time and extracting metadata from the initial text data; preprocessing each initial text data to obtain candidate text data; and storing each candidate text data and its corresponding metadata in the preset text database.
[0044] In the disclosed embodiments, initial text data is collected from at least one data source in real time through data streaming. This allows the latest text data to be stored in the text database in real time, ensuring that the latest text data is promptly reflected in search results. This improves the timeliness of information retrieval and enables the retrieval system to respond promptly to dynamic data changes.
[0045] In the prior art, text databases are usually updated through regular updates. This lacks an efficient real-time update mechanism, resulting in delayed or incomplete data integration. It cannot meet scenarios with high data update requirements. For example, information retrieval for campus scenarios. Information on campus is frequently updated, such as exam notices, event schedules, etc. Moreover, campus information usually comes from multiple data source channels, such as the school's official website, student forums, various network platforms, etc. If a regular update method is adopted, the latest text data will not be reflected in the search results in a timely manner, affecting the efficiency of campus users in obtaining the latest information. The disclosed embodiment, however, collects initial text data from multiple data sources on campus in real time through data streaming ingestion, so that the preset text database can include the latest information on campus in real time. Campus users can quickly obtain the latest campus developments, such as real-time updated exam notices or event schedules, which greatly improves the timeliness of information retrieval and meets the high timeliness requirements of campus scenarios. In addition, data from multiple data sources is collected and integrated in real time to ensure the comprehensiveness and consistency of the data. When campus users query, the system can dynamically present the latest information from all sources, such as comprehensive results including official website notifications and forum discussions, ensuring the real-time and completeness of the search results.
[0046] Furthermore, each initial text data is pre-processed, including: removing duplicate data in the initial text data, and performing word segmentation processing on each initial text data to remove stop words to obtain candidate text data.
[0047] In some embodiments, the initial text data is segmented using the Jieba word segmentation tool. The metadata of the initial text data includes the time of publication of the initial text data, the author of the initial text data, the source of the data, etc.
[0048] Optionally, searching a preset text database for the user query information according to a preset first search strategy to obtain a first search text includes: obtaining a timeliness weight corresponding to each candidate text data item based on the release time of each candidate text data item; obtaining a relevance score between each candidate text data item and the user query information based on the user query information, each candidate text data item, and each timeliness weight; and obtaining the first search text item based on the relevance score of each candidate text data item.
[0049] This approach takes into account the timeliness of each candidate text data, dynamically improving the search ranking of newly updated candidate text data and prioritizing urgent information. This is particularly true for campus scenarios, where time-sensitive content such as emergency notices and course adjustments are prioritized.
[0050] Furthermore, according to the release time of each candidate text data, the timeliness weight corresponding to each candidate text data is obtained, including: Get the timeliness weight of the lth candidate text data at the current time. lt is the timeliness weight of the lth candidate text data at the current time, t is the current time, t l is the release time of the lth candidate text data, and e is a preset constant. α is the first preset hyperparameter, and β is the second preset hyperparameter. Both α and β are adjustable hyperparameters. The hyperparameters α and β are used to adjust the timeliness weight of candidate text data, balancing the importance of new and old candidate text data. This enhances the retrieval system's ability to handle time-sensitive data. This allows the retrieval system to dynamically adjust result preferences based on different scenarios, such as querying the latest notifications or historical data, enhancing the retrieval system's adaptability and flexibility.
[0051] In some embodiments, the timeliness weight of each candidate text data is updated at intervals of a preset time, for example, the preset time is 1 hour.
[0052] Furthermore, based on the user query information, each candidate text data, and each timeliness weight, a relevance score between each candidate text data and the user query information is obtained, including: using the BM25 algorithm to calculate the user query information, the candidate text data, and the timeliness weight to obtain the relevance score between the candidate text data and the user query information. In this way, using the BM25 algorithm, through keyword matching, the first search text related to the user query information can be retrieved more quickly.
[0053] Specifically, by calculating the score l =BM25(Q, D l )·ω ltGet the relevance score between the lth candidate text data and the user query information. l is the correlation score between the lth candidate text data and the user query information, Q is the user query information, D l is the lth candidate text data, ω lt is the timeliness weight of the lth candidate text data at the current time.
[0054] Based on the traditional BM25 algorithm, a weighting mechanism is introduced by taking into account the release time of the candidate text data, giving the latest updated files a higher weight, thereby improving the timeliness of the search results and enhancing the retrieval system's adaptability to dynamic data.
[0055] In some embodiments, before calculating the user query information, candidate text data, and timeliness weight using the BM25 algorithm, the user query information is segmented and stop words are removed.
[0056] Furthermore, obtaining the first search text based on the relevance score of each candidate text data includes: sorting the candidate text data in descending order of relevance score, and determining the candidate text data ranked in the first preset rank as the first search text. For example, the first preset rank is 50. That is, the candidate text data ranked in the top 50 in relevance score is determined as the first search text. After obtaining the first search text, the first search text set and the relevance score corresponding to each first search text are stored.
[0057] Optionally, searching a preset text database for the user query information according to a preset second search strategy to obtain a second search text includes: obtaining a similarity score between each candidate text data and the user query information, and obtaining the second search text according to the similarity score of each candidate text data.
[0058] Furthermore, obtaining a similarity score between the candidate text data and the user query information includes: obtaining a candidate text data vector corresponding to the candidate text data, obtaining a user query information vector corresponding to the user query information, calculating a cosine similarity between the candidate text data vector and the user query information vector, and determining the cosine similarity as the similarity score between the candidate text data and the user query information.
[0059] The similarity score measures the semantic similarity between the candidate text data and the user's query information. By calculating the cosine similarity between the candidate text data vector and the user's query information vector, we can understand the deeper semantic needs of the user's query information. This takes into account the semantic intent of the user's query information and retrieves text data that is semantically relevant to the user, thereby improving the accuracy of the search results.
[0060] In some embodiments, the m3e embedding model is used to encode the alternative text data to generate an alternative text data vector. The alternative text data vector is a 768-dimensional vector. The m3e embedding model is used to encode the user query information to generate a user query information vector. The user query information vector is a 768-dimensional vector. In this way, the m3e embedding model can generate high-quality text embeddings, which can capture the deep semantic features of user query information and alternative text data. This can more accurately understand the user's semantic intentions. In addition, the m3e embedding model is regularly fine-tuned according to the set time to improve its adaptability to the newly updated alternative text data, so as to optimize the accuracy of the semantic representation. In this way, the m3e embedding model is regularly fine-tuned to capture the semantics of the newly added alternative text data and improve the semantic matching accuracy of long-tail queries.
[0061] Further, the cosine similarity between the candidate text data vector and the user query information vector is calculated, including: Obtain the cosine similarity between the candidate text data vector and the user query information vector. l ) is the cosine similarity between the lth candidate text data vector and the user query information vector, q is the user query information vector, d l is the lth candidate text data vector. That is, sin(q, d l ) is the similarity score between the lth candidate text data and the user query information.
[0062] Furthermore, obtaining a second search text based on the similarity scores of each candidate text data includes: sorting each candidate text data in descending order of similarity score, and determining the candidate text data ranked in the top second preset rank as the second search text. For example, the second preset rank is 50. That is, the candidate text data ranked in the top 50 in terms of similarity score are determined as the second search text. After obtaining the second search text, the second search text set and the similarity score corresponding to each second search text are stored.
[0063] Optionally, obtaining a target text based on the information type, the first search text, and the second search text includes: obtaining a matching weight based on the information type; obtaining a matching score for a candidate text based on the matching weight; wherein the candidate text includes the first search text and the second search text; and obtaining the target text based on the matching score of each candidate text.
[0064] Obtaining the first search text through the first search strategy of keyword matching can improve the search efficiency. Obtaining the second search text through the second search strategy of semantic retrieval can improve the accuracy of the search. Then determine the target text from the first search text and the second search text. In this way, through the hybrid search strategy, while ensuring a quick response, the accuracy of the search results is also significantly improved. It is more suitable for the actual needs of high-frequency queries and large-scale data in campus scenarios. In addition, when determining the target text, the matching weight is determined according to the type of user query information, and the weights of the first search text and the second search text are dynamically adjusted. In this way, the information type of the user query information is taken into account, so that the final search result has a higher adaptation rate to the user's query intention.
[0065] Optionally, obtaining the matching weight according to the information type of the user query information includes: performing a table lookup operation in a preset data table using the information type to find the matching weight corresponding to the information type. The preset data table stores the correspondence between the information type and the matching weight.
[0066] In some embodiments, Table 1 is an example table of a preset data table. As shown in Table 1, when the information type is a simple query, the corresponding matching weight is 0.7. In this case, keyword search is more suitable. When the information type is a medium query, the corresponding matching weight is 0.5. In this case, both semantic understanding and keyword search are applicable. When the information type is a complex query, the corresponding matching weight is 0.2. In this case, semantic understanding should be prioritized.
[0067] Table 1
[0068] Information Type Matching weight Simple query 0.7 Moderate query 0.5 Complex queries 0.2
[0069] Further, the matching score of the candidate text is obtained according to the matching weight, including: calculating the final_score(D j )=λ·score_sparse(D j )+(1-λ)·score_dense(D j ) to obtain the matching score of the candidate text. Among them, final_score(D j ) is the matching score of the j-th candidate text, λ is the matching weight, score_sparse(D j ) is the relevance score of the j-th candidate text, socre_dense(D j ) is the similarity score of the j-th candidate text.
[0070] Because the candidate texts are text data in the first search text set and the second search text set. In this way, the candidate text may belong to both the first search text set and the second search text set. However, it may also belong to only one of the text sets. In the case that the jth candidate text belongs to the first search text set but not the second search text set, the similarity score of the jth candidate text is zero. In the case that the jth candidate text belongs to the second search text set but not the first search text set, the relevance score of the jth candidate text is zero.
[0071] Furthermore, the target text is obtained based on the matching scores of each candidate text, including: sorting the candidate texts in descending order of matching scores, and determining the candidate texts ranked in the top three preset positions as the target texts. For example, the third preset position is 4. That is, the candidate texts ranked in the top four in matching scores are determined as the target texts. The weights of the keyword search results and the semantic search results are dynamically adjusted based on the user query information. In this way, when fusing the keyword search results and the semantic search results, the information type of the user query information is taken into account, so that the final search results have a higher adaptation rate to the user's query intent.
[0072] Optionally, after obtaining the target text, the method further includes: inputting the user query information and the target text into a preset large model to generate a natural language answer.
[0073] Optionally, after the natural language answer is generated, the natural language answer is displayed through a display interface for user viewing.
[0074] In some embodiments, the target text is concatenated into a final document in descending order of matching scores. The user query and the final document are fed into a pre-defined macro model to generate a natural language answer. The pre-defined macro model is the DeepSeek-R1 inference macro model.
[0075] For example, the user's query information is: "What are the research achievements of University A in the field of agriculture?" The retrieved target texts include: Target text A: "The research focus of University A in the field of agriculture includes the optimization of bee breeding technology. The research team has developed a new beehive design that can significantly improve the honey production and disease resistance of bees." Target text B: "The research team of the College of Agriculture of University A has made a breakthrough in rapeseed planting technology. By improving varieties and optimizing irrigation systems, rapeseed yields have increased by more than 20%." Target text C: "University A has carried out a number of studies in agricultural ecology, especially in the study of the relationship between bees and crop pollination. It has proposed the concept of 'eco-friendly agriculture' and reduced the use of chemical pesticides." Target text D: "The scientific research team of University A has developed an intelligent agricultural monitoring system that uses sensors to monitor the humidity, temperature and light conditions of farmland in real time, helping farmers manage crops more efficiently." Among them, the matching scores from large to small are target text A, target text B, target text C and target text D.
[0076] After entering the preset large model, prompt="""
[0077] You are a know-it-all from University A, so please answer the questions based on the context I provided.
[0078] Context:
[0079] A key focus of the University of Malaya's agricultural research is optimizing beekeeping techniques. The research team has developed a new beehive design that significantly increases honey production and disease resistance.
[0080] A research team from the Faculty of Agriculture at the University of Malaya has achieved a breakthrough in rapeseed cultivation technology. By improving varieties and optimizing irrigation systems, the team has increased rapeseed yields by over 20%.
[0081] The University of Alabama has conducted numerous studies in agroecology, particularly on the relationship between bees and crop pollination. This research has led to the concept of "eco-friendly agriculture," which reduces the use of chemical pesticides. A research team at the University of Alabama has developed an intelligent agricultural monitoring system that uses sensors to monitor humidity, temperature, and light conditions in farmland in real time, helping farmers manage their crops more efficiently.
[0082] Question: What are University A’s research achievements in the field of agriculture?
[0083] """
[0084] Combine Figure 2As shown, an embodiment of the present disclosure provides a device for information retrieval, including a receiving module 201, an information type acquisition module 202, a first retrieval module 203, a second retrieval module 204 and a target text acquisition module 205. The receiving module 201 is configured to receive user query information. The information type acquisition module 202 is configured to obtain the information type of the user query information. The first retrieval module 203 is configured to search the user query information in a preset text database according to a preset first retrieval strategy to obtain a first retrieval text. The second retrieval module 204 is configured to search the user query information in a preset text database according to a preset second retrieval strategy to obtain a second retrieval text. The target text acquisition module 205 is configured to obtain a target text based on the information type, the first retrieval text and the second retrieval text.
[0085] The information retrieval device provided by the embodiments of the present disclosure can utilize multiple retrieval strategies to search a preset text database for user query information, obtaining a first search text and a second search text, respectively. When obtaining a target text based on the first and second search texts, the first and second search texts are not simply superimposed; rather, the information type of the user's query information is considered, resulting in a higher degree of compatibility between the target text and the user's query intent.
[0086] The information retrieval device also includes a data update module configured to obtain a preset text database by: collecting initial text data from at least one data source in real time and extracting metadata from the initial text data; preprocessing each initial text data to obtain candidate text data; and storing each candidate text data and corresponding metadata in the preset text database. The metadata includes the release time of the initial text data.
[0087] The first retrieval module is configured to search the user query information in a preset text database according to a preset first retrieval strategy in the following manner to obtain a first retrieval text: according to the release time of each candidate text data, obtain the timeliness weight corresponding to each candidate text data; according to the user query information, each candidate text data and each timeliness weight, obtain the correlation score between each candidate text data and the user query information; according to the correlation score of each candidate text data, obtain the first retrieval text.
[0088] The second retrieval module is configured to search the user query information in a preset text database according to a preset second retrieval strategy in the following manner to obtain a second retrieval text: obtaining a similarity score between each candidate text data and the user query information; and obtaining the second retrieval text based on the similarity score of each candidate text data.
[0089] The target text acquisition module is configured to obtain the target text according to the information type, the first search text and the second search text in the following manner: obtaining a matching weight according to the information type; obtaining a matching score of the candidate text according to the matching weight; wherein the candidate text includes the first search text and the second search text; and obtaining the target text according to the matching score of each candidate text.
[0090] The target text acquisition module is configured to obtain the matching weight according to the information type in the following manner: using the information type, performing a table lookup operation in a preset data table to find the matching weight corresponding to the information type; wherein the preset data table stores the correspondence between the information type and the matching weight.
[0091] The apparatus for information retrieval further includes an answer generation module configured to input the user query information and the target text into a preset macro model to generate a natural language answer.
[0092] Combine Figure 3 As shown, an embodiment of the present disclosure provides an electronic device 300, including a processor 304 and a memory 301 storing program instructions. Optionally, the electronic device may further include a communication interface 302 and a bus 303. The processor 304, the communication interface 302, and the memory 301 may communicate with each other via the bus 303. The communication interface 302 may be used for information transmission. The processor 304 may call the program instructions in the memory 301 to execute the method for information retrieval of the above embodiment.
[0093] In addition, the logic instructions in the memory 301 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0094] Memory 301, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present disclosure. Processor 304 executes the program instructions / modules stored in memory 301 to perform functional applications and data processing, thereby implementing the information retrieval method in the above-mentioned embodiments.
[0095] The memory 301 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal device. Furthermore, the memory 301 may include high-speed random access memory and non-volatile memory.
[0096] Optionally, the electronic device includes a smartphone, a tablet, a computer, or a server.
[0097] An embodiment of the present disclosure provides a storage medium storing program instructions, wherein the program instructions are executed by a processor to implement the above-mentioned method for information retrieval.
[0098] The technical solution of the embodiments of the present disclosure may be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present disclosure. The aforementioned storage medium may be a non-transitory storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code, or a transient storage medium.
[0099] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to also include plural forms. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variations "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups of these. In the absence of further restrictions, an element defined by the sentence "comprising a..." does not exclude the presence of other identical elements in the process, method or device that includes the element. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments can be referenced to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can be found in the description of the method part.
[0100] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0101] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical functional division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, and can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of the present disclosure may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0102] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architectures, functions and operations of the systems, methods and computer program products according to the embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action, or may be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A method for information retrieval, characterized in that include: Receive user query information; Acquire the information type of the user query information; According to a preset first search strategy, the user query information is searched in a preset text database to obtain a first search text; According to a preset second search strategy, searching the user query information in a preset text database to obtain a second search text; A target text is obtained according to the information type, the first search text and the second search text.
2. The method according to claim 1, characterized in that The default text database is obtained in the following ways: Collecting initial text data from at least one data source in real time, and extracting metadata of the initial text data; Preprocessing each of the initial text data to obtain candidate text data; The candidate text data and corresponding metadata are stored in a preset text database.
3. The method according to claim 2, characterized in that The metadata includes the release time of the initial text data; according to a preset first search strategy, searching the user query information in a preset text database to obtain a first search text, including: According to the release time of each candidate text data, the timeliness weight corresponding to each candidate text data is obtained; Obtaining a relevance score between each candidate text data and the user query information based on the user query information, each candidate text data, and each timeliness weight; The first search text is obtained according to the relevance score of each candidate text data.
4. The method according to claim 2, characterized in that According to a preset second search strategy, searching the user query information in a preset text database to obtain a second search text includes: Obtaining similarity scores between each candidate text data and the user query information; A second search text is obtained based on the similarity scores of the candidate text data.
5. The method according to claim 1, wherein Obtaining a target text according to the information type, the first search text, and the second search text includes: Obtaining a matching weight according to the information type; Obtaining a matching score of the candidate text according to the matching weight; wherein the candidate text includes a first search text and a second search text; The target text is obtained based on the matching scores of each candidate text.
6. The method according to claim 5, characterized in that Obtaining a matching weight based on the information type includes: Using the information type, a table lookup operation is performed in a preset data table to find out the matching weight corresponding to the information type; wherein the preset data table stores the corresponding relationship between the information type and the matching weight.
7. The method according to any one of claims 1 to 6, characterized in that After obtaining the target text, it also includes: The user query information and the target text are input into a preset macro model to generate a natural language answer.
8. A device for information retrieval, characterized in that: include: A receiving module is configured to receive user query information; An information type acquisition module is configured to acquire the information type of the user query information; A first search module is configured to search the user query information in a preset text database according to a preset first search strategy to obtain a first search text; A second search module is configured to search the user query information in a preset text database according to a preset second search strategy to obtain a second search text; The target text acquisition module is configured to obtain a target text according to the information type, the first search text and the second search text.
9. An electronic device comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the method for information retrieval according to any one of claims 1 to 7 when running the program instructions.
10. A storage medium storing program instructions, characterized in that: The program instructions are executed by a processor to implement the method for information retrieval according to any one of claims 1 to 7.