Information retrieval method, device, medium and equipment based on multi-channel recall

By identifying the target service type in information retrieval and performing multiple recall searches, and using confidence calculation to output search results, the problem of insufficient efficiency and accuracy of unstructured data text retrieval in the prior art is solved, and more efficient and accurate information retrieval is achieved.

CN119577117BActive Publication Date: 2025-05-06HANGZHOU SHUYIXIN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510138509.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-06
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The prior art is difficult to meet the actual needs of vertical fields when processing text retrieval of unstructured data, and the search efficiency and accuracy are insufficient, especially when processing complex and changeable text data.

Method used

By identifying the target business type to which the text search content belongs, performing multiple recall searches within the corresponding search scope, combining the criterion label recall model, hierarchical label recall model and keyword recall model, the confidence of each recall data is calculated, and a preset number of recall data is output as the search result.

Benefits of technology

It improves the efficiency and accuracy of information retrieval, can more accurately match user needs, reduce the retrieval scope of irrelevant data, and enhances the understanding and utilization of semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119577117B_ABST
    Figure CN119577117B_ABST
Patent Text Reader

Abstract

The present application provides an information retrieval method, device, medium and equipment based on multi-way recall, which belongs to the field of information retrieval technology. The method includes: obtaining the user's text retrieval content; identifying the target business type to which the text retrieval content belongs; converting the text retrieval content into a retrieval vector, calling a preset multi-way recall model to search within the retrieval range corresponding to the target business type, and obtaining the recall data output by each recall model; calculating the first confidence of each recall data relative to the text retrieval content; and outputting a preset number of recall data as the retrieval result of the text retrieval content based on the first confidence. The present application can improve the accuracy and efficiency of information retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information retrieval technology, and in particular to an information retrieval method, device, medium and equipment based on multi-path recall. Background Art

[0002] With the widespread popularity of mobile terminals and industrial Internet terminals, data forms have become more diverse, and unstructured data contains a wide range of application scenarios and potential commercial value that can be mined. Traditional machine learning, deep convolutional neural networks (CNNs) and recurrent neural networks (RNNs) face huge challenges in processing such data, especially in vertical field applications of text data. These technologies are difficult to meet actual implementation needs, making the implementation of the industrial Internet face many difficulties.

[0003] In the process of retrieving and processing unstructured data, it is necessary to implement algorithms in vertical industry fields based on actual application scenarios and industry standards. Algorithm selection, the degree of adaptation of subsequent algorithm solutions, and algorithm optimization and iteration are all crucial.

[0004] Existing text classification and grading technologies are based on character rules and other labeling methods, which have poor flexibility. Specifically, 1) Rules can generally only process pre-agreed forms and are difficult to cope with complex and changeable actual text data; 2) Poor scalability. When more types of text patterns need to be processed, the complexity of regular expressions will increase dramatically, and maintenance costs will also increase; 3) Unable to understand semantic information: Regular expressions and other logic can only understand text at the character level, and cannot take into account the context and the semantic information behind it.

[0005] Some improved methods have begun to try to classify data based on semantic understanding. For example, the semantic retrieval method based on multi-way recall can improve the accuracy of retrieval to a certain extent. In the existing multi-way recall retrieval technology disclosed in CN118820400A, CN115809312B, etc., although semantic understanding is also performed on the data to be retrieved, in-depth analysis of the text data of varying lengths and the fields in which they are located is rarely performed, resulting in a large range of data to be searched and low accuracy of semantic understanding. Therefore, the retrieval efficiency and accuracy of the existing technology need to be improved. Summary of the invention

[0006] The purpose of the present invention is to provide a new information retrieval method, device, medium and equipment based on multi-way recall to solve at least one technical problem in the prior art.

[0007] In a first aspect of the present application, a multi-path recall-based information retrieval method is provided, the method comprising:

[0008] Get the user's text search content;

[0009] Identify the target business type to which the text search content belongs;

[0010] Convert the text search content into a search vector, call a preset multi-channel recall model to search in the search range corresponding to the target business type, and obtain the recall data output by each recall model;

[0011] Calculating a first confidence level of each recall data relative to the text retrieval content;

[0012] A preset number of recalled data is output as a retrieval result of the text retrieval content based on the first confidence level.

[0013] Optionally, after obtaining the text search content of the user, the method further includes:

[0014] Calculating a second confidence level that the text search content belongs to a rule-based search;

[0015] When the second confidence exceeds a preset confidence threshold, determining that the search type of the text search content is a rule-based search, and querying a preset number of search results matching the keywords in the text search content from a global database;

[0016] When the second confidence is lower than a preset confidence threshold, the search type is determined to be a model-based search, and the step of identifying the target business type to which the text search content belongs is performed.

[0017] Optionally, converting the text search content into a search vector includes:

[0018] Segmenting the text search content to form keywords corresponding to the text search content;

[0019] Calculate the information content score of each keyword, where the information content score is used to reflect the importance of the keyword in the text search content;

[0020] The keyword is serialized based on the information score to form serialized characters, and the serialized characters are converted into the retrieval vector through a preset multi-layer encoder.

[0021] Optionally, the segmenting the text search content to form keywords corresponding to the text search content includes:

[0022] Performing Jieba word segmentation on the text search content to form initial words;

[0023] Performing synonym expansion on the initial word;

[0024] The initial word and the corresponding expanded synonyms are randomly inserted and replaced and deleted to obtain keywords corresponding to the text search content.

[0025] Optionally, calling a preset multi-channel recall model to search within a search range corresponding to the target business type to obtain recall data output by each recall model includes:

[0026] For one or more of the multiple recall models, the search vector is used as an entry point of the top layer in a pre-set hierarchical navigable small-world algorithm, and a neighboring point search of the current layer is performed from the entry point;

[0027] The retrieved adjacent points of the current layer are used to enter the next layer to continue the adjacent point search until the adjacent points of the last layer are found, and the data corresponding to the adjacent points of the last layer are used as the recalled data.

[0028] Optionally, the calculating of a first confidence level of each recalled data relative to the text retrieval content includes:

[0029] The text search content is spliced ​​with each recalled data, the text semantic information of the spliced ​​data formed by each recalled data is identified, and the confidence corresponding to each recalled data is calculated based on the text semantic information.

[0030] Optionally, the multi-way recall model includes a criterion label recall model, a hierarchical label recall model and a keyword recall model.

[0031] Optionally, calling a preset multi-channel recall model to search within a search range corresponding to the target business type to obtain recall data output by each recall model includes:

[0032] Calling the criterion label recall model to match the search vector with the criterion label to obtain a target criterion label that matches the search vector, and querying a first number of recall data that matches the target criterion label from the search scope; and / or

[0033] Calling the hierarchical label recall model to determine the target hierarchical level to which the search vector belongs, and querying a second amount of recall data matching the target hierarchical level from the search scope; and / or

[0034] The keyword recall model is called to split the search vector into keywords, and a third number of recalled data matching the keywords are searched in the search scope.

[0035] In a second aspect of the present application, a multi-path recall-based information retrieval device is provided, the device comprising:

[0036] A search content acquisition module is used to obtain the user's text search content;

[0037] A business type identification module, used to identify the target business type to which the text search content belongs;

[0038] A recall data acquisition module, used to convert the text search content into a search vector, call a preset multi-channel recall model to search in the search range corresponding to the target business type, and obtain the recall data output by each recall model;

[0039] The retrieval result output module is used to calculate a first confidence level of each recalled data relative to the text retrieval content; and output a preset number of recalled data as retrieval results of the text retrieval content based on the first confidence level.

[0040] In a third aspect of the present application, a computer-readable storage medium is provided, on which executable instructions are stored. When the executable instructions are executed by a processor, the processor executes the information retrieval method based on multi-way recall in any embodiment of the present application.

[0041] In a fourth aspect of the present application, an electronic device is provided, including: one or more processors;

[0042] A memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors execute the information retrieval method based on multi-way recall in any embodiment of the present application.

[0043] The information retrieval method, apparatus, medium and device based on multi-channel recall in the present application identify the target business type to which the text retrieval content belongs, and perform retrieval within the retrieval range corresponding to the target business type, without the need to search in all data, thereby improving retrieval efficiency; and in the retrieval process, a multi-channel recall method is adopted, and confidence calculation is performed on the recall data output by each recall model, and an appropriate number of recall data are selected as retrieval results based on the confidence, which can improve the accuracy of the retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope of the present application.

[0045] Figure 1 is a flowchart of an information retrieval method based on multi-channel recall in one embodiment;

[0046] Figure 2 A schematic diagram of a process of converting text search content into a search vector in one embodiment;

[0047] Figure 3 A schematic diagram of a process of segmenting text search content to form keywords corresponding to the text search content in one embodiment;

[0048] Figure 4 is a structural block diagram of an information retrieval device based on multi-way recall in one embodiment;

[0049] Figure 5 is a structural block diagram of an information retrieval device based on multi-path recall in another embodiment;

[0050] Figure 6 FIG. 1 is a schematic diagram of the structure of an electronic device in an embodiment. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0052] All terms (including technical and scientific terms) used in this application have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used here should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0053] For example, the terms "first", "second", etc. used in this application are only used to distinguish similar objects and to differentiate the first object from another object, rather than to describe a specific order or sequence, and cannot be understood as indicating or implying relative importance.

[0054] This application proposes an information retrieval method based on multi-way recall, combined with Figure 1 As shown, the method includes:

[0055] Step 110, obtaining the user's text search content.

[0056] In this embodiment, the text search content is the text to be searched input by the user. The user can enter the text to be searched in the relevant input interface, for example, if the user enters "risk mitigation information", the "risk mitigation information" is the corresponding text search content.

[0057] Step 120: Identify the target business type to which the text search content belongs.

[0058] In this embodiment, the electronic device can search for data related to the text search content in a preset database, and the database can be a database formed by relevant data in a preset vertical field. Among them, the corresponding business type is set in advance according to the characteristics of the vertical field in which the relevant data is located, and different business types correspond to different search scopes. Among them, the number of business types can be divided according to the data characteristics of the vertical field, and each business type can be further divided into multiple levels of subcategories, and the number of types in each level of subcategories is not necessarily the same. Similarly, each level of subcategories corresponds to a different search scope. The business type to which the text search content belongs is the target business type.

[0059] In one embodiment, a pre-trained semantic vector model is called to divide text data in a vertical field into business types to form multiple business types. Optionally, the semantic vector model is a sbert_textcnn model.

[0060] For example, text data in the financial field can be divided into 9 business types. Taking the text search content of "risk mitigation information" as an example, it can be determined that the business type to which the text search content belongs is the "bank risk management" type.

[0061] Step 130, converting the text search content into a search vector, calling a preset multi-channel recall model to search in the search range corresponding to the target business type, and obtaining the recall data output by each recall model.

[0062] In this embodiment, natural language processing technology (such as word embedding, BERT, etc.) is used to segment the text search content, and word vector conversion is performed for each segmented word, and the converted word vectors are merged to form a corresponding search vector. Among them, word vector conversion can be performed on each or part of the segmented words to form a corresponding search vector.

[0063] Specifically, semantic recognition is performed on the text retrieval content to identify the importance of each word in the text retrieval content, and word vector conversion is performed based on the importance, so that the formed retrieval vector can accurately reflect the data expected to be obtained from the text retrieval content.

[0064] The retrieval vector is the input of the relevant recall model. The electronic device is pre-set with multiple recall models, each of which has different model parameters and can obtain corresponding different retrieval results (ie, recall data).

[0065] By calling each recall model to search within the data range corresponding to the target business type and obtaining the recall data output by each recall model respectively, the comprehensiveness of the search results can be improved.

[0066] In one embodiment, the electronic device pre-sets the number of recall data output by each recall model. For example, there are N recall models, and the number of recall data output by each recall model is K. Then the total recall data obtained is N×K.

[0067] Among them, the recall model can be one or more of a criterion label recall model, a hierarchical label recall model and a keyword recall model. The criterion label recall model is a model set based on preset standard criteria, which can be a national standard criterion for the field to which the retrieval data belongs. The criterion label recall model can be better applied in the field to which it belongs. The hierarchical label recall model is to classify and grade the business based on the field to which the retrieval data belongs. For example, each of the above business types is divided into multiple subcategories, and the minimum level of the text retrieval content under the target business type is identified. Data retrieval is performed at the minimum level to capture the appropriate data range. Keyword recall is to perform semantic understanding based on the text retrieval content, and to obtain a more accurate basis for judgment by setting data rules through business and screening word frequency.

[0068] Taking the hierarchical label recall model as an example, for a certain business type, the two 4-level hierarchical labels it contains are "business-account information-account basic information-account basic information overview" and "customer-unit-unit basic information-unit basic information overview". When the text search content is "basic information overview", it can be identified and located to which of the hierarchical labels such as "business-account information-account basic information-account basic information overview", "customer-unit-unit basic information-unit basic information overview", etc. it belongs, and then it can be searched within the data range corresponding to the identified hierarchical label to obtain the recalled data. Compared with the traditional character rule-based search that searches within all ranges involving "basic information", the hierarchical label recall model can accurately obtain the target data to be retrieved.

[0069] Each recall model can use a corresponding appropriate algorithm. For example, one or more of the three recall models (criteria label recall model, hierarchical label recall model and keyword recall model) can be retrieved using proximity search and hierarchical navigable small world (HNSW) algorithms.

[0070] Step 140, calculating a first confidence level of each recalled data relative to the text retrieval content.

[0071] In this embodiment, confidence calculation is performed on the recall data output by each recall model. The first confidence is used to reflect the accuracy of the corresponding recall data relative to the text retrieval content. The higher the confidence, the higher the accuracy of the retrieved data as the retrieval result expected by the user.

[0072] Optionally, the electronic device pre-sets a corresponding confidence calculation model, and after obtaining each recall data, calculates the corresponding first confidence by using the confidence calculation model.

[0073] Step 150: output a preset number of recalled data as a retrieval result of the text retrieval content based on the first confidence level.

[0074] In this embodiment, the electronic device may predetermine the number of data that needs to be output (ie, a preset number), and select the preset number of recalled data with the highest confidence as the retrieval result.

[0075] The information retrieval method based on multi-channel recall in the present application identifies the target business type to which the text retrieval content belongs, and performs retrieval within the retrieval range corresponding to the target business type. There is no need to search in all the data, thereby improving the retrieval efficiency. In addition, a multi-channel recall method is adopted in the retrieval process, and the confidence calculation is performed on the recall data output by each recall model. Based on the confidence, an appropriate number of recall data are selected as the retrieval result, which can improve the accuracy of the retrieval.

[0076] In one embodiment, before step 120, the method further includes: identifying the search type to which the text search content belongs, and when it is a rule-based search, querying a preset number of search results that match the keywords in the text search content from a global database; when it is a model-based search, executing identification of the target business type to which the text search content belongs.

[0077] In this embodiment, the search type may include rule-based search and model-based search. The rule may include keyword matching, Boolean logic operations, grammatical rules, etc. Rule-based search represents a search type in which the information reflected by the text search content is clear, and model-based search represents a search type in which the search information reflected by the text search content is unclear. When the search information of the text search content is identified to be clear enough, it is determined to be a rule-based search, otherwise, it is determined to be a model-based search.

[0078] In one embodiment, a second confidence level is calculated that the text search content belongs to a rule-based search; when the second confidence level exceeds a preset confidence threshold, the search type of the text search content is determined to be a rule-based search, and a preset number of search results matching the keywords in the text search content are queried from a global database; when the second confidence level is lower than a preset confidence threshold, the search type is determined to be a model-based search, and the target business type to which the text search content belongs is identified.

[0079] In this embodiment, the second confidence level is used to reflect the clarity of the search information reflected by the text search content. The higher the second confidence level, the clearer the search information reflected by the text search content. When it exceeds the preset confidence level threshold, it is directly determined to be a rule-based search.

[0080] When it is a rule-based search, the data matching the keywords in the text search content is directly queried from the global database, and the corresponding preset number of search results are output. For example, when the text search content is "Salary of XX employees in XX department in October 2023", very clear search content is identified from the text search content, and the corresponding second confidence calculated is higher than the preset confidence threshold, it is determined to be a rule-based search, and the data that meets the text search content can be directly retrieved. When the text search content is the above-mentioned "risk mitigation information", the search information is not clear enough, and the calculated second confidence is lower than the preset confidence threshold, it is determined to be a model-based search.

[0081] By identifying the search type, data can be directly retrieved for text search content with very clear search content, and the search data required by the user can be output accurately and quickly. When the search content is not clear enough, multi-way recall search is performed to improve the accuracy of the search data output.

[0082] In one embodiment, Figure 2 As shown, the text retrieval content is converted into a retrieval vector, including:

[0083] Step 210, segmenting the text search content to form keywords corresponding to the text search content.

[0084] In this embodiment, the word segmentation method may include a combination of one or more analysis methods such as Jieba word segmentation, forward maximum matching method word segmentation, reverse maximum matching method, hidden Markov model word segmentation, etc. By segmenting the text search content, one or more corresponding keywords can be obtained.

[0085] For example, the text search content "risk mitigation information" is segmented, and the keywords obtained after segmentation are "risk", "mitigation", and "information". Optionally, all the words obtained by the segmentation can be used as keywords, and the words obtained by the segmentation can also be expanded, and the expanded words can also be used as corresponding keywords.

[0086] Step 220, calculating the information score of each keyword.

[0087] The information score is used to reflect the importance of keywords in text retrieval content. In text retrieval content, different keywords have different corresponding importance, and some keywords may even interfere with the retrieval results.

[0088] Specifically, the corresponding information score can be calculated by combining one or more algorithms such as the preset TF-IDF (term frequency-inverse document frequency) algorithm and the TextRank algorithm. Taking the TF-IDF algorithm as an example, the term frequency (TF) of each keyword (such as the above-mentioned keywords "risk", "mitigation", and "information") in the text retrieval content can be counted, and then the inverse document frequency (IDF) of each keyword in the entire corpus can be calculated. Finally, TF and IDF are multiplied to obtain the TF-IDF value of each keyword, and the TF-IDF value is used as the information score.

[0089] Taking the TextRank algorithm as an example, the keywords formed can be converted into a vocabulary graph, where each node represents a keyword and the edges between nodes represent the co-occurrence relationship between keywords. The co-occurrence relationship can be determined by setting a window size (such as 2, 3 or 5 words), that is, if two words are adjacent in the text and are located in the same window, there is an edge between them. For example, for the text retrieval content "How to effectively manage enterprise resources", the window size is set to 2, and the following vocabulary graph is constructed: 1. How -> Effective; 2. Effective -> Management; 3. Management -> Enterprise; 4 Enterprise -> Resources.

[0090] For the constructed vocabulary graph, the score of each node is calculated separately, and this score is the information score. Specifically, the calculation process includes two processes: initialization score and iterative update. In the process of initialization score, an initial score is assigned to each node. For example, the initialization score of each node can be uniformly set to 1. In the process of iterative update, for each node, the score of its neighboring nodes is considered when calculating its score. The following formula can be called to calculate the score of each iteration:

[0091]

[0092] in, Indicates the node in the current iteration The score,d is the damping coefficient, which can be set to 0.85, for example. Indicates that the node The node set of Represents a slave node The set of edges pointing to other nodes. express The number of edges.

[0093] In each iteration, the score of each node is updated according to the above formula until the scores of all nodes converge (that is, the difference between the scores obtained in the previous and next iterations is less than the preset score threshold) or the number of iterations reaches the preset iteration threshold. When the iteration is terminated, the score output by each node in the last iteration is used as the information score of the corresponding keyword.

[0094] Step 230, serialize the keywords based on the information score to form serialized characters, and convert the serialized characters into retrieval vectors through a preset multi-layer encoder.

[0095] Optionally, the length of the serialized characters corresponding to each keyword is determined based on the information score. The greater the information score, the longer the length of the corresponding serialized characters, and the more information the corresponding keyword embodies.

[0096] In one embodiment, the length of the serialized characters corresponding to each keyword is kept consistent. Keywords with information volume scores lower than a preset score threshold are discarded, and only keywords with information volume scores higher than the score threshold are serialized.

[0097] Get the common segmentation text of each keyword in sentencePiece , based on the universal segmentation text, calculate the corresponding serialized characters .in, It represents the word vector encoding of the i-th keyword. It represents the serialized characters of the ith keyword. The ith keyword is the ith keyword that is retained and needs to be serialized.

[0098] For the serialized characters, call BCEmbedding to perform multi-layer encoder inference to obtain the corresponding retrieval vector V. , It represents the search vector of the ith keyword of the text after being extracted by the Transformer encoder. The dimension of each search vector can be a pre-set dimension, for example, each search vector can be set is a vector of 768 dimensions.

[0099] In one embodiment, Figure 3 As shown, the text search content is segmented to form keywords corresponding to the text search content, including:

[0100] Step 310, perform Jieba word segmentation on the text search content to form initial words.

[0101] In this embodiment, the word obtained by jieba segmentation is used as the initial word, and the initial word is a word directly segmented from the text search content.

[0102] Step 320, perform synonym expansion on the initial word.

[0103] In this embodiment, synonym expansion includes expansion in the same language and expansion in different languages. For example, when the text search content is Chinese content, synonym expansion in the same language can be performed on the segmented Chinese initial words, and the Chinese initial words are translated into a foreign language, and further expansion is performed on the translated foreign language, so as to obtain one or more synonyms of each initial word.

[0104] Step 330 , randomly inserting, replacing and deleting the initial word and the corresponding expanded synonyms to obtain keywords corresponding to the text search content.

[0105] In this embodiment, EDA data enhancement (Easy Data Augmentation) is performed on the obtained initial words and synonyms. Specifically, different data expansions are obtained by randomly inserting synonyms, randomly replacing synonyms, exchanging the order of phrases, and randomly deleting word units, and the data is merged and the order is disrupted to make the output after data enhancement. In the process of synonym replacement, Sohu and Sina financial vocabulary characters (Sohu_financial.char / Sgns.financial.char) are used for replacement preparation, and the threshold is calculated by the TF-IDF algorithm to avoid the replacement of effective keywords in the text retrieval content, and to prepare the data set for subsequent keyword matching\classification model training.

[0106] By performing EDA data enhancement, more diverse keyword combinations can be generated, improving the comprehensiveness of subsequent searches.

[0107] In one embodiment, a preset multi-channel recall model is called to perform a search in a search range corresponding to the target business type to obtain recall data output by each recall model, including: for one or more recall models in the multi-channel recall model, a search vector is used as an entry point of the top layer in a preset hierarchical navigable small-world algorithm, and adjacent points of the current layer are retrieved from the entry point; the adjacent points of the retrieved current layer are used to enter the next layer to continue adjacent point retrieval until the adjacent points of the last layer are found, and the data corresponding to the adjacent points of the last layer are used as recall data.

[0108] In this embodiment, one or more of the recall models uses a hierarchical navigable small world network (HNSW) algorithm model to retrieve recall data. Specifically, for the corresponding recall model, the number of layers of the graph of the HNSW algorithm model is set, and based on the corresponding search range, the nodes in each layer are set, and each node corresponds to a recall data. Based on the search vector, a node is determined as an entry point in the nodes in the top layer of the HNSW algorithm model, and the node distance between other nodes in the same layer as the entry point and the entry point is calculated, and the neighboring points of the entry point are determined based on the node distance.

[0109] Based on the determined adjacent points, enter the next layer of the adjacent points and continue to search for the corresponding adjacent points in the next layer. Search for adjacent points from the top layer to the next layer until the adjacent points in the bottom layer (i.e. the last layer) are found. The data corresponding to the adjacent points in the last layer are used as the recall data.

[0110] Optionally, the number of adjacent points determined in each layer may be one or more, and the specific number of adjacent points may be set as needed. For example, when K recall data need to be output, the total number of adjacent points in the last layer of the final output is set to K, thereby obtaining K recall data.

[0111] In one embodiment, the relevant parameters of the HNSW algorithm model integrated in each recall model are different, so that the recall data output by each recall model is not necessarily exactly the same.

[0112] By adopting the HNSW algorithm model to perform recall data retrieval, the accuracy of recall data retrieval can be improved. For example, the recall models adopted include the above-mentioned criterion label recall model, hierarchical label recall model and keyword recall model. The three recall models all adopt the HNSW algorithm model to output an appropriate amount of recall data.

[0113] In one embodiment, for each HNSW algorithm model in the recall model, the number of layers in each HNSW algorithm model and the node distribution in each layer can be preset. The node is a node determined based on the search scope corresponding to the target business type, so that the search is only performed within the search scope, and the entry point corresponding to the search vector is also within the node in the determined topmost layer.

[0114] Specifically, each node can be inserted into the corresponding layer by random insertion or probability distribution. The probability of a node being inserted into a higher layer is lower than the probability of being inserted into a lower layer. Starting from the higher layer, the k points closest to the point to be inserted are retrieved layer by layer to the lower layer search layer and a connection is established. During the retrieval process, a maximum degree is set for each node in each layer. If the degree of the node exceeds the maximum degree after the node is inserted, the connection between the node and the farthest point in the layer will be discarded. After the corresponding node is inserted, the entry point corresponding to the text vector is also determined at the top layer of the graph.

[0115] In one embodiment, the calling of the preset multi-channel recall model performs a search in the search scope corresponding to the target business type to obtain the recall data output by each recall model, including: calling the criterion label recall model to match the search vector with the criterion label to obtain the target criterion label matching the search vector, and querying a first number of recall data matching the target criterion label from the search scope; calling the hierarchical label recall model to determine the target hierarchical level to which the search vector belongs, and querying a second number of recall data matching the target hierarchical level from the search scope; calling the keyword recall model to split the search vector into keywords, and querying a third number of recall data matching the keywords from the search scope.

[0116] For the criterion label recall model, all data in the database are set with corresponding criterion labels. For the search vector, the criterion label recall model can identify the criterion label (i.e., the target criterion label) that the search vector is adapted to. The target criterion label includes one or more.

[0117] After obtaining one or more target criterion labels, all data corresponding to the target criterion labels are queried from the search range to form a first data sub-range. A hierarchical navigable small-world algorithm is used to perform a search in the first data sub-range to obtain a first amount of recalled data.

[0118] For the hierarchical label recall model, similarly, all data in the database are set with corresponding hierarchical labels. The hierarchical label recall model is called to identify the lowest hierarchical label (i.e., the target hierarchical label) corresponding to the search vector, and all data corresponding to the target hierarchical label are queried in the search range to form a second data sub-range. The hierarchical navigable small-world algorithm is used to search in the second data sub-range to obtain a second amount of recalled data.

[0119] For the keyword tag recall model, the search vector is converted into the corresponding keyword through the keyword recall model, and all search data matching the keyword are queried from the search range. A third data sub-range is formed based on all the queried search data, and a hierarchical navigable small-world algorithm is used to search within the third data sub-range to obtain a third amount of recall data.

[0120] In this embodiment, the number of recall data (i.e., the first number, the second number, and the third number) obtained by the three recall models (i.e., the criterion label recall model, the hierarchical label recall model, and the keyword recall model) can be any pre-set appropriate number, and the numbers can be the same or different. For example, the numbers are all K.

[0121] For the data sub-ranges determined by each recall model (i.e., the first data sub-range, the second data sub-range, and the third data sub-range), based on the pre-constructed layers, a node is determined as the entry point in the nodes in the top layer of the HNSW algorithm model, and the node distances between other nodes in the same layer as the entry point (other nodes are nodes corresponding to the data in the corresponding data sub-range, and nodes that do not belong to the data in the data sub-range do not participate in the node distance calculation and adjacent point selection) and the entry point are calculated, and the adjacent points of the entry point are determined based on the node distances.

[0122] Taking the determined adjacent points as the benchmark, enter the next layer of the adjacent points, and continue to search for the corresponding adjacent points in the next layer. By continuously searching for adjacent points from the top layer to the next layer, until the adjacent points in the bottom layer (i.e. the last layer) are found, the data corresponding to the adjacent points in the last layer are used as the recall data. The data corresponding to the adjacent points in each layer are all within the data sub-range. That is, the data corresponding to the adjacent points determined by the HNSW algorithm model in the criterion label recall model in each corresponding layer are all within the first data sub-range; the data corresponding to the adjacent points determined by the HNSW algorithm model in the hierarchical label recall model in each corresponding layer are all within the second data sub-range; the data corresponding to the adjacent points determined by the HNSW algorithm model in the keyword recall model in each corresponding layer are all within the third data sub-range.

[0123] By setting up three recall models including the criterion label recall model, the hierarchical label recall model and the keyword recall model, each recall model finds an appropriate amount of recall data from different dimensions, which can improve the comprehensiveness of the retrieval.

[0124] In one embodiment, the first confidence of each recalled data relative to the text retrieval content is calculated, including: splicing the text retrieval content with each recalled data, identifying the text semantic information of the spliced ​​data formed by each recalled data, and calculating the first confidence corresponding to each recalled data based on the text semantic information.

[0125] In this embodiment, the confidence of each recall data is calculated by a cross encoder, the recall data is sorted based on the calculated confidence, and the top K recall data (that is, a preset number of recall data) are used as retrieval results.

[0126] Specifically, the text search content and the recall data are spliced ​​and cross-coded to form spliced ​​data, and the first confidence of the text semantic information of the spliced ​​data relative to the text search content is identified. The higher the first confidence, the higher the accuracy of the corresponding recall data, and the more it conforms to the search results expected by the user. The electronic device can set a corresponding semantic recognition model, input the spliced ​​data into the semantic recognition model for semantic recognition, and output the corresponding first confidence to improve the accuracy of the confidence. The semantic recognition model can be any suitable recognition model, for example, it can be a large model semantic recognition model, such as a retrieval enhancement generation (RAG) model.

[0127] By using a cross-coding method to perform the first confidence calculation of the recalled data, the accuracy of the confidence calculation can be improved.

[0128] In one implementation, a multi-path recall-based information retrieval device is provided, such as Figure 4 As shown, the device comprises:

[0129] The search content acquisition module 410 is used to acquire the text search content of the user.

[0130] The business type identification module 420 is used to identify the target business type to which the text search content belongs.

[0131] The recall data acquisition module 430 is used to convert the text search content into a search vector, call a preset multi-channel recall model to search in the search range corresponding to the target business type, and obtain the recall data output by each recall model.

[0132] The retrieval result output module 440 is used to calculate the first confidence of each recalled data relative to the text retrieval content; and output a preset number of recalled data as the retrieval result of the text retrieval content based on the first confidence.

[0133] In one embodiment, Figure 5 As shown, the device also includes a retrieval type identification module 450, which is used to calculate the second confidence that the text retrieval content belongs to a rule-based retrieval; when the second confidence exceeds a preset confidence threshold, the retrieval type of the text retrieval content is determined to be a rule-based retrieval; when the second confidence is lower than the preset confidence threshold, the retrieval type is determined to be a model-based retrieval.

[0134] The search result output module 440 is further configured to query a preset number of search results matching the keywords in the text search content from the global database when the second confidence exceeds a preset confidence threshold.

[0135] The business type identification module 420 is further configured to identify the target business type to which the text retrieval content belongs when the second confidence level is lower than a preset confidence level threshold.

[0136] The recall data acquisition module 430 is also used to segment the text retrieval content to form keywords corresponding to the text retrieval content; calculate the information score of each keyword, which is used to reflect the importance of the keyword in the text retrieval content; serialize the keywords based on the information score to form serialized characters, and convert the serialized characters into retrieval vectors through a preset multi-layer encoder.

[0137] In one embodiment, the recall data acquisition module 430 is also used to perform Jieba word segmentation on the text search content to form initial words; perform synonym expansion on the initial words; and perform random insertion and replacement and deletion processing on the initial words and the corresponding expanded synonyms to obtain keywords corresponding to the text search content.

[0138] In one embodiment, the recall data acquisition module 430 is also used to retrieve the neighboring points of the current layer from the entry point for one or more of the recall models in the multi-way recall model, using the retrieval vector as the entry point of the top layer in the pre-set hierarchical navigable small-world algorithm; and to enter the next layer with the retrieved neighboring points of the current layer to continue the neighboring point retrieval until the neighboring points of the last layer are found, and the data corresponding to the neighboring points of the last layer are used as the recall data.

[0139] In one embodiment, the search result output module 440 is further used to identify text semantic information of the spliced ​​data formed by each recalled data by splicing the text search content with each recalled data, and calculate the first confidence corresponding to each recalled data based on the text semantic information.

[0140] In one embodiment, the multi-way recall model includes a criterion label recall model, a hierarchical label recall model, and a keyword recall model.

[0141] In one embodiment, a computer storage medium is provided, on which executable instructions are stored. When the instructions are executed by a processor, the processor executes the steps in the above-mentioned information retrieval method embodiments based on multi-way recall.

[0142] In one embodiment, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the information retrieval method based on multi-way recall in any of the above embodiments.

[0143] In one embodiment, an electronic device is provided, which may be a terminal or a server. Figure 6 As shown, the electronic device 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0144] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage section 608 as needed.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0146] In addition, those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present application and form different embodiments. For example, all of the above embodiments may be used in any combination. The information disclosed in this background technology section is intended only to deepen the understanding of the overall background technology of the present application and should not be regarded as an admission or in any form to imply that the information constitutes prior art already known to those skilled in the art.

Claims

1. An information retrieval method based on multi-way recall, characterized in that: The method comprises: Get the user's text search content; Calculating a second confidence level that the text search content belongs to a rule-based search; When the second confidence exceeds a preset confidence threshold, determining that the search type of the text search content is the rule-based search, querying a preset number of search results that match the keywords in the text search content from a global database; When the second confidence level is lower than a preset confidence level threshold, determining that the search type is a model-based search, and identifying the target business type to which the text search content belongs; Convert the text search content into a search vector, call a preset multi-channel recall model to search in the search range corresponding to the target business type, and obtain the recall data output by each recall model; Calculating a first confidence level of each recall data relative to the text retrieval content; Outputting a preset number of recalled data as a search result of the text search content based on the first confidence level; The multi-channel recall model includes a criterion label recall model, a hierarchical label recall model and a keyword recall model. The calling of the preset multi-channel recall model performs a search in the search range corresponding to the target business type to obtain the recall data output by each recall model, including: calling the criterion label recall model to match the search vector with the criterion label to obtain the target criterion label matching the search vector, querying a first number of recall data matching the target criterion label from the search range, calling the hierarchical label recall model to determine the target level to which the search vector belongs, querying a second number of recall data matching the target level from the search range, calling the keyword recall model to split the search vector into keywords, and querying a third number of recall data matching the keywords in the search range.

2. The information retrieval method based on multi-way recall according to claim 1 is characterized in that: The converting the text search content into a search vector comprises: Segmenting the text search content to form keywords corresponding to the text search content; Calculate the information content score of each keyword, where the information content score is used to reflect the importance of the keyword in the text search content; The keyword is serialized based on the information score to form serialized characters, and the serialized characters are converted into the retrieval vector through a preset multi-layer encoder.

3. The information retrieval method based on multi-way recall according to claim 2 is characterized in that: The segmenting of the text search content to form keywords corresponding to the text search content includes: Performing Jieba word segmentation on the text search content to form initial words; Performing synonym expansion on the initial word; The initial word and the corresponding expanded synonyms are randomly inserted and replaced and deleted to obtain keywords corresponding to the text search content.

4. The information retrieval method based on multi-way recall according to claim 1 is characterized in that: The calling of the preset multi-channel recall model to search within the search range corresponding to the target business type to obtain the recall data output by each recall model includes: For one or more of the multiple recall models, the search vector is used as an entry point of the top layer in a pre-set hierarchical navigable small-world algorithm, and a neighboring point search of the current layer is performed from the entry point; The retrieved adjacent points of the current layer are used to enter the next layer to continue the adjacent point search until the adjacent points of the last layer are found, and the data corresponding to the adjacent points of the last layer are used as the recalled data.

5. The information retrieval method based on multi-way recall according to claim 1, characterized in that: The calculating of the first confidence of each recall data relative to the text retrieval content comprises: The text search content is spliced ​​with each recalled data, the text semantic information of the spliced ​​data formed by each recalled data is identified, and the confidence corresponding to each recalled data is calculated based on the text semantic information.

6. An information retrieval device based on multi-way recall, characterized in that: The device comprises: A search content acquisition module is used to obtain the user's text search content; A search type identification module, used to calculate a second confidence level that the text search content belongs to a rule-based search, and when the second confidence level exceeds a preset confidence level threshold, determine that the search type of the text search content is the rule-based search, and when the second confidence level is lower than the preset confidence level threshold, determine that the search type is a model-based search; A business type identification module, used to identify the target business type to which the text search content belongs; A recall data acquisition module, used to convert the text search content into a search vector, call a preset multi-channel recall model to search in the search range corresponding to the target business type, and obtain the recall data output by each recall model; A retrieval result output module, used to calculate a first confidence level of each recalled data relative to the text retrieval content; and output a preset number of recalled data as retrieval results of the text retrieval content based on the first confidence level; The multi-way recall model includes a criterion label recall model, a hierarchical label recall model and a keyword recall model. The recall data acquisition module is further used to call the criterion label recall model to match the search vector with the criterion label, obtain a target criterion label matching the search vector, query a first number of recall data matching the target criterion label from the search scope, call the hierarchical label recall model to determine the target hierarchy to which the search vector belongs, query a second number of recall data matching the target hierarchy from the search scope, call the keyword recall model to split the search vector into keywords, and query a third number of recall data matching the keywords from the search scope; The search result output module is further configured to query a preset number of search results matching the keywords in the text search content from a global database when the second confidence exceeds a preset confidence threshold.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores executable instructions, and when the executable instructions are executed by a processor, the processor executes the information retrieval method based on multi-way recall as described in any one of claims 1 to 5.

8. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to execute the information retrieval method based on multi-way recall as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • A search recall method based on multi-path recall

    CN115809312B

  • Multi-path recall retrieval enhancement generation method and device based on real-time data index

    CN118820400A

  • Data matching method and device and electronic equipment

    CN114153962A

  • Information retrieval method and device, information recommendation method and device and electronic equipment

    CN116821440A