Data retrieval method, electronic equipment and computer readable storage medium
By using target word segmentation pairs in data retrieval to match the identification information in the information database, the problem of poor search results in the prior art is solved, and more accurate search results are achieved.
Patent Information
- Application Number
- CN202411911261.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art has the problem of poor search results in data retrieval, mainly because the keyword extraction method cannot accurately reflect the user's search intention, resulting in the search results deviating from the user's original intention.
By obtaining the target word segmentation pairs of the search text, reflecting their contextual relationships, and matching these word segmentation pairs in the information library to obtain identification information, thus obtaining more accurate search results.
This method can more accurately reflect the user's search intention, improve the accuracy of the search results, and solve the problem of poor search results.
Smart Images

Figure CN119917638A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of Internet, and specifically relates to a data retrieval method, an electronic device and a computer-readable storage medium. Background Art
[0002] With the advent of the big data era, the amount of information on the Internet is growing rapidly. Users need to locate the required content from massive data through data retrieval. Therefore, how to optimize data retrieval methods is an issue worthy of attention.
[0003] The related technology usually extracts keywords from the search text input by the user, and obtains the search results of the search text by searching the keywords. However, the search results obtained in this way are easy to deviate from the user's original intention, and there is a problem of poor data retrieval results. Summary of the invention
[0004] The embodiments of the present application provide a data retrieval method, an electronic device, and a computer-readable storage medium, which can solve the problem of poor data retrieval results existing in the related art.
[0005] In a first aspect, an embodiment of the present application provides a data retrieval method, the method comprising: Acquire a target word segmentation pair of a first search text, wherein the target word segmentation pair is used to reflect the contextual relationship of the first search text; Determining first identification information matching the target word segmentation pair in a first information database, wherein the first information database stores a mapping relationship between word segmentation pairs and identification information; Based on the first identification information, a first search result is obtained.
[0006] In a second aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0007] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed, the steps of the method described in the first aspect are implemented.
[0008] In a fourth aspect, an embodiment of the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect.
[0009] At least one of the above technical solutions provided in the embodiments of the present application can achieve the following technical effects: In an embodiment of the present application, a target word segmentation pair of a first search text is obtained, and the target word segmentation pair is used to reflect the contextual relationship of the first search text; first identification information matching the target word segmentation pair is determined in a first information base, and the first information base stores the mapping relationship between the word segmentation pair and the identification information; based on the first identification information, a first search result is obtained. In this way, in the process of searching the first search text, the first identification information matching the target word segmentation pair can be determined in the first information base, and then the first search result is obtained through the first identification information. Compared with the method of extracting keywords for retrieval in the related technology, since the target word segmentation pair can reflect the contextual relationship of the first search text rather than isolated keywords, the first search result obtained through the first identification information matching the target word segmentation pair is more in line with the original intention of the first search text, and the first search result obtained is more accurate, which solves the problem of poor data retrieval results in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0011] Figure 1 is a flow chart of a data retrieval method provided in an embodiment of the present application; Figure 2 is a flow chart of another data retrieval method provided in an embodiment of the present application; Figure 3 It is a schematic diagram of a retrieval processing method provided in an embodiment of the present application; Figure 4 is a flow chart of another data retrieval method provided in an embodiment of the present application; Figure 5 is a schematic diagram of a cache mechanism provided in an embodiment of the present application; Figure 6 is a flow chart of another data retrieval method provided in an embodiment of the present application; Figure 7 is a schematic diagram of a preprocessing process of a text to be stored provided in an embodiment of the present application; Figure 8 It is a specific flow chart of a data retrieval method provided in an embodiment of the present application; Fig. 9 It is a structural block diagram of a data retrieval device provided in an embodiment of the present application; Fig.10 It is a structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0012] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0013] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0014] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0015] The data retrieval method provided in the embodiment of the present application is applied to data retrieval technology, for example, a user quickly locates the required content in a large amount of information by inputting keywords, and an enterprise uses data retrieval technology to find specific business data in an internal database, etc. Specifically, in the process of searching the first search text, the first identification information matching the target word segmentation pair can be determined in the first information library, and then the first search result is obtained through the first identification information.
[0016] The data retrieval method provided in the embodiment of the present application can be executed by a target device, wherein the target device can be one electronic device or multiple electronic devices. In other words, the data retrieval method provided in the embodiment of the present application can be executed by an electronic device, wherein the electronic device can be a terminal device such as a desktop computer, a laptop computer, a mobile phone, a tablet, etc., or a server, such as an independent physical server, a server cluster composed of multiple servers, and a cloud server capable of cloud computing. In the case where the data retrieval method provided in the embodiment of the present application is executed by multiple electronic devices, these multiple electronic devices can form a service cluster, which cooperate with each other to complete each step.
[0017] The data retrieval method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0018] See also Figure 1 , Figure 1 is a flow chart of a data retrieval method provided in an embodiment of the present application. Figure 1 As shown, the method comprises the following steps: Step 110: Obtain a target word segmentation pair of the first search text, wherein the target word segmentation pair is used to reflect the contextual relationship of the first search text.
[0019] In the embodiment of the present application, the first search text may include a term to be searched or a text to be searched, and the first search text may be a text input by a user, such as "The weather is very good today". The target segmented word pair is a segmented word pair composed of segmented words in the first search text, and two or more segmented words with adjacent relationships in the first search text may be spliced to obtain the target segmented word pair, and the contextual relationship of the first search text is reflected through the adjacent relationship between the segmented words.
[0020] Specifically, in one embodiment of the present application, obtaining the target word segmentation pair of the first search text in step 110 includes: performing word segmentation processing on the first search text to obtain multiple word segments; determining a target word segmentation pair from the multiple word segments, and the target word segmentation pair includes two adjacent word segments in the first search text.
[0021] In the embodiment of the present application, the first search text can be segmented by a word segmenter (such as a Tokenizer) to divide the first search text into multiple word segments, and a search word segmentation sequence can be obtained. . Wherein, m is the number of segmented words of the first search text, and the segmented words used include at least one of rule-based segmented words, statistical-based segmented words, and deep learning-based segmented words. For example, if the first search text is "The weather is very good today", three segmented words "today", "weather" and "very good" can be obtained.
[0022] After obtaining multiple segmented words, two adjacent segmented words can be obtained from the multiple segmented words according to the adjacent relationship of the multiple segmented words, and the two adjacent segmented words are spliced to obtain the target segmented word pair. Here, a connector such as "+" can be used to splice the two adjacent participles, or the two adjacent participles can be directly spliced. For example, for the three participles "today", "weather" and "very good" in the above example, two participle pairs "today's weather" and "the weather is very good" can be obtained.
[0023] Step 120: Determine first identification information matching the target word segmentation pair in a first information database, wherein the first information database stores a mapping relationship between word segmentation pairs and identification information.
[0024] In the embodiment of the present application, the first information base is used to store the mapping relationship between the word segmentation pair and the identification information, the identification information is used to identify the text containing the word segmentation pair, and the first information base can use a data structure such as a data table or a dictionary. The word segmentation pair matching the target word segmentation pair can be queried in the first information base, and then the first identification information is obtained through the mapping relationship between the word segmentation pair and the identification information stored in the first information base.
[0025] Step 130: Obtain a first search result based on the first identification information.
[0026] In an embodiment of the present application, a target text matching the first identification information may be determined in the second information base, and the target text may be determined as a first search result. The target text may be in the form of a document, the first identification information may be an identification of the target text, the second information base is used to store text carrying the identification information, and the first search result may be a search content related to the first search text. After obtaining the first search result, the first search result may be returned to the user.
[0027] In order to ensure data security during storage and transmission, the second information database may store encrypted text carrying identification information. Specifically, in step 130, obtaining a first search result based on the first identification information includes: determining an encrypted text matching the first identification information from the second information database, decrypting the encrypted text, and obtaining the first search result.
[0028] In an embodiment of the present application, based on the first identification information, an encrypted text matching the first identification information can be queried from the second information base. Then, the encrypted text can be decrypted to obtain a first search result. Among them, the first search result can be a text obtained by decrypting the encrypted text. The first identification information may include one or more text identifiers. In the case where the first identification information includes multiple text identifiers, multiple encrypted texts corresponding to the multiple text identifiers can be determined in the second information base. The first search result can be multiple texts obtained by decrypting the multiple encrypted texts, and the multiple texts are in original readable form.
[0029] In one embodiment of the present application, the data retrieval method, in addition to steps 110 to 130, further includes: obtaining a single word segment of a second search text; determining second identification information matching the single word segment in a third information database, wherein the third information database stores a mapping relationship between the word segment and the identification information; and obtaining a third search result based on the second identification information.
[0030] In an embodiment of the present application, since the search text may only contain one participle, there may not necessarily be a participle pair. Therefore, an embodiment of the present application provides a data retrieval method for a single participle. Specifically, the participle result of the second search text may only include a single participle, and the third information base may use a data structure such as a data table or a dictionary. The second identification information that matches the single participle can be determined in the third information base, and the second identification information may include identification information that matches the single participle, which may be, for example, a set of text identifiers. After obtaining the second identification information, the text that matches the text identifier contained in the second identification information can be determined, and the text that matches the second identification information can be used as the third search result.
[0031] In an embodiment of the present application, a target word segmentation pair of a first search text is obtained, and the target word segmentation pair is used to reflect the contextual relationship of the first search text; first identification information matching the target word segmentation pair is determined in a first information base, and the first information base stores the mapping relationship between the word segmentation pair and the identification information; based on the first identification information, a first search result is obtained. In this way, in the process of searching the first search text, the first identification information matching the target word segmentation pair can be determined in the first information base, and then the first search result is obtained through the first identification information. Compared with the method of extracting keywords for retrieval in the related technology, since the target word segmentation pair can reflect the contextual relationship of the first search text rather than isolated keywords, the first search result obtained through the first identification information matching the target word segmentation pair is more in line with the original intention of the first search text, and the first search result obtained is more accurate, which solves the problem of poor data retrieval results in the related technology.
[0032] See also Figure 2 , Figure 2 FIG. 1 is a flow chart of another data retrieval method provided in an embodiment of the present application. Figure 2 As shown, the method comprises the following steps: Step 210: Obtain target word segmentation pairs of the first search text, wherein the target word segmentation pairs are used to reflect the contextual relationship of the first search text, and the target word segmentation pairs include N word segmentation pairs sequentially sorted in the first search text.
[0033] In the embodiment of the present application, after obtaining N word segmentation pairs of the first search text, the N word segmentation pairs can be sorted according to the order of the N word segmentation pairs in the first search text. For example, if the first search text is "The weather is really good today", three word segmentation pairs "Today's weather", "The weather is really" and "Really good" can be obtained. According to the order in the first search text, three word segmentation pairs can be obtained: "Today's weather", "The weather is really" and "Really good".
[0034] Step 220: Select M word segment pairs from the N word segment pairs in order, where M is an integer less than or equal to N, and N is a positive integer.
[0035] In the embodiment of the present application, M consecutive word pairs can be selected from the N word pairs, where M is an integer less than or equal to N. For example, M can be equal to N. For the three word pairs in the above example, three word pairs "today's weather", "the weather is really" and "really good" can be selected. Alternatively, M can be less than N. For the three word pairs in the above example, two word pairs "the weather is really" and "really good" can be selected.
[0036] Step 230: For each of the M word segmentation pairs, determine a text identifier set corresponding to the word segmentation pair in a first information database, wherein the first information database stores a mapping relationship between the word segmentation pairs and identifier information.
[0037] In an embodiment of the present application, for the ith word segmentation pair among the M word segmentation pairs, i is a positive integer less than or equal to M, and identification information corresponding to the ith word segmentation pair can be determined in the first information base, and the identification information can be a text identifier of a document, such as a document ID. Specifically, the ith word segmentation pair can be accurately matched in the first information base to obtain the identification information corresponding to the ith word segmentation pair. The identification information corresponding to the ith word segmentation pair may include one identifier or multiple identifiers, which can constitute a text identifier set of the ith word segmentation pair. For example, for word segmentation pair A, the text identifier set {document ID1, document ID2, document ID3}" corresponding to word segmentation pair A can be determined in the first information base.
[0038] Step 240: Based on the M word pairs, obtain M text identifier sets.
[0039] In an embodiment of the present application, for M word segmentation pairs, M text identification sets corresponding to the M word segmentation pairs can be determined through the first information base, and the M text identification sets correspond one-to-one to the M word segmentation pairs. For example, for three word segmentation pairs (word segmentation pair A, word segmentation pair B, and word segmentation pair C), the text identification set of word segmentation pair A is {document ID1, document ID2, document ID3}, the text identification set of word segmentation pair B is {document ID2, document ID3, document ID4}, and the text identification set of word segmentation pair C is {document ID2, document ID3}, and three text identification sets corresponding one-to-one to the three word segmentation pairs can be obtained.
[0040] Step 250: Obtain the intersection of the M text identification sets as the first identification result.
[0041] In the embodiment of the present application, the intersection of the M text identifier sets, that is, the common text identifiers in the M text identifier sets, can be obtained, and the obtained common text identifiers can be used as the first identification result. Taking the three text identifier sets in the above example as an example, the first identification result can be {document ID2, document ID3}.
[0042] Step 260: Based on the first identification result, a common text identification is obtained, and the common text identification is used as the first identification information.
[0043] In an embodiment of the present application, the first identification result may be directly used as the shared text identification. Alternatively, the shared text identification may be determined from the first identification result based on the first identification result. After obtaining the shared text identification, the shared text identification may be used as the first identification information, and specific reference may be made to the following embodiments.
[0044] Exemplarily, in one embodiment of the present application, M=N-1, the M word pairs are the first N-1 word pairs of the N word pairs, and N is an integer greater than 1. The step 260 of obtaining a common text identifier based on the first identifier result includes: performing prefix matching on the Nth word pair of the N word pairs to obtain a first text identifier set; obtaining the intersection of the first identifier result and the first text identifier set as the common text identifier.
[0045] In the embodiment of the present application, N is an integer greater than 1, M=N-1, and the first N-1 word segmentation pairs of the N word segmentation pairs can be used as M word segmentation pairs, and the M word segmentation pairs are accurately matched in the first information base to obtain M text identification sets corresponding to the M word segmentation pairs one by one, and the intersection of the M text identification sets is the first identification result. For the Nth word segmentation pair of the N word segmentation pairs, the Nth word segmentation pair is prefix matched in the first information base to obtain 1 first text identification set. In other words, the prefix of the Nth word segmentation pair can be accurately matched in the first information base, and the prefix of the Nth word segmentation pair can be, for example, the first word segmentation in the Nth word segmentation pair, or can be the prefix part of the first word segmentation in the Nth word segmentation pair.
[0046] After obtaining the first identification result and the first text identification set, the intersection of the first identification result and the first text identification set can be determined, that is, the text identifications shared by the first identification result and the first text identification set. If the intersection is empty, it can be determined that the search result for the first search text is empty, that is, no content related to the first search text can be matched.
[0047] Exemplarily, in one embodiment of the present application, M=N-1, the M word pairs are the last N-1 word pairs of the N word pairs, and N is an integer greater than 1. The step 260 of obtaining a common text identifier based on the first identifier result includes: performing suffix matching on the first word pair of the N word pairs to obtain a second text identifier set; obtaining the intersection of the first identifier result and the second text identifier set as the common text identifier.
[0048] In an embodiment of the present application, N is an integer greater than 1, and the last N-1 word segmentation pairs can be selected from the N word segmentation pairs as M word segmentation pairs, and the M word segmentation pairs can be accurately matched in the first information base to obtain M text identification sets corresponding to the M word segmentation pairs one by one, and the intersection of the M text identification sets is the first identification result. For the first word segmentation pair among the N word segmentation pairs, the first word segmentation pair can be suffix matched in the first information base to obtain a second text identification set. In other words, the suffix of the first word segmentation pair can be accurately matched in the first information base, and the suffix of the first word segmentation pair can be, for example, the second word segmentation in the first word segmentation pair, or the suffix part of the second word segmentation in the first word segmentation pair.
[0049] After obtaining the first identification result and the second text identification set, the intersection of the first identification result and the second text identification set can be determined, that is, the text identifications shared by the first identification result and the second text identification set. If the intersection is empty, it can be determined that the search result for the first search text is empty, that is, no content related to the first search text can be matched.
[0050] Exemplarily, in one embodiment of the present application, M=N-2, the M word pairs are N-2 word pairs in the middle of the N word pairs, and N is an integer greater than 2. The step 260 of obtaining a common text identifier based on the first identifier result includes: performing suffix matching on the first word pair of the N word pairs to obtain a first identifier set; performing prefix matching on the Nth word pair of the N word pairs to obtain a second identifier set; obtaining the intersection of the first identifier result, the first identifier set, and the second identifier set as the common text identifier.
[0051] In the embodiment of the present application, N is an integer greater than 2, and N-2 word segmentation pairs (i.e., from the 2nd word segmentation pair to the N-1th word segmentation pair) among the N word segmentation pairs can be selected as M word segmentation pairs, and the M word segmentation pairs are precisely matched in the first information library to obtain M text identification sets corresponding to the M word segmentation pairs one by one, and the intersection of the M text identification sets is the first identification result. For the first word segmentation pair among the N word segmentation pairs, a suffix match can be performed on the first word segmentation pair in the first information library to obtain a first identification set. For the Nth word segmentation pair among the N word segmentation pairs, a prefix match can be performed on the Nth word segmentation pair in the first information library to obtain a second identification set.
[0052] After obtaining the first identification result, the first identification set and the second identification set, the intersection of the first identification result, the first identification set and the second identification set, that is, the text identifications shared by the first identification result, the first identification set and the second identification set, can be determined. In the case where the intersection is empty, it can be determined that the search result for the first search text is empty, that is, no content related to the first search text can be matched.
[0053] Taking the three word pairs "Today's weather", "The weather is really" and "Really good" in the above example as examples, N=3, M=1, the second word pair "The weather is really" can be selected as the M word pairs, and the second word pair "The weather is really good" can be accurately matched to obtain the first identification result. The first word pair "Today's weather" can be suffix matched to obtain the first identification set. The third word pair "Really good" can be prefix matched to obtain the second identification result.
[0054] In the embodiment of the present application, prefix matching and suffix matching are introduced, which can match richer search results, handle more complex searches, and implement variant searches of the search content input by the user, thereby improving the flexibility of the search.
[0055] In one embodiment of the present application, the data retrieval method, in addition to steps 210 to 270, also includes: obtaining a single word segment of the second search text; determining second identification information matching the single word segment in a third information database, wherein the third information database stores a mapping relationship between the word segment and the identification information; and obtaining a third search result based on the second identification information.
[0056] Please refer to Figure 3 , Figure 3 Schematic diagram of a search processing method provided by an embodiment of the present application. Figure 3 As shown, when the search process starts, the cache can be queried first to determine whether the first search text is in the cache library. If the first search text exists in the cache library, the search result of the first search text can be obtained from the cache library. If the first search text does not exist in the cache library, the first search text can be segmented, and the search method is determined according to the number of segmented words obtained by the segmentation process. If the number of segmented words is equal to one, the target segmented word pair can be obtained according to the segmentation result, and the third information library can be queried using a single segmented word search method. The first search text is fuzzy matched in the third information library to obtain a text identifier that matches the first search text.
[0057] When the number of word segments is greater than one, the first information base can be queried using a multi-word search method. For the N word segmentation pairs contained in the target word segmentation pairs, the first N-1 word segmentation pairs of the N word segmentation pairs can be accurately matched, and the last word segmentation pair can be prefix matched to obtain a set of N text identifiers obtained by matching the N word segmentation pairs. The N text identifier sets are subjected to an intersection operation to obtain a text identifier that matches the first search text. It should be noted that Figure 3 The method of performing exact matching on the first N-1 segmented word pairs among the N segmented word pairs and performing prefix matching on the Nth segmented word pair among the N segmented word pairs is only an example. It is also possible to perform exact matching on the last N-1 segmented word pairs among the N segmented word pairs and perform suffix matching on the first segmented word pair among the N segmented word pairs. Alternatively, suffix matching is performed on the first segmented word pair among the N segmented word pairs, exact matching is performed on the N-2 segmented word pairs in the middle of the N segmented word pairs, and prefix matching is performed on the last segmented word pair among the N segmented word pairs.
[0058] Step 270: Obtain a first search result based on the first identification information.
[0059] In an embodiment of the present application, after obtaining the common text identifier, the common text identifier can be used as the first identification information, the target text matching the first identification information can be determined in the second information library, and the target text can be used as the first search result corresponding to the first search text.
[0060] In an embodiment of the present application, M word segmentation pairs may be selected in order from N word segmentation pairs, and the M word segmentation pairs may be matched to obtain M text identification sets, and the intersection of the M text identification sets may be determined as a first identification result, and the first identification information may be obtained based on the first identification result. Since the M word segmentation pairs are selected in order from the N word segmentation pairs, and the N word segmentation pairs are sorted according to the order in the first search text, the M word segmentation pairs may be used to represent a coherent text in the first search text, which is consistent with the original intention of the first search text, so that the first search result obtained is more consistent with the original intention of the user's search.
[0061] See also Figure 4 , Figure 4 FIG. 1 is a flow chart of another data retrieval method provided in an embodiment of the present application. Figure 4 As shown, the method comprises the following steps: Step 410: When the first search text does not exist in the cache library, obtain a target word segmentation pair of the first search text, the cache library includes a mapping relationship between historical search texts and identification information of the historical search texts, and the target word segmentation pair is used to reflect the contextual relationship of the first search text.
[0062] In an embodiment of the present application, the cache library can be used to cache historical search texts and identification information corresponding to the historical search texts, and the historical search texts can be texts searched before searching the first search text. After searching the historical search texts, the identification information of the historical search texts can be obtained. Since the historical search texts may still be queried in the future, the historical search texts and the identification information corresponding to the historical search texts can be cached, so that the identification information can be directly obtained from the cache in the case of repeated queries, thereby improving the efficiency of the query.
[0063] In the case that the first search text does not exist in the cache library, it can be determined that the identification information corresponding to the first search text is not cached in the cache library. At this time, the first search text can be further searched. The first search text is segmented, and based on the segmentation result of the first search text, the target segmentation pair of the first search text is obtained.
[0064] In one embodiment of the present application, when there is a first search text in the cache library, first identification information matching the first search text is obtained from the cache library, wherein the first identification information is used to obtain a second search result for the first search text.
[0065] In the embodiment of the present application, if the first search text exists in the historical search texts stored in the cache library, it can be said that the first search text has been searched in the historical time period, and the identification information of the first search text exists in the cache library. Therefore, the first identification information matching the first search text can be directly obtained from the cache library without subsequent search, and the second search result of the first search text can be directly obtained based on the first identification information, that is, the text matching the first identification information can be directly determined in the second information library, and the matching text is used as the second search result.
[0066] In this way, when the first search text hits the cache library, the first identification information matching the first search text can be directly obtained from the cache library without subsequent search, which can achieve fast return of search results and avoid repeated searches, further improving search efficiency.
[0067] In an embodiment of the present application, the initialization process of the cache library is as follows: a cache data structure, such as a hash table, can be created in the cache library. The cache data structure can be used to store relevant information of the historical search text. The relevant information of the historical search text may include the specific content of the historical search text, the timestamp of the most recent search of the historical search text, the number of times the historical search text was searched, and the number of text identifiers matched by the historical search text. Since the identification information matched by the historical search text usually has a lot of content, it is not convenient to store the identification information of the historical search text in the above-mentioned cache data structure. The mapping relationship between the historical search text and the identification information matched by the historical search text can be stored separately. Specifically, a data table can be created in the cache library, and the data table is used to store the mapping relationship between the historical search text and the identification information matched by the historical search text.
[0068] In one embodiment of the present application, the cache library may be regularly cleaned, performance monitored, and logged. Specifically, the effective time period of the cache item may be preset, and if the time period between the last timestamp of the cache item being retrieved and the current timestamp exceeds the effective time period, the cache item may be deleted from the cache library. The validity of the cache items in the cache library may be regularly checked, and outdated or no longer used cache items may be removed in a timely manner.
[0069] In addition, the hit rate and response time of cache items in the cache library can be monitored in real time to evaluate the effectiveness of the cache library's cache strategy. The work log of cache item update, invalidation, and replacement events in the cache library can be recorded in a timely manner, which can be used for subsequent troubleshooting and performance analysis. In this way, a query cache system with a clear structure, efficient management, and fast response can be built. The cache content in the cache library can be dynamically adjusted according to the retrieval status of historical retrieval texts and document changes, thereby optimizing retrieval performance and resource utilization.
[0070] Step 420: Determine first identification information matching the target word segmentation pair in a first information database, wherein the first information database stores a mapping relationship between word segmentation pairs and identification information.
[0071] In one embodiment of the present application, after determining the first identification information that matches the target word segmentation pair in the first information library, if there is an available cache item in the cache item, the mapping relationship between the first search text and the first identification information can be stored in the cache library. If there is no available cache item in the cache library, the cache probability of the mapping relationship and the cache probability of the original cache item in the cache library can be obtained. If the cache probability of the mapping relationship is greater than the cache probability of the target cache item in the cache library, the mapping relationship is stored in the cache library to replace the target cache item, and the target cache item is the original cache item with the smallest cache probability in the cache library.
[0072] In an embodiment of the present application, after determining the first identification information that matches the target word segmentation pair, it may be considered to store the first identification information that matches the first search text with the first search text in the cache library, so that in the case of subsequent repeated retrieval of the first search text, the first identification information that matches the first search text can be directly obtained from the cache library. However, the capacity of the cache library is limited, and it is possible to first determine whether there are any available cache items in the cache library. Specifically, the cache number of the cache library can be preset to 100, and the quantitative relationship between the number of cache items in the cache library and the preset cache number can be used to determine whether there are any available cache items in the cache library. In the case where the number of cache items in the cache library is less than the cache number of the cache library, it can be determined that there are available cache items in the cache library, and the identification information that matches the first search text with the first search text can be directly stored in the cache library as a cache item.
[0073] When the number of cache items in the cache library is equal to the number of cache items in the cache library, it can be determined that there are no available cache items in the cache library, and the cache probability of the mapping relationship between the first search text and the first identification information and the cache probability of the original cache items in the cache library can be obtained to determine whether it is necessary to select an original cache item from the cache library for deletion and replace the original cache item with the mapping relationship between the first search text and the first identification information. The specific replacement strategy can be determined based on the cache probability of the cache item, and the cache probability of the cache item can be used to indicate the probability of the cache item being stored in the cache library or continuing to be stored in the cache library.
[0074] In an embodiment of the present application, through the above-mentioned replacement strategy, cache items with a higher cache probability can be stored in the cache library. Here, the high cache probability of the cache item can be used to indicate that the cache item is frequently retrieved, the cache item has been retrieved recently, or the cache item matches richer identification information, etc. The method for determining the cache item can be determined according to the specific scenario and strategy.
[0075] According to the cache probability of the original cache items in the cache library, the original cache item with the lowest cache probability can be used as the target cache item. If the cache probability of the mapping relationship between the first search text and the first identification information is greater than the cache probability of the target cache item, the target cache item is deleted from the cache library, and the mapping relationship between the first search text and the first identification information is stored in the cache library. If the cache probability of the mapping relationship between the first search text and the first identification information is less than or equal to the cache probability of the target cache item, no replacement is required. In the process of selecting a target cache item from the original cache items in the cache library, if there are multiple original cache items in the cache library with the same cache probability, and all of them are cache items with the lowest cache probability, then one cache item can be selected from multiple cache items with the same cache probability as the target cache item, or the target cache item can be determined according to other strategies (such as least recently used).
[0076] In one embodiment of the present application, obtaining the cache probability of the mapping relationship and the cache probability of the original cache item in the cache library includes: obtaining the current timestamp, the first parameter of the mapping relationship and the second parameter of the original cache item in the cache library, the first parameter including: at least one of the first identification number contained in the first identification information, the first search number of the first search text and the first timestamp when the first search text was last searched; the second parameter including: at least one of the second identification number contained in the original cache item, the second search number of the original cache item and the second timestamp when the original cache item was last searched; based on the first parameter and the current timestamp, determining the cache probability of the mapping relationship; based on the second parameter and the current timestamp, determining the cache probability of the original cache item in the cache library.
[0077] In an embodiment of the present application, in the process of obtaining the cache probability of the mapping relationship between the first search text and the first identification information and the cache probability of the original cache item in the cache library, the cache probability of the mapping relationship and the cache probability of the original cache item can be determined in the same manner. For the cache probability of the mapping relationship between the first search text and the first identification information, a first parameter of the mapping relationship can be obtained, and the first parameter includes at least one of the number of first identifications contained in the first identification information, the first search times of the first search text, and the first timestamp of the last search of the first search text. The cache probability of the mapping relationship can be determined based on the current timestamp and the first parameter. Among them, since the first search text is not in the cache library, it can be determined that the first search times of the first search text is 1, and the first timestamp of the last search of the first search text is the current timestamp.
[0078] Correspondingly, for the i-th original cache item among the Q original cache items in the cache library, Q is a positive integer, i is a positive integer less than or equal to Q, the second parameter of the i-th original cache item can be obtained, and the second parameter includes at least one of the number of second identifiers contained in the i-th original cache item, the second search number of the i-th original cache item, and the second timestamp of the last search of the i-th original cache item, and the second parameter can be read from the cache data structure in the cache library. The cache probability of the i-th original cache item can be determined based on the current timestamp and the second parameter of the i-th original cache item. Based on the Q original cache items, Q cache probabilities corresponding to the Q original cache items can be determined.
[0079] Specifically, the first parameter may include the first search number of the first search text and the first timestamp of the last search of the first search text, and the second parameter may include the second search number of the original cache item and the second timestamp of the last search of the original cache item. Based on the current timestamp, the first search number of the first search text and the first timestamp of the last search of the first search text, the cache probability of the mapping relationship may be determined by the following formula: ; in, is the cache probability of the mapping relationship, k is a preset weight coefficient, is the first search number of the first search text, is the first timestamp of the last search of the first search text, e is a natural constant, is the preset attenuation coefficient, and CT is the current timestamp.
[0080] The cache probability of the i-th original cache item can be determined by the following formula based on the current timestamp, the second search number of the i-th original cache item, and the second timestamp when the i-th original cache item was last retrieved: ; in, is the cache probability of the i-th original cache item, k is the preset weight coefficient, is the second retrieval number of the i-th original cache item, is the second timestamp of the last retrieval of the i-th original cache item, e is a natural constant, is the preset attenuation coefficient, and CT is the current timestamp.
[0081] In one embodiment of the present application, the first parameter includes: the first identification number contained in the first identification information, the first search times of the first search text, and the first timestamp of the last search of the first search text, and the second parameter includes: the second identification number contained in the original cache item, the second search times of the original cache item, and the second timestamp of the last search of the original cache item. Based on the current timestamp and the first parameter of the mapping relationship, the cache probability of the mapping relationship can be determined by the following formula: ; For the i-th original cache item among the Q original cache items, the cache probability of the i-th original cache item may be determined by the following formula based on the current timestamp and the second parameter of the i-th original cache item: ; in, is the cache probability of the mapping relationship, is the cache probability of the i-th original cache item, k is a preset weight coefficient, T is a preset constant, is the number of first identifications included in the first identification information, is the number of second identifiers contained in the i-th original cache item, is the first search number of the first search text, is the second retrieval number of the i-th original cache item, is the first timestamp of the last time the first search text was searched, is the second timestamp of the last retrieval of the i-th original cache item, e is a natural constant, is the preset attenuation coefficient, and CT is the current timestamp.
[0082] In the embodiment of the present application, it can be known from the above formula that the cache probability can be composed of two distributions, the first part is: , the second part is . The first part is used to represent the cache probability related to the number of text identifiers (DF). log(DF) is used to reflect the frequency of occurrence of the retrieved text in the document. Taking the logarithm of DF can reduce the direct numerical difference in the case of high DF values, making the calculation of the cache probability more balanced and reducing the scale sensitivity. T is a preset constant that can be used to adjust the impact range of the number of text identifiers; different retrieval systems may have different sensitivities to the number of text identifiers. For example, some retrieval systems may want the cache library to be more inclined to cache retrieval texts with a high number of text identifiers, and some retrieval systems may want to achieve more even processing of retrieval texts with different numbers of text identifiers. The value of T can be adjusted according to these characteristics.
[0083] The second part is related to the number of times the text has been retrieved and the timestamp of the last time it was retrieved. is an exponential decay function that can be used to reduce the impact of the last retrieved timestamp, and is used to reflect the impact of the last retrieved timestamp on the cache probability of the cached item. k is a preset weight coefficient, and the value of k is between 0 and 1, which is used to balance the importance of the number of text identifiers and the retrieval frequency (the number of times the retrieved text is retrieved and the last retrieved timestamp). When k is close to 0, the determination of the cache probability tends to consider the number of text identifiers more, and when k is close to 1, the determination of the cache probability tends to consider the retrieval frequency (the number of times the retrieved text is retrieved and the last retrieved timestamp).
[0084] in, is a preset attenuation coefficient, A positive number used to control the impact of the last retrieved timestamp on the cache probability. A value of 0 means that the importance of the last retrieved timestamp on the cache probability decreases faster.
[0085] In one embodiment of the present application, when there is a first search text in the cache library, after obtaining the first identification information that matches the first search text from the cache library, the first search text is used as an original cache item in the cache library, and after retrieving the original cache item, the second parameter of the original cache item can be updated in time. Specifically, the second timestamp of the last retrieval of the original cache item can be determined as the current timestamp, and the second retrieval number of the original cache item is increased by one.
[0086] In one embodiment of the present application, when there is a text to be stored in the second information library, the historical search text can be obtained from the cache library; when the text to be stored contains the historical search text, the mapping relationship between the historical search text and the identification information of the text to be stored is written into the cache library. At the same time, the second parameter of the historical search text can be synchronously updated, the cache probability of the historical search text can be re-determined, and the cache item replacement strategy can be executed as needed.
[0087] Step 430: Obtain a first search result based on the first identification information.
[0088] See also Figure 5 , Figure 5 Schematic diagram of a cache mechanism provided by an embodiment of the present application. Figure 5 As shown, the cache library can be initialized before data retrieval. When a new search (the search text is not in the cache library) occurs, the identification information matching the new search is determined through the first information library or the third information library, and the cache library can be updated at this time. The cache probability of the new search and the cache probability of the original cache items in the cache library can be determined, and the replacement strategy is executed to update the cache items in the cache library. The specific replacement strategy is as follows: if the cache library is not full, the new search is directly cached into the cache library; if the cache library is full, determine whether to replace it through the cache probability of the new search and the cache probability of the original cache items in the cache library. In the case of replacement, the original cache item with the lowest cache probability is determined as the replacement item.
[0089] In addition, if Figure 5 As shown, when the new text is stored in the second information library, the original cache items in the cache library are checked, the identification information of the original cache items in the cache library is updated in time, and the second parameters of the original cache items (identification number and last retrieval timestamp) are synchronously updated, and cache invalidation is processed.
[0090] In an embodiment of the present application, a cache mechanism is introduced in data retrieval, and the user's historical search texts and identification information matching the historical search texts can be cached. When the first search text hits the cache, there is no need for subsequent search, and the first identification information matching the first search text can be directly obtained from the cache library, which can achieve fast retrieval and return retrieval results, thereby improving retrieval efficiency.
[0091] See also Figure 6 , Figure 6 FIG. 1 is a flow chart of another data retrieval method provided in an embodiment of the present application. Figure 6 As shown, the method comprises the following steps: Step 610: Based on the text to be stored, a designated word pair is obtained, where the designated word pair includes two adjacent word pairs in the text to be stored.
[0092] In the embodiment of the present application, the text to be stored may include text that already exists in the second information database, or new text that has not yet been stored in the second information database. The text to be stored here may be unencrypted text (such as a document). For the text to be stored, the text to be stored may be segmented to obtain a segmentation sequence. , n is the number of segmentations, and the mapping relationship between each segmentation in the segmentation sequence and the text to be stored can be stored in the third information database, and the third information database can use a dictionary data structure (TermDict) to store the mapping relationship between the segmentation and the text identifier.
[0093] Specifically, for the i-th word segment in the word segmentation sequence, it can be first determined whether the i-th word segment exists in the third information database. In the case that the i-th word segment exists in the third information database, a text identifier set matching the i-th word segment can be obtained from the third information database. In the case that the text identifier of the text to be stored exists in the text identifier set, no other processing is performed; in the case that the text identifier of the text to be stored does not exist in the text identifier set, the text identifier of the text to be stored can be added to the text identifier set, and the text identifier set containing the text identifier of the text to be stored can be reassigned to the text identifier set matching the i-th word segment. In the case that the i-th word segment does not exist in the third information database, a new mapping relationship can be created in the third information database to store the mapping relationship between the i-th word segment and the text to be stored.
[0094] When there is new text to be stored in the second information base, the new text can be processed in the above manner to ensure that the third information base (TermDict) can reflect the mapping relationship between the word segmentation and text identifier of the latest text.
[0095] In the embodiment of the present application, the segmented words in the segmented word sequence are sorted according to the order in the text to be stored. Afterwards, a specified word pair can be obtained based on the word sequence. Two adjacent words in the word sequence can be combined into a word pair. , i is a positive integer less than or equal to n-1. Specifically, according to the word segmentation sequence T(D), adjacent word segmentation pairs can be identified and extracted, which can be done by simply traversing the word segmentation sequence and selecting each pair of consecutive word segmentations. For example, for the sequence T(D), the first word segmentation pair can be , the second participle pair can be , and so on, until the end of the word segmentation sequence.
[0096] For each word pair, we can combine them into a single entity. It can be a simple string (such as characters "+", "-", etc.) to split the word and participle Based on the n segmentation words in the segmentation sequence, n-1 segmentation word pairs can be obtained, that is, the specified segmentation word pairs can be obtained.
[0097] Step 620: Store the mapping relationship between the specified word segmentation pair and the identifier of the text to be stored in the first information database.
[0098] In the embodiment of the present application, for the i-th word pair in the specified word pair, , it can be determined whether the i-th word segmentation pair exists in the first information database, and the first information database is used to store the mapping relationship between the word segmentation pair and the identification information. In the case that the i-th word segmentation pair does not exist in the first information database, a new mapping relationship can be created in the first information database to store the mapping relationship between the i-th word segmentation pair and the text identifier (such as the document ID) of the text to be stored. In the case that the i-th word segmentation pair exists in the first information database, a text identifier set matching the i-th word segmentation pair can be obtained from the first information database. In the case that the text identifier set matching the i-th word segmentation pair includes the text identifier of the text to be stored, no other processing can be performed. In the case that the text identifier set matching the i-th word segmentation pair does not include the text identifier of the text to be stored, the text identifier of the text to be stored can be added to the text identifier set matching the i-th word segmentation pair, and the text identifier set including the text identifier of the text to be stored can be reassigned to the text identifier set matching the i-th word segmentation pair.
[0099] When there is new text to be stored in the second information base, the new text can be processed in the above manner to ensure that the first information base (CombineDict) can reflect the mapping relationship between adjacent word pairs and text identifiers of the latest text.
[0100] Step 630: Obtain a target word segmentation pair of the first search text, wherein the target word segmentation pair is used to reflect the contextual relationship of the first search text.
[0101] Step 640: Determine first identification information matching the target word segmentation pair in a first information database, wherein the first information database stores a mapping relationship between word segmentation pairs and identification information.
[0102] Step 650: Obtain a first search result based on the first identification information.
[0103] In one embodiment of the present application, the data retrieval method, in addition to steps 610 to 650, further includes: encrypting the text to be stored to obtain a target encrypted text; and storing the target encrypted text in a second information database.
[0104] In the embodiment of the present application, in order to ensure the security of data transmission and storage, the text to be stored can be encrypted to obtain the target encrypted text, and the encrypted target encrypted text can be stored in the second information library, and the second information library is used to store the mapping relationship between the encrypted text and the text identifier. The encrypted text matching the text identifier can be obtained from the second information library based on the text identifier.
[0105] In one embodiment of the present application, the data retrieval method, in addition to steps 610 to 650, further includes: obtaining historical search text from a cache library; and when the text to be stored contains the historical search text, writing the mapping relationship between the historical search text and the identification information of the text to be stored into the cache library.
[0106] In an embodiment of the present application, the cache library is used to store the mapping relationship between the historical search text and the identification information of the historical search text. In the case where there is a text to be stored in the second information library, since the historical search text may appear in the text to be stored, the identification information of the historical search text may be updated. In the case where the identification information of the historical search text is updated, the mapping relationship between the historical search text and the identification information of the text to be stored can be written into the cache library in time to ensure that the cache content in the cache library is based on the latest data.
[0107] Specifically, for the i-th original cache item among the Q original cache items in the cache library, it can be determined whether the i-th original cache item exists in the text to be stored. In the case that the i-th original cache item exists in the text to be stored, the text identifier set corresponding to the i-th original cache item can be obtained from the cache library, and the text identifier of the text to be stored is added to the text identifier set corresponding to the i-th original cache item. In the case that the i-th original cache item does not exist in the text to be stored, no other processing is performed.
[0108] See also Figure 7 , Figure 7 is a schematic diagram of a preprocessing process of a text to be stored provided by an embodiment of the present application. Figure 7As shown, for the text to be stored, the text to be stored can be segmented. Based on the segmentation result, a TermDict mapping is constructed, and adjacent segmentations are combined to obtain a specified segmentation pair. Based on the specified segmentation pair of the text to be stored, a CombineDict mapping is constructed. While the text to be stored is segmented, the text to be stored can be encrypted, and the encrypted target encrypted text is stored in the second information database, and the text identifier of the text to be stored is saved.
[0109] In an embodiment of the present application, a process for preprocessing the text to be stored is provided, and the mapping relationship between the identifier of the text to be stored and the specified word segmentation pair of the text to be stored can be stored in the first information database for subsequent matching of the target word segmentation pair using the first information database, and the identifier of the text to be stored and the specified word segmentation pair of the text to be stored stored in the first information database are based on the original text of the text to be stored, rather than the encrypted text. In the process of matching using the first information database, there is no need for decryption, which optimizes the retrieval query process.
[0110] See also Figure 8 , Figure 8 This is a specific flow chart of a data retrieval method provided in an embodiment of the present application. Figure 8 As shown, the method comprises the following steps: Step 810: Based on the text to be stored, obtain a specified word segmentation pair, where the specified word segmentation pair includes two adjacent word segments in the text to be stored.
[0111] In the embodiment of the present application, when the text to be stored is obtained, the text to be stored can be encrypted to obtain a target encrypted text, and the target encrypted text can be stored in the second information library. In addition, the historical search text can be obtained from the cache library; when the text to be stored contains the historical search text, the mapping relationship between the identification information of the historical search text and the text to be stored is written into the cache library.
[0112] Step 820: Store the mapping relationship between the specified word segmentation pair and the identifier of the text to be stored in the first information database.
[0113] Step 830: When the first search text does not exist in the cache library, obtain the target word segmentation pairs of the first search text; wherein the cache library includes a mapping relationship between historical search texts and identification information of the historical search texts, and the target word segmentation pairs include N word segmentation pairs arranged in sequence in the first search text.
[0114] In an embodiment of the present application, obtaining a target word segmentation pair for a first search text includes: performing word segmentation processing on the first search text to obtain multiple word segments; determining a target word segmentation pair from the multiple word segments, wherein the target word segmentation pair includes two adjacent word segments in the first search text.
[0115] In the case that the first search text exists in the cache library, first identification information matching the first search text may be obtained from the cache library, wherein the first identification information is used to obtain a second search result for the first search text.
[0116] Step 840: Select M word segment pairs from the N word segment pairs in order, where M is an integer less than or equal to N, and N is a positive integer.
[0117] Step 850: For each of the M word segmentation pairs, determine a text identifier set corresponding to the word segmentation pair in a first information database, wherein the first information database stores a mapping relationship between the word segmentation pairs and identifier information.
[0118] Step 860: Based on the M word segmentation pairs, obtain M text identifier sets.
[0119] Step 870: Obtain the intersection of the M text identification sets as the first identification result.
[0120] Step 880: Based on the first identification result, a common text identification is obtained, and the common text identification is used as the first identification information.
[0121] In the embodiments of the present application, the following three specific examples are provided. However, it should be noted that M is a positive integer less than or equal to N, and the value of M is not limited to the values in the following three examples.
[0122] Exemplarily, in one embodiment of the present application, M=N-1, the M word segmentation pairs are the first N-1 word segmentation pairs among the N word segmentation pairs, and N is an integer greater than 1; obtaining a common text identifier based on the first identification result includes: performing prefix matching on the Nth word segmentation pair among the N word segmentation pairs to obtain a first text identifier set; obtaining the intersection of the first identification result and the first text identifier set as the common text identifier.
[0123] Exemplarily, in one embodiment of the present application, M=N-1, the M word segmentation pairs are the last N-1 word segmentation pairs of the N word segmentation pairs, and N is an integer greater than 1; the first identification result obtains a common text identification, including: performing suffix matching on the first word segmentation pair of the N word segmentation pairs to obtain a second text identification set; obtaining the intersection of the first identification result and the second text identification set as the common text identification.
[0124] Exemplarily, in one embodiment of the present application, M=N-2, the M word segmentation pairs are the N-2 word segmentation pairs among the N word segmentation pairs, and N is an integer greater than 2; the common text identifier is obtained based on the first identification result, including: performing suffix matching on the first word segmentation pair among the N word segmentation pairs to obtain a first identification set; performing prefix matching on the Nth word segmentation pair among the N word segmentation pairs to obtain a second identification set; obtaining the intersection of the first identification result, the first identification set and the second identification set as the common text identifier.
[0125] In the embodiment of the present application, after the common text identifier is used as the first identifier information, if there is an available cache item in the cache library, the mapping relationship between the first search text and the first identifier information is stored in the cache library. If there is no available cache item in the cache library, the cache probability of the mapping relationship and the cache probability of the original cache item in the cache library are obtained; if the cache probability of the mapping relationship is greater than the cache probability of the target cache item in the cache library, the mapping relationship is stored in the cache library to replace the target cache item, and the target cache item is the original cache item with the smallest cache probability in the cache library.
[0126] Specifically, obtaining the cache probability of the mapping relationship and the cache probability of the original cache item in the cache library includes: obtaining the current timestamp, the first parameter of the mapping relationship and the second parameter of the original cache item in the cache library, the first parameter including: at least one of the first identification number contained in the first identification information, the first search number of the first search text and the first timestamp when the first search text was last searched; the second parameter including: at least one of the second identification number contained in the original cache item, the second search number of the original cache item and the second timestamp when the original cache item was last searched; based on the first parameter and the current timestamp, determining the cache probability of the mapping relationship; based on the second parameter and the current timestamp, determining the cache probability of the original cache item in the cache library.
[0127] The cache probability of the mapping relationship can be determined by the following formula: ; The cache probability of the original cache items in the cache library can be determined by the following formula: ; in, is the cache probability of the mapping relationship, is the cache probability of the original cache item, k is a preset weight coefficient, T is a preset constant, is the number of first identifications included in the first identification information, is the number of second identifiers contained in the original cache item, is the first search number of the first search text, is the second retrieval number of the original cache item, is the first timestamp of the last time the first search text was searched, is the second timestamp of the last retrieval of the original cache item, e is a natural constant, is the preset attenuation coefficient, and CT is the current timestamp.
[0128] Step 890: Obtain a first search result based on the first identification information.
[0129] In an embodiment of the present application, obtaining a first search result based on the first identification information includes: determining an encrypted text matching the first identification information from a second information database; and decrypting the encrypted text to obtain the first search result.
[0130] In an embodiment of the present application, the data retrieval method, in addition to steps 810 to 890, also includes: obtaining a single word segment of a second search text; determining second identification information matching the single word segment in a third information database, wherein the third information database stores a mapping relationship between word segments and identification information; and obtaining a third search result based on the second identification information.
[0131] In an embodiment of the present application, a target word segmentation pair of a first search text is obtained, and the target word segmentation pair is used to reflect the contextual relationship of the first search text; first identification information matching the target word segmentation pair is determined in a first information base, and the first information base stores the mapping relationship between the word segmentation pair and the identification information; based on the first identification information, a first search result is obtained. In this way, in the process of searching the first search text, the first identification information matching the target word segmentation pair can be determined in the first information base, and then the first search result is obtained through the first identification information. Compared with the method of extracting keywords for retrieval in the related technology, since the target word segmentation pair can reflect the contextual relationship of the first search text rather than isolated keywords, the first search result obtained through the first identification information matching the target word segmentation pair is more in line with the original intention of the first search text, and the first search result obtained is more accurate, which solves the problem of poor data retrieval results in the related technology.
[0132] It is important to understand that Figures 1 to 8 The explanations of the same or corresponding steps in the above may refer to each other. For example, Figure 1 The explanation of step 120 and step 130 in Figure 4 Steps 420 and 430 in .
[0133] At the same time, it should be understood that a data retrieval method provided by the embodiment of the present application can have the following beneficial effects: First, the retrieval efficiency is optimized. By using the dual mapping structure (TermDict and CombineDict), the present application optimizes the query processing flow, improves the retrieval efficiency of the encrypted text, and the retrieval results are more in line with the original intention of the retrieval text, and improves the accuracy of the retrieval results, especially when processing multi-word queries, it can locate relevant documents more quickly. Second, cache management optimization. The cache management mechanism of the present application dynamically adjusts the cache content based on the frequency of use of query words, document frequency and time decay factor through an intelligent eviction algorithm, reduces unnecessary computing and storage resource consumption, improves the efficiency and response speed of the cache, and meets the dual requirements of information security and retrieval performance. Third, the flexibility is enhanced. The introduction of the prefix matching algorithm enables the present application to handle more complex queries, including possible variant queries of users, and improves the flexibility of the system. Fourth, real-time guarantee. By monitoring database change events, the present application can update the cache in real time and perform invalidation processing, realize the consistency of cache data, and ensure the real-time and accuracy of query results.
[0134] In addition, with data leakage and privacy protection receiving increasing attention today, this application meets the market demand for high-security storage solutions through an encrypted storage mechanism. The optimized query processing flow and real-time guarantee enable users to quickly obtain accurate results, enhance user satisfaction and loyalty, and improve user experience. Intelligent cache management and resource optimization reduce server load and maintenance costs, reduce operating costs, and provide enterprises with a more cost-effective operating solution. The retrieval system provided by this application has the ability to expand services. The easy maintenance and scalability of the system provide convenience for adding new functions and services in the future, which is conducive to the long-term development and innovation of enterprises. This application supports big data applications. With the advent of the big data era, this application can process large-scale data sets and provide technical support for big data analysis and intelligent information retrieval. This application can be combined with multiple fields such as cloud storage, data analysis, and artificial intelligence to promote technological progress and economic growth in related industrial chains and promote the development of related industrial chains. In summary, this application has huge commercial value and is expected to bring economic and social benefits to enterprises or organizations.
[0135] See also Fig. 9 , Fig. 9 is a structural block diagram of a data retrieval device provided in an embodiment of the present application. Fig. 9 As shown, an embodiment of the present application provides a data retrieval device 900 , and the data retrieval device 900 includes: an acquisition module 910 , a determination module 920 and a retrieval module 930 .
[0136] The acquisition module 910 is used to acquire a target word segmentation pair of the first search text, wherein the target word segmentation pair is used to reflect the contextual relationship of the first search text; The determination module 920 is used to determine the first identification information matching the target word segmentation pair in a first information database, wherein the first information database stores a mapping relationship between word segmentation pairs and identification information; The retrieval module 930 is used to obtain a first retrieval result based on the first identification information.
[0137] In an embodiment of the present application, a target word segmentation pair of a first search text is obtained, and the target word segmentation pair is used to reflect the contextual relationship of the first search text; first identification information matching the target word segmentation pair is determined in a first information base, and the first information base stores the mapping relationship between the word segmentation pair and the identification information; based on the first identification information, a first search result is obtained. In this way, in the process of searching the first search text, the first identification information matching the target word segmentation pair can be determined in the first information base, and then the first search result is obtained through the first identification information. Compared with the method of extracting keywords for retrieval in the related technology, since the target word segmentation pair can reflect the contextual relationship of the first search text rather than isolated keywords, the first search result obtained through the first identification information matching the target word segmentation pair is more in line with the original intention of the first search text, and the first search result obtained is more accurate, which solves the problem of poor data retrieval results in the related technology.
[0138] The data retrieval device provided in the embodiment of the present application can implement each process implemented in the above method embodiment, and will not be described again here to avoid repetition.
[0139] like Fig.10As shown, an embodiment of the present application also provides an electronic device 1000. The electronic device 1000 includes: a processor 1010 and a memory 1020, and the memory 1020 stores a program or instruction, and the program or instruction implements the steps of any of the methods described above when executed by the processor 1010. For example, when the program is executed by the processor 1010, the following process is implemented: obtaining a target word segmentation pair of a first search text, and the target word segmentation pair is used to reflect the contextual relationship of the first search text; determining first identification information matching the target word segmentation pair in a first information base, and the first information base stores a mapping relationship between word segmentation pairs and identification information; based on the first identification information, obtaining a first search result. In this way, in the process of searching the first search text, the first identification information matching the target word segmentation pair can be determined in the first information library, and then the first search result can be obtained through the first identification information. Compared with the method of extracting keywords for retrieval in related technologies, since the target word segmentation pair can reflect the contextual relationship of the first search text rather than isolated keywords, the first search result obtained through the first identification information matching the target word segmentation pair is more in line with the original intention of the first search text, and the first search result obtained is more accurate, which solves the problem of poor data retrieval results in related technologies.
[0140] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of each embodiment of the data retrieval method are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0141] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0142] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0143] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0144] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0145] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0146] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A data retrieval method, characterized in that: include: Acquire a target word segmentation pair of a first search text, wherein the target word segmentation pair is used to reflect the contextual relationship of the first search text; Determining first identification information matching the target word segmentation pair in a first information database, wherein the first information database stores a mapping relationship between word segmentation pairs and identification information; Based on the first identification information, a first search result is obtained.
2. The method according to claim 1, characterized in that The step of obtaining a target word pair of the first search text includes: Performing word segmentation processing on the first search text to obtain multiple word segments; A target word segmentation pair is determined from the multiple word segments, where the target word segmentation pair includes two adjacent word segments in the first search text.
3. The method according to claim 1, characterized in that Based on the first identification information, a first search result is obtained, including: Determining, from a second information base, an encrypted text matching the first identification information; The encrypted text is decrypted to obtain a first search result.
4. The method according to claim 1, characterized in that: The target word segmentation pairs include N word segmentation pairs arranged sequentially in the first search text; Determining first identification information matching the target word pair in a first information database includes: Select M word pairs from N word pairs in order, where M is an integer less than or equal to N, and N is a positive integer; For each of the M word segmentation pairs, determining a text identifier set corresponding to the word segmentation pair in the first information database; Based on M word pairs, obtain M text identifier sets; Obtaining the intersection of the M text identification sets as a first identification result; Based on the first identification result, a common text identification is obtained, and the common text identification is used as the first identification information.
5. The method according to claim 4, characterized in that M=N-1, the M word pairs are the first N-1 word pairs among the N word pairs, and N is an integer greater than 1; the common text identification is obtained based on the first identification result, including: Perform prefix matching on the Nth word segmentation pair among the N word segmentation pairs to obtain a first text identifier set; The intersection of the first identification result and the first text identification set is obtained as the common text identification.
6. The method according to claim 4, characterized in that M=N-1, the M word pairs are the last N-1 word pairs of the N word pairs, and N is an integer greater than 1; the first identification result, obtaining a common text identification, includes: Perform suffix matching on the first word pair of the N word pairs to obtain a second text identifier set; The intersection of the first identification result and the second text identification set is obtained as the common text identification.
7. The method according to claim 4, characterized in that M=N-2, the M word pairs are N-2 word pairs in the middle of the N word pairs, and N is an integer greater than 2; the common text identification is obtained based on the first identification result, including: Perform suffix matching on the first word pair of the N word pairs to obtain a first identifier set; Perform prefix matching on the Nth word pair among the N word pairs to obtain a second identifier set; An intersection of the first identification result, the first identification set, and the second identification set is obtained as a common text identification.
8. The method according to claim 1, characterized in that: Obtaining the target word pair of the first search text includes: If the first search text does not exist in the cache library, obtaining a target word segmentation pair of the first search text; The cache library includes a mapping relationship between historical search texts and identification information of the historical search texts.
9. The method according to claim 8, characterized in that The method further comprises: In the case that the first search text exists in the cache library, target identification information matching the first search text is obtained from the cache library, wherein the target identification information is used to obtain a second search result for the first search text.
10. The method according to claim 8, characterized in that After determining in the first information database the first identification information matching the target word segmentation pair, the method further includes: In the case where there is an available cache item in the cache library, storing the mapping relationship between the first search text and the first identification information in the cache library; When there is no available cache item in the cache library, the cache probability of the mapping relationship and the cache probability of the original cache item in the cache library are obtained; when the cache probability of the mapping relationship is greater than the cache probability of the target cache item in the cache library, the mapping relationship is stored in the cache library to replace the target cache item, and the target cache item is the original cache item with the smallest cache probability in the cache library.
11. The method according to claim 10, characterized in that Obtaining the cache probability of the mapping relationship and the cache probability of the original cache items in the cache library includes: Obtaining a current timestamp, a first parameter of the mapping relationship, and a second parameter of an original cache item in the cache library, wherein the first parameter includes: at least one of the first identification number included in the first identification information, the first search number of the first search text, and the first timestamp when the first search text was last searched; and the second parameter includes: at least one of the second identification number included in the original cache item, the second search number of the original cache item, and the second timestamp when the original cache item was last searched; Determining a cache probability of the mapping relationship based on the first parameter and a current timestamp; Based on the second parameter and the current timestamp, a cache probability of the original cache item in the cache library is determined.
12. The method according to claim 11, characterized in that The cache probability of the mapping relationship is determined by the following formula: ; The cache probability of the original cache items in the cache library is determined by the following formula: ; in, is the cache probability of the mapping relationship, is the cache probability of the original cache item, k is a preset weight coefficient, T is a preset constant, is the number of first identifications included in the first identification information, is the number of second identifiers contained in the original cache item, is the first search number of the first search text, is the second retrieval number of the original cache item, is the first timestamp of the last time the first search text was searched, is the second timestamp of the last retrieval of the original cache item, e is a natural constant, is the preset attenuation coefficient, and CT is the current timestamp.
13. The method according to claim 1, characterized in that The method further comprises: Obtain a single word segment of the second search text; Determining second identification information matching the single segmented word in a second information database, wherein the second information database stores a mapping relationship between the segmented words and the identification information; Based on the second identification information, a third search result is obtained.
14. The method according to claim 1, characterized in that Before obtaining the target word pair of the first search text, the method further includes: Based on the text to be stored, obtaining a specified word pair, wherein the specified word pair includes two adjacent word pairs in the text to be stored; The mapping relationship between the specified word segmentation pair and the identifier of the text to be stored is stored in the first information database.
15. The method according to claim 14, characterized in that The method further comprises: Get the historical search text from the cache library; In the case that the text to be stored includes the historical search text, the mapping relationship between the historical search text and the identification information of the text to be stored is written into the cache library.
16. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction running on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 15 are implemented.
17. A computer-readable storage medium, characterized in that: The medium stores a program or an instruction, and when the program or the instruction is executed, the steps of the method according to any one of claims 1 to 15 are implemented.
18. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 15 when executed by a processor.