Method, System and Electronic Device for Querying Keywords with a Multi-Thesaurus Shared Data Structure

By constructing a shared data structure of multiple thesaurus and using a dictionary tree to merge keywords from multiple thesaurus, the complexity and storage space occupation problems caused by the increase in the number of vocabulary under traditional methods are solved, and efficient keyword query and resource sharing are achieved.

CN118550932BActive Publication Date: 2025-06-10GUANGZHOU AIYOU INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410628699.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2025-06-10
Estimated Expiration
2044-05-21

AI Technical Summary

Technical Problem

Traditional string search methods lead to a sharp increase in the number of vocabulary when processing multiple projects, increasing the complexity of data management and maintenance, and occupying a lot of storage space.

Method used

By building a shared data structure of multiple thesaurus, using a dictionary tree to merge keywords from multiple thesaurus into one data structure, efficient keyword query is achieved, and storage space is optimized through dynamic data structures and segmented storage.

Benefits of technology

It realizes efficient integration and resource sharing of multiple vocabulary libraries, reduces the use of storage space, and improves the accuracy and efficiency of information query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118550932B_ABST
    Figure CN118550932B_ABST
Patent Text Reader

Abstract

The present invention provides a method, a system and an electronic device for querying keywords using a shared data structure of multiple thesauruses, belonging to the field of information technology. The method includes: obtaining the keywords of a single project and constructing a first trie of the keywords of the single project; obtaining the keywords of all projects, inserting the keywords of all projects into the first trie to obtain a second trie; traversing all nodes of the second trie and constructing the nodes for jump on failed match in the second trie to obtain a third trie. Among them, a dynamic data structure is used to store the child nodes of the third trie, and each node of the obtained third trie is stored in segments; the third trie is used for keyword query. This method can achieve efficient integration and resource sharing of multiple thesauruses, improve the accuracy and efficiency of information query, and avoid the redundancy of storing keywords separately for each thesaurus, which can greatly reduce the required storage space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and particularly to a method, a system, and an electronic device for querying keywords using a shared data structure for multiple thesauruses. Background Art

[0002] In the fields of information retrieval, natural language processing, and data mining, keyword query is a crucial technology. Traditional string search methods mainly rely on specific algorithms to encapsulate multiple keywords into a data structure form convenient for searching, so as to achieve efficient string matching. When it is necessary to query whether a certain string contains a specified keyword, the system will compare the string to be queried with the constructed data structure to determine whether it contains the specified keyword.

[0003] This traditional string search method has significant limitations in practice. Especially when dealing with multiple projects, each project often has its own independent data structure of the thesaurus. This means that as the number of projects increases, the number of thesauruses to be constructed will also increase sharply. This not only increases the complexity of data management and maintenance, but more critically, a large amount of thesaurus data occupies valuable storage space, making the system resources not effectively utilized.

[0004] Therefore, it is necessary to improve the existing keyword query method to overcome the defects of the existing technology. Summary of the Invention

[0005] To overcome the problems existing in the related art, one of the objectives of the present invention is to provide a method for querying keywords using a shared data structure for multiple thesauruses. This method can achieve efficient integration and resource sharing of multiple thesauruses, improve the accuracy and efficiency of information query, and avoid the redundancy of storing keywords separately for each thesaurus, which can greatly reduce the required storage space.

[0006] A method for querying keywords using a shared data structure for multiple thesauruses includes:

[0007] Obtain the keywords of a single project and construct a first trie for the keywords of the single project;

[0008] Obtain the keywords of all projects and insert the keywords of all projects into the first trie to obtain a second trie;

[0009] Traverse all nodes of the second trie and construct the mismatch jump nodes in the second trie to obtain a third trie; wherein, a dynamic data structure is used to store the child nodes of the third trie, and the nodes of the obtained third trie are stored in segments;

[0010] Use the third trie to query keywords.

[0011] In a preferred technical solution of the present invention, inserting the keywords of all items into the first trie further includes:

[0012] Marking the keywords of different items with different marks.

[0013] In a preferred technical solution of the present invention, marking the keywords of different items with different marks includes:

[0014] Constructing the end of each keyword;

[0015] Hanging a data structure at the end of each keyword; the data structure is used to mark item information.

[0016] In a preferred technical solution of the present invention, obtaining the keyword of a single item includes:

[0017] Obtaining the keyword of a single item;

[0018] Preprocessing the obtained keyword;

[0019] Extracting the keyword of the preprocessed keyword.

[0020] In a preferred technical solution of the present invention, preprocessing the obtained keyword includes:

[0021] Removing duplicates, merging synonyms, and processing morphological variations.

[0022] In a preferred technical solution of the present invention, constructing the failure jump node in the second trie includes:

[0023] Adding a failure pointer to each node in the second trie except the root node; the failure pointer points to the node to jump to when the match fails at the current node;

[0024] Among them, during the process of adding the failure pointer, breadth-first search is used to construct the failure pointer.

[0025] In a preferred technical solution of the present invention, inserting the keywords of all items into the first trie to obtain the second trie includes:

[0026] During the insertion process, if an existing character node is encountered, share the existing character node.

[0027] In a preferred technical solution of the present invention, using the third trie for keyword query includes:

[0028] Constructing a unified query interface in the third trie;

[0029] Perform keyword search through a unified query interface.

[0030] A second object of the present invention is to provide a system for querying keywords with a shared data structure of a thesaurus, the system being used to implement the method for querying keywords with a shared data structure of multiple thesauruses as described above. The system includes:

[0031] A shared data structure unit, which is used to store keyword information in multiple thesauruses, and by constructing a trie tree, fuse the keywords in multiple thesauruses into one trie tree;

[0032] Process the keywords, remove duplicates in the keywords, merge synonyms, and process morphological variations;

[0033] A keyword query unit, which is used to traverse all nodes of the trie tree in the order of the characters of the keyword to search for matching keywords;

[0034] A result return unit, which is used to feedback the search result of the keyword query unit to the user;

[0035] A thesaurus update unit, which is used to update the thesaurus of the shared data structure unit and modify the trie tree in the shared data structure unit.

[0036] A third object of the present invention is to provide an electronic device, including:

[0037] A processor; and

[0038] A memory, on which executable code is stored. When the executable code is executed by the processor, the processor executes the method for querying keywords with a shared data structure of multiple thesauruses as described above.

[0039] The beneficial effects of the present invention are:

[0040] A method for querying keywords with a shared data structure of multiple thesauruses provided by the present invention includes: obtaining keywords of a single project and constructing a first trie tree of the keywords of the single project; obtaining keywords of all projects, inserting the keywords of all projects into the first trie tree to obtain a second trie tree; traversing all nodes of the second trie tree and constructing a jump node for failed matching in the second trie tree to obtain a third trie tree; and using the third trie tree to query keywords. This method fuses the keywords of multiple thesauruses into one trie tree, realizes the sharing and reuse of thesauruses, and saves storage space. Moreover, the more projects there are and the larger the repeated thesaurus is, the more space can be saved by the method of the present application. In addition, since the trie tree is a data structure based on prefix matching, when querying keywords, it is not necessary to traverse the entire thesaurus, and only need to go down along the structure of the trie tree, which can greatly improve the query efficiency.

[0041] The present application also provides a system and an electronic device for querying keywords using a shared data structure for multiple thesauruses. This system can achieve efficient integration and resource sharing of multiple thesauruses, reduce the work of constructing different thesauruses, and reduce the space occupied by the data structures of different thesauruses. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a flowchart of the method for querying keywords using a shared data structure for multiple thesauruses provided by the present invention;

[0043] Figure 2 is a flowchart of preprocessing keywords provided by the present invention;

[0044] Figure 3 is a schematic diagram of the trie tree constructed for the "stars" project provided in Embodiment 1 of the present invention;

[0045] Figure 4 is a schematic diagram of the trie tree constructed for the "flags" project provided in Embodiment 1 of the present invention;

[0046] Figure 5 is a schematic diagram of the trie tree for constructing the jump section for failed matches in the trie trees of the "stars" and "flags" projects provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.

[0048] In the fields of information retrieval, natural language processing, and data mining, keyword query is a crucial technology. Traditional string search methods mainly rely on specific algorithms to encapsulate multiple keywords into a data structure form convenient for searching to achieve efficient string matching. When it is necessary to query whether a certain string contains a specified keyword, the system will compare the string to be queried with the constructed data structure to determine whether it contains the specified keyword.

[0049] This traditional string search method has significant limitations in practice. Especially when dealing with multiple projects, each project often has its own independent thesaurus data structure. This means that as the number of projects increases, the number of thesauruses to be constructed will also increase sharply. This not only increases the complexity of data management and maintenance, but more importantly, a large amount of thesaurus data occupies valuable storage space, making the system resources not effectively utilized.

[0050] Based on this, the present application provides a method for querying keywords using a common data structure for multiple thesauruses.

[0051] Embodiment 1

[0052] As Figure 1 - Figure 2 shown, this embodiment provides a method for querying keywords using a common data structure for multiple thesauruses, including the following steps:

[0053] S100. Obtain the keywords of a single project and construct a first trie tree for the keywords of the single project;

[0054] S200. Obtain the keywords of all projects, insert the keywords of all projects into the first trie tree to obtain a second trie tree;

[0055] S300. Traverse all nodes of the second trie tree and construct matching failure jump nodes in the second trie tree to obtain a third trie tree; wherein, a dynamic data structure is used to store the child nodes of the third trie tree, and each node of the obtained third trie tree is stored in segments;

[0056] Specifically, the dynamic data structure of the present application includes linked lists, dynamic arrays (such as std::vector in C++ or list in Python), and hash tables (such as std::unordered_map in C++ or dict in Python);

[0057] More specifically, when storing child nodes, it is necessary to define a node class (or structure) of the trie tree, which includes a character field to store the character of the current node and a dynamic data structure field pointing to the child nodes.

[0058] For example, in Python, a node can be defined as follows:

[0059] class TrieNode:

[0060] def __init__(self):

[0061] self.char = None

[0062] self.children = {} # Use a dictionary as a dynamic data structure to store child nodes

[0063] self.is_end_of_word = False

[0064] In C++, the node is defined as follows:

[0065] class TrieNode {

[0066] public:

[0067] char ch;

[0068] std::unordered_map<char, TrieNode*> children; / / Use a hash table to store child nodes

[0069] bool is_end_of_word;

[0070] TrieNode(char c) : ch(c), is_end_of_word(false) {}

[0071] };

[0072] After defining the node class of the trie, the trie is segmented according to the defined node class of the trie. The segmentation strategy can be based on factors such as the height of the trie, the number of nodes, memory limitations, or access patterns. For example, segment by the level of the trie, or segment according to the number of child nodes of each node.

[0073] According to the selected segmentation strategy, the trie is split into multiple smaller parts or "segments". Each segment can contain a part of the nodes and subtrees of the trie. Ensure that the size of each segment is appropriate, neither too large to cause memory problems nor too small to cause management complexity.

[0074] After segmentation, each segment is stored as a separate file or memory block. Specifically, the file system can be used to store files, or memory-mapped files, databases, or other memory management techniques can be used to store memory blocks. It should be noted that a node mapping table needs to be established during segmentation to record which segment each node is in and its position within the segment. In this way, during search, the corresponding segment where the node is located can be found according to the node mapping table, and the search can be performed within the segment. Or create an index to track the node information contained in each segment. The index can be a simple list of files, or a more complex data structure such as a hash table or a B-tree, for quickly locating the segment containing a specific node. The index should contain sufficient information to quickly locate the segment containing the target node.

[0075] After segment - by - segment completion and storage, if node insertion and deletion are required, first determine the segment where the new node or the node to be deleted is located. Then, load that segment and perform the corresponding operations. In the insertion operation, if the new node causes the current segment to be too large, the new node needs to be split into two or more segments. Similarly, in the deletion operation, if the deletion of a node causes the current segment to be too small, it needs to be merged with other segments. Additionally, in the later stage, the trie stored in segments needs to be regularly checked and optimized. This may include re - balancing the segment sizes, updating the indexes, cleaning up segments that are no longer needed, etc. Ensure that the trie stored in segments remains efficient and manageable.

[0076] Preferably, the present application also provides a method for quickly finding each node file after segment - by - segment storage of the dictionary e - tree, which specifically includes:

[0077] Create a separate index file, where the index file records information about each segment file. The index file can contain information such as the starting node, ending node, file path, or file handle of each segment. When a certain node needs to be found, first query the index file to determine which segment file the node is located in, and then load the corresponding segment file.

[0078] It can also include encoding the paths of the trie so that the position information of the nodes can be directly calculated through the encoding. For example, the depth of the node and the order of pre - order traversal can be used as the encoding. When a certain node needs to be found, directly calculate the segment file where it is located according to the encoding of the node. This method does not require an additional index file, but it is necessary to ensure that the encoding scheme is efficient and reliable.

[0079] During the process of encoding the paths of the trie:

[0080] First, an encoding structure that can uniquely identify each node path in the trie is required. This encoding structure usually includes the depth of the node in the trie and its relative position at the current level (for example, the order from left to right)

[0081] Then use depth - first traversal: Starting from the root node of the trie, use depth - first traversal (DFS) to assign a unique encoding to each node. This encoding can be an integer, a string, or any data structure that can uniquely identify the path. The encoding can consist of two parts: The first is the segment identifier: indicating the segment where the node is located. This can be a simple segment index (for example, the name or number of the segment file), or segment information based on the node depth. The first is the node position: indicating the position of the node within its segment. This can be a relative position index based on the level, or an encoding based on a certain order (such as lexicographical order).

[0082] Store the encoding of each node in a data structure for quick lookup. This data structure can be a hash table that maps the encoding to a reference to the node or node information

[0083] When looking up a node in the trie, first generate the corresponding encoding based on the given keyword; use the encoding to look up the node information in the hash table; locate the corresponding segmented file based on the node information and find the specific node in that file.

[0084] For example: Suppose we have a trie that is divided into two segments by depth: the first segment contains nodes with depth 0 and 1, and the second segment contains nodes with depth 2 and above. Then assign an encoding to each node as follows:

[0085] Root node (depth 0): Encoding is 0:root (where 0 represents the segment identifier and root represents the special identifier of the root node)

[0086] First layer nodes (depth 1): Encoding is 1:A, 1:B,... (where 1 represents the segment identifier and A, B, etc. represent the characters of the nodes)

[0087] Second layer and deeper nodes: Encoding is 2:A.B, 2:A.C,... (where 2 represents the segment identifier and A.B, A.C, etc. represent the path from the root node to this node)

[0088] In this way, through the encoding, the segment where any node is located and its position within the segment can be quickly located. It should be noted that the encoding scheme should be designed according to the specific trie structure and segmentation strategy to ensure the effectiveness and efficiency of the encoding.

[0089] S400. Use the third trie for keyword query.

[0090] Specifically, obtaining the keyword of a single item includes:

[0091] S110. Obtain the keyword of a single item;

[0092] S120. Preprocess the obtained keyword; where preprocessing the obtained keyword includes:

[0093] Remove duplicates, merge synonyms, and handle morphological variations.

[0094] Removing duplicates means removing all duplicates of the keywords in a single item.

[0095] During the process of merging synonyms, when synonyms are encountered, merge them into the same node. This node can contain a list of synonyms and a pointer to the shared child node at the next level.

[0096] Processing inflectional changes helps to construct an accurate and efficient trie. Specifically, the methods for processing inflectional changes include but are not limited to the following:

[0097] Stemming: Stemming is a process of reducing a word to its base form or stem. For example, reducing "runn i ng", "ran", and "runs" to "run". By this method, only the stems can be stored in the trie, thus reducing the storage space and improving the query efficiency.

[0098] Lemmatization: Lemmatization is more precise than stemming. It attempts to reduce a word to its base form or lemma in the dictionary. This usually relies on language rules and dictionary resources. Lemmatization can better handle irregular verbs and plural forms of nouns.

[0099] Using regular expressions: Before constructing the trie, regular expressions can be used to preprocess the input word forms. For example, regular expressions can be used to remove special characters, punctuation marks from words or convert them to lowercase. This can ensure that the standardized word forms are stored in the trie.

[0100] Writing general processing rules: For some common inflectional changes (such as adding or removing prefixes and suffixes), general processing rules can be written. These rules can be automatically applied when constructing the trie to handle common inflectional changes.

[0101] S130. Extract the keywords of the preprocessed keywords. Extracting the preprocessed keywords as the input of the trie can reduce the redundant nodes of the trie, streamline the structure of the trie, and reduce the memory.

[0102] The process of constructing the first trie for the keywords of a single item is as follows:

[0103] Start constructing the trie from an empty root node. Each node represents a character, and the path from the root node to any node represents a specific keyword prefix. Each node represents a character, and the path from the root node to any node represents a specific keyword prefix.

[0104] Add the preprocessed keywords to the trie one by one. For each keyword, start from the root node and traverse down along the tree structure. If there is a node in the children of the current node that matches the current character of the keyword, continue traversing down; otherwise, create a new child node and use the current character as the label of this node.

[0105] Repeat the above process until all keywords are added to the trie to obtain the data structure of the first trie.

[0106] After obtaining the first trie, the keywords of all items are preprocessed using the same preprocessing method, and the preprocessed keywords are used as input and inserted into the data structure of the first trie.

[0107] More preferably, the process of inserting the keywords of all items into the first trie to obtain the second trie includes:

[0108] During the insertion process, if an existing character node is encountered, the existing character node is shared.

[0109] By fusing the keywords of multiple libraries into a shared trie, the redundancy of storing keywords separately for each library is avoided. This greatly reduces the required storage space.

[0110] In a preferred technical solution of the present invention, the insertion of the keywords of all items into the first trie further includes:

[0111] Using different markers to mark the keywords of different items.

[0112] Marking the keywords of different items helps to distinguish and identify the vocabulary from different items, ensuring that the results can be accurately associated with the corresponding items during the query and retrieval processes. This is crucial for data management and analysis in a multi-item environment and can avoid confusion and misunderstanding. Additionally, marking different items can also improve the query efficiency. In the trie, through item marking, the system can quickly locate the vocabulary of a specific item, thus avoiding traversing the entire trie and reducing unnecessary calculations and resource consumption. Moreover, item marking also helps to achieve data isolation and security. In a multi-user or multi-item scenario, different items may involve sensitive or private information. By setting a unique marker for each item, it can be ensured that only authorized users can access and operate the vocabulary of a specific item, thereby protecting the integrity and security of the data. Finally, item marking can also facilitate data analysis. By statistically analyzing the vocabulary of different items, information such as the vocabulary usage and popular words of each item can be understood, providing data support for the optimization and improvement of the project.

[0113] This application also provides a method for marking the keywords of different items: the use of different markers to mark the keywords of different items includes:

[0114] Construct the end of each keyword; specifically, mark the node corresponding to the last character of each keyword to indicate that the node is the end of a string. This can be achieved by setting a boolean variable _isEnd or adding a special marker, such as an end marker, to the node structure.

[0115] Hang a data structure at the end of each keyword; the data structure is used to mark item information.

[0116] The data structure can be an integer, an enumeration value, a string, or other appropriate data types, depending on the number of items and the identification method.

[0117] The data structure hung at the end of each keyword can be an object, a list, or other containers to store item information related to the keyword. For example, if implemented in C++, an unordered_map can be added to the TrieNode structure to store item information.

[0118] After hanging the data structure at the end of each keyword, the keyword can be queried and updated: when querying or updating, the corresponding keyword node can be found by traversing the trie, and then the operation can be performed through the hung data structure.

[0119] Specifically, constructing the failure jump nodes in the second trie includes:

[0120] Add a failure pointer to each node in the second trie except the root node; the failure pointer points to the node to jump to when the current node fails to match.

[0121] In the process of adding the failure pointer, breadth-first search is used to construct the failure pointer. The process of constructing the failure pointer by breadth-first search is as follows: starting from the child nodes of the root node, process layer by layer downward. For each node, if the child node of a certain character does not exist, then the failure pointer corresponding to that character points to the root node. If the child node of a certain character exists, then the failure pointer corresponding to that character points to the child node of the same character in the node pointed to by the failure pointer of its parent node (if it exists). If not, continue to search upward along the failure pointer until a matching character is found or the root node is reached.

[0122] In a better implementation, the construction process of the failure pointer can also be optimized and extended. For example, additional information can be added to accelerate the matching process, or more complex matching patterns can be processed.

[0123] In one implementation, using the third trie for keyword query includes:

[0124] Construct a unified query interface in the third trie;

[0125] Perform keyword search through the unified query interface.

[0126] Specifically, the process of constructing the unified query interface is as follows:

[0127] First, design a query interface. By defining a unified query interface, this interface receives the user's query request and returns the query result. This interface can be a function, an API call, or other forms of interaction.

[0128] Then, parse the query request through the query interface. In the query interface, parse the query request entered by the user. This may include steps such as word segmentation, removing irrelevant characters, and converting to a unified format to ensure that the query request matches the vocabulary in the trie.

[0129] In the query interface, a project identifier can be introduced to distinguish the thesauruses of different projects. Specifically, a project identifier can also be introduced in the query request. This identifier can be the name, ID, or other unique identifier of the project, and this identifier corresponds to the tag of the project.

[0130] During the query process, according to the project identifier provided by the user, locate the thesaurus corresponding to the project in the trie. Then, perform a query operation in this thesaurus and use the search algorithm of the trie to find the matching vocabulary.

[0131] The above method for querying keywords using a shared data structure for multiple thesauruses realizes the sharing and reuse of thesauruses by integrating the keywords of multiple thesauruses into one trie, saving storage space. Moreover, the more projects there are and the larger the repeated thesauruses are, the more space can be saved by the method of this application. In addition, since the trie is a data structure based on prefix matching, when querying keywords, it is not necessary to traverse the entire thesaurus, but only need to go down along the structure of the trie, which can greatly improve the query efficiency.

[0132] The following provides a detailed implementation of the method for querying keywords using a shared data structure for multiple thesauruses:

[0133] Step 1: Construct an n-keyword trie for a single project and mark the keyword records of the project. See Figure 3 , taking the "Stars" project as an example, add the search keywords "not good-looking, good-looking, easy to do?, good-looking but not up to standard" to construct the trie.

[0134] Step 2: Repeat Step 1 to construct the keywords of all projects into the same trie. See Figure 4 , taking the "Flags" project as an example, add the keywords "look good, not good-looking, easy to do?" to the same trie.

[0135] Step 3: Loop through each node in the trie and construct the jump nodes for failed matches. See Figure 5 , the first dotted line is the jump point for failed matches of the "Stars" project, and the second dotted line is the jump point for failed matches of the "Flags" project.

[0136] After the above three steps, a data structure of a common query trie integrating the thesauruses of two projects can be obtained. There are a total of seven keywords in the two projects, and two of them are common, namely "not good-looking, easy to do or not". As a result, the generated query structure tree has only four branches, which can effectively reduce the space occupation.

[0137] Query process: When querying "not good-looking, what about the upper part", if the "Stars" project team conducts the query, it will first query the branch of "not -> good -> looking" to find the keyword "not good-looking", and then jump to the "looking" node of the branch of "good -> looking -> upper -> not" to find the keyword "good-looking". Then, it further queries to obtain the keyword "good-looking upper not", and then jumps to the branch of "not -> good -> looking", but there is no matching branch anymore. Therefore, the final results are the three keywords "not good-looking, good-looking, good-looking upper not".

[0138] If the "Flags" project team queries the same string "not good-looking, what about the upper part", it will also first query the branch of "not -> good -> looking" to find the keyword "not good-looking", and then jump to the branch of "looking -> good", but no keyword is found. Then, it enters the branch of "not -> good -> looking", and still no keyword is found. The final result is only the one keyword "not good-looking".

[0139] In practical applications, assuming the number of projects is m and the number of thesauruses is n, the space occupied by the query is m * n. If 90% of the thesauruses in these m projects are duplicate content, then the space occupied by this method is only n * 90% + (m * n * 10%). The more duplicate thesauruses there are and the larger the number of projects, the more space this method can save.

[0140] Embodiment 2

[0141] This embodiment provides a system for querying keywords in a shared thesaurus data structure. The system for querying keywords in a shared thesaurus data structure is used to implement the method for querying keywords in a shared thesaurus data structure as described above.

[0142] The system includes:

[0143] A shared data structure unit, which is used to store the keyword information in multiple thesauruses, and by constructing a trie, integrate the keywords in multiple thesauruses into one trie; process the keywords to remove duplicate items, merge synonyms, and process morphological variations;

[0144] The shared data structure unit uses a trie as the basic structure. By inserting keywords from different lexical databases, a shared trie containing multiple lexical databases is formed. Meanwhile, this module is also responsible for handling keyword conflicts and duplicates to ensure the correctness and consistency of the shared trie.

[0145] The keyword query unit is used to traverse all nodes of the trie in the order of the characters of the keyword to search for the matching keyword; it queries in the trie of the shared data structure unit. During the query process, the module first starts from the root node of the trie in the shared data structure unit and traverses layer by layer according to the character order of the keyword until it finds the matching keyword or reaches the leaf node. To improve the query efficiency, this unit also adopts the failure pointer technique. When the current character does not exist in the children nodes of the current node, it jumps to other possible nodes through the failure pointer to continue the query.

[0146] The result return unit is used to feedback the search result of the keyword query unit to the user;. When the user's query is successful, the result return unit returns the matching keyword and its related information; when the query fails, the result return unit returns the corresponding prompt information. Meanwhile, the result return unit also supports functions such as multi-keyword query and fuzzy query to meet the needs of different users.

[0147] The lexical database update unit is used to update the lexical database of the shared data structure unit and modify the trie in the shared data structure unit. When a new lexical database or keyword is added, the lexical database update unit inserts it into the trie of the shared data structure unit; when a lexical database or keyword is deleted, the lexical database update unit deletes the corresponding node from the trie of the shared data structure unit. In addition, the lexical database update unit can also support the batch import and export functions of the lexical database to facilitate users to manage and maintain the lexical database.

[0148] Embodiment 3

[0149] This embodiment provides an electronic device, which includes a memory and a processor.

[0150] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), application-specific integrated circuits

[0151] (Application Specific Integrated Circuit, ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0152] The memory may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices.

[0153] Executable code is stored on the memory, and when the executable code is processed by the processor, it can cause the processor to execute some or all of the methods for querying keywords in the multi-lexicon shared data structure described above.

[0154] Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present application. In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.

[0155] It should be understood that the spatially relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is inverted, the device described as "above other devices or structures" or "on top of other devices or structures" will then be oriented "below other devices or structures" or "beneath other devices or structures". Thus, the exemplary term "above" can include both orientations of "above" and "below". The device may also be oriented in other different ways (rotated 90 degrees or at other orientations), and the spatially relative descriptions used herein are to be interpreted accordingly.

[0156] In addition, it should be noted that the use of terms such as "first" and "second" to limit components is only for the convenience of distinguishing the corresponding components. Without further declaration, the above terms have no special meaning and thus should not be construed as limiting the protection scope of the present application.

[0157] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for searching keywords using a multi-lexicon shared data structure, characterized in that: include: Get the keywords of a single item and construct the first dictionary tree of the keywords of the single item; Obtain keywords of all items, insert the keywords of all items into the first dictionary tree, and obtain a second dictionary tree; specifically, during the insertion process, if an existing character node is encountered, the existing character node is shared; Traverse all nodes of the second dictionary tree, and construct a match failure jump node in the second dictionary tree to obtain a third dictionary tree; wherein, a dynamic data structure is used to store the child nodes of the third dictionary tree, and each node of the obtained third dictionary tree is stored in segments; specifically, the construction of the match failure jump node in the second dictionary tree includes: adding a failure pointer to each node except the root node in the second dictionary tree; the failure pointer points to the node to jump to when the current node fails to match; wherein in the process of adding the failure pointer, a breadth-first search is used to construct the failure pointer; and the third dictionary tree is used to perform keyword search.

2. The method for searching keywords using a multi-lexicon shared data structure according to claim 1, characterized in that: The method of inserting the keywords of all items into the first dictionary tree further includes: Use different tags to mark the keywords of different projects.

3. The method for searching keywords using a multi-lexicon shared data structure according to claim 2, characterized in that: The method of using different tags to mark keywords of different items includes: Build the end of each keyword; A data structure is hung at the end of each keyword; the data structure is used to mark the item information.

4. The method for searching keywords using a multi-lexicon shared data structure according to claim 1, characterized in that: The keywords for obtaining a single item include: Get keywords for a single item; Preprocessing the acquired keywords; Extract keywords from preprocessed keywords.

5. The method for searching keywords using a multi-lexicon shared data structure according to claim 4, characterized in that: The preprocessing of the acquired keywords includes: Remove duplicates, merge synonyms, and handle inflection.

6. The method for searching keywords using a multi-lexicon shared data structure according to any one of claims 1 to 5, characterized in that: The keyword search using the third dictionary tree includes: Construct a unified query interface in the third dictionary tree; Perform keyword search through the unified query interface.

7. A system for searching keywords using a common data structure of a vocabulary library, characterized in that: The system is used to implement the method for querying keywords using a multi-lexicon shared data structure as described in any one of claims 1 to 6; the system comprises: A shared data structure unit is used to store keyword information in multiple word libraries, and to integrate keywords in multiple word libraries into one dictionary tree by constructing a dictionary tree; Process keywords, remove duplicates, merge synonyms, and handle word form changes; A keyword query unit, used to traverse all nodes of the dictionary tree according to the character sequence of the keyword and search for matching keywords; A result returning unit, used to feed back the search results of the keyword query unit to the user; The word library updating unit is used to update the word library of the shared data structure unit and modify the dictionary tree in the shared data structure unit.

8. An electronic device, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method for querying keywords using a multi-lexicon shared data structure as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Dictionary tree-based search processing method and device, equipment and storage medium

    CN111460311A

  • New word discovery method and device based on left and right information entropy and mutual information

    CN114330336A