Property noun prompting method, system and equipment and storage medium

By building a dictionary tree and real-time query to match proper nouns, the problems of slow input speed and high error rate of proper nouns are solved, and the efficiency and quality of text editing are improved.

CN120491839APending Publication Date: 2025-08-15SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510559491.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

During the text editing process, memorizing and entering a large number of proper nouns leads to slow editing speed and prone to errors, affecting file quality.

Method used

Build a dictionary tree, listen to the characters entered by the user in real time, query matching proper nouns from the pre-constructed dictionary tree, and use a double-ended queue to cache characters, update the dictionary tree according to the predecessor and successor relationship between characters, providing auxiliary editing of proper nouns.

Benefits of technology

It realizes rapid matching and auxiliary editing of proper nouns, improves the efficiency and quality of file editing, and simplifies the input process of proper nouns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491839A_ABST
    Figure CN120491839A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, and particularly provides a proper noun prompting method, system and device and a storage medium, and the method comprises the steps that a target character field is dynamically intercepted from edited text content, and the target character field comprises a plurality of characters; querying a plurality of proper nouns matched with the target character field from a pre-constructed dictionary tree; calculating the similarity between the target character field and the plurality of proper nouns, and screening out the highest similarity; if the highest similarity reaches a set similarity threshold value, popping up a window for displaying a proper noun corresponding to the highest similarity at a corresponding position according to the position of a target character field; the dictionary tree comprises a plurality of nodes, and the nodes comprise characters, subsequent node pointer data and entry ID arrays ending with the nodes. According to the method and the device, auxiliary editing of the proper nouns is realized, and the efficiency and the quality of editing the file with the proper nouns are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a proper noun prompting method, system, device and storage medium. Background Art

[0002] With the development of information technology, people increasingly use electronic devices for text input in their daily communications. However, some articles often contain a large number of proper nouns. This large number of proper nouns, covering a wide range of fields, including historical figures, place names, and institutional names, can make memorization relatively challenging. Furthermore, some proper nouns may have multiple spellings or abbreviations, further increasing the difficulty of memorization.

[0003] Therefore, when editing some files containing proper nouns, not only is the editing speed slow, but there are also cases of proper noun errors, resulting in reduced file quality. Summary of the Invention

[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a proper noun prompting method, system, device and storage medium to solve the above-mentioned technical problems.

[0005] In a first aspect, the present invention provides a method for prompting a proper noun, comprising: Dynamically intercepting a target character segment from the edited text content, wherein the target character segment includes a plurality of characters; Querying a pre-built dictionary tree for a plurality of proper nouns matching the target character segment; Each node in the dictionary tree represents a character in a word. There is a predecessor and successor relationship between adjacent characters of the same proper noun. The nodes include: Character field, used to store the word represented by the node; The associated field is used to store the array of successor node pointers of the node, and the associated field array is sorted by node characters to maintain order when inserting new successor nodes; The terminator field is used to store the array of term IDs ending with this node.

[0006] In an optional embodiment, a target character segment is dynamically intercepted from the edited text content, wherein the target character segment includes a plurality of characters, including: Monitor the latest characters input by the user in real time; Querying the dictionary tree for a node matching the character; If a node matching the character is found, the character is cached in a double-ended queue; If no node matching the character is found, the character is ignored.

[0007] In an optional embodiment, the method further comprises: If the character is a punctuation mark, or the character is the end character of a proper noun, the double-ended queue is cleared.

[0008] In an optional embodiment, the method for updating the dictionary tree includes: Parsing the characters and character arrangement order of the noun to be updated, and determining the predecessor and successor relationships between the characters based on the character arrangement order; If all characters of the noun to be updated exist in the dictionary tree, updating the associated fields of the corresponding characters according to the predecessor and successor relationships between the characters; If all characters of the noun to be updated do not exist in the dictionary tree, differential characters of the noun to be updated are created in the dictionary tree, where the differential characters are characters of the noun to be updated that do not exist in the dictionary tree.

[0009] In an optional embodiment, updating the associated fields of corresponding characters according to the predecessor and successor relationships between the characters includes: The insertion position is determined by binary search, and the pointer of the node where the successor character of the current character is located is inserted into the corresponding insertion position of the associated field of the node where the current character is located.

[0010] In an optional embodiment, the method further comprises: Receiving a click command from the user on the window; The target character segment is deleted, and the proper noun in the window is written into the position of the target character segment.

[0011] In a second aspect, the present invention provides a proper noun prompt system, comprising: An interception module, configured to dynamically intercept a target character segment from the edited text content, wherein the target character segment includes a plurality of characters; A query module, configured to query a pre-built dictionary tree for a plurality of proper nouns matching the target character segment; Each node in the dictionary tree represents a character in a word. There is a predecessor and successor relationship between adjacent characters of the same proper noun. The nodes include: Character field, used to store the word represented by the node; The associated field is used to store the array of successor node pointers of the node, and the associated field array is sorted by node characters to maintain order when inserting new successor nodes; The terminator field is used to store the array of term IDs ending with this node.

[0012] According to a third aspect, a device is provided, comprising: a memory for storing a proper noun prompt program; The processor is configured to implement the steps of the proper noun prompt method provided in the first aspect when executing the proper noun prompt program.

[0013] In a fourth aspect, a computer-readable storage medium is provided, on which a proper noun prompt program is stored. When the proper noun prompt program is executed by a processor, the steps of the proper noun prompt method provided in the first aspect are implemented.

[0014] The beneficial effect of the present invention is that the proper noun prompt method, system, device and storage medium provided by the present invention can quickly match proper nouns or proper noun fragments by redundantly constructing a proper noun dictionary tree, thereby meeting the real-time requirements of text editing, and displaying the matching proper nouns through a window, thereby realizing auxiliary editing of proper nouns and improving the efficiency and quality of file editing with proper nouns.

[0015] In addition, the present invention has a reliable design principle, a simple structure and a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.

[0018] Figure 2 1 is a schematic diagram of the structure of a dictionary tree of a method according to an embodiment of the present invention.

[0019] Figure 3 2 is a schematic diagram of updating a dictionary tree according to a method of an embodiment of the present invention.

[0020] Figure 4 This is a schematic diagram of a proper noun query method according to an embodiment of the present invention.

[0021] Figure 5 This is a schematic diagram of a first proper noun query scenario of a method according to an embodiment of the present invention.

[0022] Figure 6 This is a schematic diagram of a second proper noun query scenario of a method according to an embodiment of the present invention.

[0023] Figure 7 FIG. 4 is a schematic block diagram of a system according to an embodiment of the present invention.

[0024] Figure 8A schematic structural diagram of a device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0027] The proper noun prompting method provided by the embodiment of the present invention is executed by a computer device. Accordingly, the proper noun prompting system runs in the computer device.

[0028] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention. Figure 1 The execution subject can be a proper noun prompt system. According to different requirements, the order of the steps in the flow chart can be changed, and some steps can be omitted.

[0029] like Figure 1 As shown, the method includes: S1. Dynamically intercepting a target character segment from the edited text content, wherein the target character segment includes multiple characters; S2. Querying multiple proper nouns that match the target character segment from a pre-built dictionary tree; Among them, each node in the dictionary tree represents a character in a word, and there is a predecessor and successor relationship between adjacent characters of the same proper noun; the node includes: a character field, which is used to store the character represented by the node; an association field, which is used to store an array of successor node pointers of the node, and the association field array is sorted by node characters to maintain order when inserting a new successor node; a terminator field, which is used to store an array of entry IDs ending with the node.

[0030] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.

[0031] S101. Monitor the latest characters input by the user in real time.

[0032] Set up a dedicated input listening mechanism in the user interaction layer of the application. For example, in graphical interface applications, find controls such as text boxes and text areas that receive user input. Register an event listener for these controls. When the user enters content in the control, the listener will be triggered. Once triggered, the system immediately captures the most recently entered character and uses it as the input data source for subsequent processing. In mobile applications, the input event interface provided by the operating system is used to connect to user operations on the keyboard and accurately obtain the characters entered after each key press. In command-line applications, the standard input stream is continuously read, and whenever new input enters, the most recently entered characters are parsed.

[0033] S102. Query the dictionary tree for a node that matches the character.

[0034] Clarify the data structure of the dictionary tree. Each node contains multiple child node pointers. These pointers are set according to different character sets. For example, for English characters, there may be 26 pointers corresponding to different letters. Start the query from the root node of the dictionary tree, traverse the child nodes of the current node, and check whether there is a child node whose identification character is the same as the most recently entered character. During the traversal process, a sequential search or a hash table-optimized search method can be used (if hash-related technology is used when constructing the dictionary tree node). If a matching child node is found, a reference to the child node is obtained for subsequent operations; if no matching character is found after traversing all child nodes, it means that there is no node in the dictionary tree that matches the character.

[0035] S103. If a node matching the character is found, the character is cached in a double-ended queue.

[0036] If a node matching the input character is successfully found in the dictionary tree, a double-ended queue data structure is used. Double-ended queues typically have the ability to insert and delete elements at both ends. Here, we choose to use the tail-end insertion operation of the double-ended queue (in most double-ended queue implementations, the time complexity of the insertion operation is O (1)). The most recently entered and successfully matched character is added as a new element to the tail of the double-ended queue. The purpose of this is to maintain a character cache queue arranged in the order of input, so that more complex processing may be performed based on the cached characters later, such as generating associated words.

[0037] S104. If no node matching the character is found, the character is ignored.

[0038] If, during a trie query, no node matches the input character, no additional processing is performed on that character. This means the system neither stores the character in the cache nor initiates any additional processing logic based on it. Instead, the system simply skips the character and waits for the next input from the user, repeating the entire process from listening for input to querying the trie.

[0039] S105. If the character is a punctuation mark, or the character is the end character of a proper noun, clear the double-ended queue.

[0040] To determine whether an input character is a punctuation mark, the system needs to predefine a punctuation mark set, which contains common punctuation marks such as periods, commas, exclamation marks, etc. After obtaining the input character, it is compared one by one with the elements in the punctuation mark set (the comparison can be done by traversing the set). If they match, it is determined to be a punctuation mark. To determine whether it is the end character of a proper noun, this depends on the design of the dictionary tree node. A flag may be set in the node. When the dictionary tree is constructed, the flag is set to a specific value (such as True) for the node corresponding to the last character of the proper noun. When the node flag queried for the input character is this specific value, it is considered to be the end character of the proper noun. Once it is determined that the input character meets one of the above two conditions, the double-ended queue of the cached characters is cleared to remove all characters stored in the queue in preparation for the next possible character input cache.

[0041] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.

[0042] Each node in the dictionary tree represents a character in a word. There is a predecessor and successor relationship between adjacent characters of the same proper noun. The node includes: a character field, which is used to store the character represented by the node; an association field, which is used to store an array of successor node pointers of the node, and the association field array is sorted by node characters to maintain order when inserting new successor nodes; a terminator field, which is used to store an array of entry IDs ending with the node.

[0043] For details, please refer to the structure of the dictionary tree Figure 2 ,include: Each node in the tree represents a character in a word. For two adjacent characters in a word, the first character is the predecessor node of the second character, and the second character is the successor node of the first character.

[0044] A node in the tree consists of the word represented by the node (data field), an array of pointers to the successor nodes of the node (nexts field), and an array of entry IDs ending with the node (ends_vocabulary).

[0045] Each node has any number of successor nodes, and the pointers to the successor nodes are stored in the nexts array. When adding a successor node B to a certain node A, the pointer of B is inserted into the nexts array of node A, and it is necessary to ensure that the nexts array is always sorted.

[0046] The elements in the Nexts array are sorted according to the characters of the nodes pointed to by the elements. When inserting a pointer to a successor node into the nexts array, the index position of the pointer to the successor node to be inserted in the nexts array can be determined by the binary search method.

[0047] Ensuring that the nexts array is sorted can quickly locate the position of a specific successor node of a certain node by the binary search method when searching for it.

[0048] Each node includes a list of entries ending with the node (ends_vocabulary field). This field is used to quickly retrieve the complete entries ending with the node. The reasons for needing this field include two aspects: on the one hand, since it is possible that the prefix of a proper noun is also a proper noun, the entry ending node is not necessarily at the leaf node of the tree, and this field is needed to identify the situation where an intermediate node is the entry ending node (for example, when both "World Trade" and "World Trade Organization" are inserted into the trie as proper nouns, the node of the character "Yi" is also an ending node but not at the leaf node); on the other hand, since it is convenient to search for candidate proper nouns through incomplete fragments, when inserting a certain proper noun, the incomplete words (any suffixes) after removing any prefixes of the proper noun are also inserted. Therefore, the word matched starting from the root node cannot be guaranteed to be a complete word, so the ending node needs to record the original entry ID.

[0049] The update method of the trie includes: 1. Analyze the characters and the character arrangement order of the noun to be updated, and determine the predecessor and successor relationships between the characters based on the character arrangement order.

[0050] Input: The entry to be inserted (such as "World Trade") and its unique ID (such as id = 1).

[0051] Operations: Decompose the entry into a list of characters (such as ["Shi", "Jie", "Mao", "Yi"]).

[0052] Determine the predecessor-successor relationship between the characters: "World" → "Trade" (complete path).

[0053] Insert all suffixes (such as "Trade", "Trade", "Trade") at the same time to support incomplete fragment matching.

[0054] 2. If all characters of the noun to be updated do not exist in the trie tree, create the different characters of the noun to be updated in the trie tree, and the different characters are the characters of the noun to be updated that do not exist in the trie tree.

[0055] Traverse the trie tree: start from the root node and match layer by layer in character order.

[0056] Branch logic: Case 1: The character already exists: directly reuse the existing node.

[0057] Update the ends_vocabulary field of the node and add the current entry ID.

[0058] Execute step 3 to maintain the order of the nexts array.

[0059] Case 2: The character does not exist: Create a new node, initialize data (character), nexts (empty array), ends_vocabulary (including the current entry ID).

[0060] Insert the pointer of the new node into the nexts array of the parent node in character order.

[0061] 3. If all characters of the noun to be updated exist in the trie tree, update the associated fields of the corresponding characters according to the predecessor and successor relationships between the characters; Determine the insertion position by binary search, and insert the pointer of the node where the successor character of the current character is located into the corresponding insertion position of the associated field of the node where the current character is located.

[0062] Order of the nexts array: <第 0000176 行>When inserting a new successor node, determine the insertion position by binary search.

[0063] For example: the nexts array of the parent node currently stores ["A", "C"], and "B" needs to be inserted, then it is inserted at the position of index 1. <00001'80>Update of ends_vocabulary: Each node records all the entry IDs that end with this node, supporting overlapping words (such as "World Trade" and "World Trade Organization" sharing a prefix).

[0065] In a specific example, the entries "World Trade" (id = 1) and "World Trade Organization" (id = 2) are inserted into an empty trie tree.

[0066] Step 1: Insert "World Trade" (id = 1). Insert the complete entry and all its suffixes: "World Trade", "orld Trade", "Trade", "rade".

[0067] Node changes: Root node → "Wor" → "orld" → "rld" → "ld".

[0068] The ends_vocabulary of each node contains id = 1 (since all suffixes belong to entry 1).

[0069] Step 2: Insert "World Trade Organization" (id = 2). Insert the complete entry and all its suffixes: "World Trade Organization", "orld Trade Organization", "Trade Organization", "rade Organization", "Organization", "rganization".

[0070] Node changes: Reuse the existing path: Root → "Wor" → "orld" → "rld" → "ld".

[0071] Create a new path: Add new nodes "Gro" → "rld" after the "ld" node.

[0072] Update the associated fields: The ends_vocabulary of the "ld" node adds id = 2 (since "ld" is the suffix end of "World Trade Organization").

[0073] The ends_vocabulary of the "Gro" and "rld" nodes contains id = 2.

[0074] Please refer to Figure 3 In another scenario of inserting proper nouns into a trie tree, including: (1) Record the position of the predecessor node, which is the root node at the beginning. <​​​​​​​​

[0078] (5) If the character is not found, construct a tree node with the current character, insert the node into the nexts array of the predecessor node, at the insertion position recorded in the binary search step, and record the newly inserted node as the predecessor node.

[0079] (6) When inserting the last character, record the entry ID into the ends_vocabulary array of the node corresponding to the last character. (7) Not only complete and correct proper nouns need to be inserted, but also incomplete proper nouns with prefixes removed need to be inserted for quick fragment matching. For example, if the complete entry is "Proper Noun", in addition to inserting "Proper Noun", "roper Noun", and "Noun" also need to be inserted.

[0080] Please refer to Figure 4 , and query multiple proper nouns that match the target character segment from the pre-constructed trie.

[0081] The query idea includes: starting from the root node, layer by layer matching the characters of the target segment. If any character in the middle does not match, immediately terminate; after successful matching, traverse all subsequent paths starting from the last character of the target segment to collect complete entries; utilize the orderliness of the trie (the nexts array is sorted by characters), and accelerate the search through binary search to avoid ineffective traversal. Specifically, it includes the following steps: 1. Layer-by-layer matching of the target segment. Starting from the root node, for each character char in the target segment (such as "Trade", "Easy"), match the target segment character by character: Use the binary search method to locate the character position in the nexts array of the parent node. If node.nexts[idx].data == char, move to the next lower-level node. If the character is not found, directly return an empty list.

[0082] 2. Collect candidate entries.

[0083] Initialize the traversal stack, which stores the current node and the matched path (used to generate complete entries, but in the actual scenario, only the IDs need to be collected).

[0084] Use the breadth-first search (BFS) method to traverse all subsequent paths starting from the last character of the target segment.

[0085] 3. Result deduplication and mapping.

[0086] ID reverse lookup, convert the entry ID to the original entry through an external mapping table.

[0087] If the same entry is repeated due to the insertion of multiple suffixes, use a set to deduplicate.

[0088] In an example query, the following steps are involved: 1. Query the target segment "trade": Layer-by-layer matching: Root → "world" → "world" → "trade" → "yi" (completely matches the target segment).

[0089] Collect candidate IDs: Starting from the "Yi" node, traverse the subsequent path ("Yi" → "Group" → "Organization").

[0090] Collection IDs: [1,2] (from the "easy" and "weave" nodes).

[0091] Mapping and deduplication: id=1→"World Trade", id=2→"World Trade Organization".

[0092] Final result: ["world trade","World Trade Organization"].

[0093] 2. Query the target segment "Yi Group": Layer-by-layer matching: Root → "world" → "world" → "trade" → "yi" → "group" (completely matches the target fragment).

[0094] Collect candidate IDs: Starting from the "Group" node, traverse the subsequent path ("Group" → "Organize").

[0095] Collected ID: [2] (from "weave" node only).

[0096] Mapping results: Final result: ["World Trade Organization"].

[0097] In one embodiment of the present invention, the candidate distinguished name search method includes: If the input prefix is correct: Starting from the root node, search each character until all characters in the prefix are matched. Then, recursively traverse all subsequent nodes and obtain all entries in the "ends_vocabulary" of the subsequent nodes as candidate proper nouns. The candidate proper noun list can be returned to complete the search.

[0098] If the input prefix contains extra characters, missing characters, or missing characters: Match each character of the input word sequentially from the dictionary tree, and stop matching when a certain character is encountered.

[0099] If a word that cannot be matched is encountered and <2 nodes have been matched, the remaining words will be used as input to search again.

[0100] When encountering a character that cannot be matched and having already matched >= 2 nodes (a word or part of a word), recursively traverse the successor nodes and obtain all the entries in ends_vocabulary of the successor nodes as part of the candidate proper noun list.

[0101] Then use the remaining words as the new input to continue searching the candidate proper noun list.

[0102] Since when constructing the dictionary tree, the proper nouns have been redundantly inserted after removing any prefixes, it is possible to discard any number of (incorrect) characters in front of the proper noun and complete the search and matching of the candidate proper noun relying only on the middle partial string.

[0103] Please refer to Figure 5 , Example of complete prefix matching: When the input is "zhuan you" (专有), perform binary search layer by layer from the root node, and it is possible to quickly match the first two green nodes in the figure. Then traverse all the successor nodes and return all the entry IDs in the ends_vocabulary array of the successor nodes as the candidate proper nouns.

[0104] Please refer to Figure 6 , Example of incomplete prefix matching: When the input is "zhuan you ming" (转有名) or "zhuan g you ming" (专g有名) or "you ming" (有名), since the prefix before "you ming" cannot match a partial word in the dictionary tree, it can be directly skipped. Continue to match the subsequent string "you ming". Since when inserting "zhuan you ming ci" (专有名词), "you ming ci" (有名词) without the prefix has also been inserted. So it is still possible to match the candidate proper noun.

[0105] In one embodiment of the present invention, based on step S3, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.

[0106] One input may search for multiple candidate proper nouns. Therefore, it is necessary to sort them according to the similarity between the candidate proper nouns and the input word. It is possible to sort them according to the length of the longest common subsequence (LCS) between the candidate proper noun and the input. The longest common subsequence is obtained using the classical dynamic programming algorithm.

[0107] S301. Input and output. Input: Target segment (such as "trade"), candidate word list (such as ["World Trade", "World Trade Organization"]); Output: Candidate word list sorted in descending order of LCS length (such as ["World Trade", "World Trade Organization"]).

[0108] S302. Calculate the LCS length.

[0109] Principle of dynamic programming: Dynamic programming avoids repeated computation by breaking down large problems into smaller ones and saving the solutions to the smaller ones. To calculate the LCS length of two strings, let the target segment length be m and the candidate word length be n. Create a two-dimensional array dp of size (m+1)x(n+1) to store the solutions to the subproblems.

[0110] dp[i][j] represents the LCS length of the first i characters of the target segment and the first j characters of the candidate word.

[0111] initialization: When i = 0 or j = 0, it means that one of the strings is empty, and the LCS length is 0. Therefore, the first row and the first column of the dp array are initialized to 0, that is, dp[0][j] = 0 (0 <= j <= n) and dp[i][0] = 0 (0 <= i <= m).

[0112] State transition equation: If the i-th character (index starts from 1) of the target segment is the same as the j-th character of the candidate word, then dp[i][j] = dp[i-1][j-1] + 1. This is because the current characters are the same and they can be part of the LCS, so the LCS length increases by 1 based on the previous subproblem.

[0113] If the i-th character of the target segment is different from the j-th character of the candidate word, then dp[i][j] = max(dp[i-1][j],dp[i][j-1]). That is, the maximum LCS length is taken when the i-th character of the target segment is removed or the j-th character of the candidate word is removed.

[0114] Calculation process: Traverse the dp array from i=1 to m, j=1 to n, and fill the array according to the state transition equation. Finally, dp[m][n] is the LCS length of the target segment and the candidate word.

[0115] S303. Sort candidate words.

[0116] Sort by LCS length in descending order. If the lengths are the same, sort by candidate length in ascending order (or keep the original order).

[0117] Sorting rules First, sort the candidate words in descending order by LCS length, i.e., candidates with larger LCS lengths are ranked higher. This can be achieved by using the built-in sorting functions of the programming language and customizing the comparison function.

[0118] If two candidate words have the same LCS length, they are sorted in ascending order according to their length, with the shorter candidate word at the front. You can also choose to keep their order in the original list.

[0119] Implementation steps To record the corresponding LCS length for each candidate word, you can use a list of tuples, where each tuple contains a candidate word and its LCS length, for example [(candidate word 1, LCS length 1), (candidate word 2, LCS length 2), ...].

[0120] Use the sorting function to sort the tuple list. According to the custom comparison function, first compare the LCS lengths. If they are the same, then compare the candidate word lengths or keep the original order.

[0121] Finally, candidate words are extracted from the sorted tuple list to form the final sorted result list.

[0122] In an embodiment of the present invention, based on step S4, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.

[0123] S401. Threshold setting and decision logic. Threshold type: supports absolute length (such as LCS ≥ 2) or relative ratio (such as LCS length / input length ≥ 0.8).

[0124] S402. Pop-up interaction design. Trigger timing: Delay 200ms after the user stops typing (for anti-shake processing) to avoid frequent pop-ups.

[0125] Pop-up content: Displays the most similar terms and provides an "Accept" or "Close" button.

[0126] Positioning method: Dynamically calculate the pop-up window position based on the screen coordinates of the input box (such as directly below).

[0127] S403. Proper noun display process. Listen for input events: Capture user input in real time.

[0128] Query candidate words: call the dictionary tree query interface to obtain a list of candidate words.

[0129] Calculate LCS and sort: Sort by LCS length to obtain the terms with the highest similarity.

[0130] Threshold determination: Check whether the maximum LCS length meets the threshold.

[0131] Pop-up control: Show / hide the pop-up window based on the judgment result, and update the position and content.

[0132] In a specific scenario, for example: 1. User enters "trade". Candidate word list: ["world trade", "World Trade Organization"].

[0133] LCS calculation: The LCS length of both words and the input is 2.

[0134] Threshold determination: 2≥2→condition met.

[0135] Pop-up effect: Displayed below the input box, what you are looking for is: World Trade.

[0136] 2. The user continues to enter "Yi Group".

[0137] Candidate word list: ["World Trade Organization"].

[0138] LCS calculation: The LCS length of "Yi Group" and "World Trade Organization" is 2.

[0139] Pop-up effect: Shows that what you are looking for is: World Trade Organization.

[0140] In order to further improve editing efficiency, in one embodiment, upon receiving a user click command on the window, the target character segment is deleted, and the proper noun in the window is written in the position of the target character segment. Specifically, the following steps are included: Event Listening: Add a click event listener to the window to capture user clicks. In web applications, you can use the JavaScript addEventListener method to listen for click events. In desktop applications, different GUI libraries have different event listening mechanisms, such as the bind method in Tkinter.

[0141] Event handling: When the user clicks on the window, the event listener triggers the corresponding event handling function. In the event handling function, you can perform some necessary verification and processing, such as checking whether the clicked position is within the valid area of the window.

[0142] Text deletion: Delete a target character segment from a text based on its location. In text processing, string manipulation functions can be used to implement deletion operations, such as string slicing in Python.

[0143] Text Insertion: Inserts the proper noun in the window into the original location of the target segment. String manipulation functions are also used to implement the insertion operation, ensuring that the proper noun accurately replaces the target segment.

[0144] Interface Update: After completing text deletion and insertion operations, update the text display on the interface so that users can see the latest text content. In web applications, you can update the text display by modifying the content of the DOM element; in desktop applications, use the corresponding GUI library methods to update the content of the text component.

[0145] In some embodiments, the proper noun prompt system may include multiple functional modules composed of computer program segments. The computer program of each program segment in the proper noun prompt system may be stored in a memory of a computer device and executed by at least one processor to perform (see Figure 1 Description) The function of the proper noun prompt.

[0146] In this embodiment, the proper noun prompt system can be divided into multiple functional modules according to the functions it performs, such as Figure 7 As shown. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can perform fixed functions, which are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0147] An interception module, configured to dynamically intercept a target character segment from the edited text content, wherein the target character segment includes a plurality of characters; A query module, configured to query a pre-built dictionary tree for a plurality of proper nouns matching the target character segment; a comparison module, configured to calculate the similarity between the target character segment and the plurality of proper nouns, and select the one with the highest similarity; A display module configured to pop up a window displaying the proper noun corresponding to the highest similarity at a corresponding position according to the position of the target character segment if the highest similarity reaches a set similarity threshold; The dictionary tree includes multiple nodes, each node includes a character, a successor node pointer data, and an entry ID array ending with the node.

[0148] Figure 8The proper noun prompt method provided for the embodiment of the present application can be applied to a device. Those skilled in the art will understand that the device structure involved in the embodiment of the present invention does not constitute a limitation on the device, and the device may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently. In an embodiment of the present invention, the device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.

[0149] The device 800 may include a processor 810, a memory 820, and a communication unit 830. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention. The server structure may be a bus structure or a star structure, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0150] The memory 820 can be used to store execution instructions of the processor 810. The memory 820 can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 820 are executed by the processor 810, the device 800 can perform some or all of the steps in the above-described method embodiments.

[0151] The processor 810 is the control center of the storage device, which uses various interfaces and lines to connect various parts of the entire electronic device. It executes various functions of the electronic device and / or processes data by running or executing software programs and / or modules stored in the memory 820, and calling data stored in the memory. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 810 can only include a central processing unit (CPU). In an embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.

[0152] The communication unit 830 is configured to establish a communication channel so that the storage device can communicate with other devices, receive user data sent by other devices, or send user data to other devices.

[0153] The present invention also provides a computer storage medium, wherein the computer storage medium may store a program that, when executed, may include some or all of the steps of each embodiment provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0154] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software and a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code, and includes instructions for causing a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0155] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0156] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or modules, and can be electrical, mechanical or other forms.

[0157] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of the present embodiment according to actual needs.

[0158] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0159] Although the present invention has been described in detail with reference to the accompanying drawings and in conjunction with preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, persons of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and such modifications or substitutions shall be within the scope of the present invention. Any changes or substitutions that can be easily conceived by persons skilled in the art within the technical scope disclosed in the present invention shall be within the scope of protection of the present invention.

Claims

1. A proper noun prompting method, characterized in that: include: Dynamically intercepting a target character segment from the edited text content, wherein the target character segment includes a plurality of characters; Querying a pre-built dictionary tree for proper nouns matching the target character segment; Each node in the dictionary tree represents a character in a word. There is a predecessor and successor relationship between adjacent characters of the same proper noun. The nodes include: Character field, used to store the word represented by the node; The associated field is used to store the array of successor node pointers of the node, and the associated field array is sorted by node characters to maintain order when inserting new successor nodes; The terminator field is used to store the array of term IDs ending with this node.

2. The method according to claim 1, characterized in that The method further comprises: Calculating similarities between the target character segment and the plurality of proper nouns, and selecting the one with the highest similarity; If the highest similarity reaches a set similarity threshold, a window displaying the proper noun corresponding to the highest similarity will pop up at the corresponding position according to the position of the target character segment; The dictionary tree includes multiple nodes, each node includes a character, a successor node pointer data, and an entry ID array ending with the node.

3. The method according to claim 1, characterized in that Dynamically intercept a target character segment from the edited text content, wherein the target character segment includes multiple characters, including: Monitor the latest characters input by the user in real time; Querying the dictionary tree for a node matching the character; If a node matching the character is found, the character is cached in a double-ended queue; If no node matching the character is found, the character is ignored.

4. The method according to claim 3, characterized in that The method further comprises: If the character is a punctuation mark, or the character is the end character of a proper noun, the double-ended queue is cleared.

5. The method according to claim 4, characterized in that The method for updating the dictionary tree includes: Parsing the characters and character arrangement order of the noun to be updated, and determining the predecessor and successor relationships between the characters based on the character arrangement order; If all characters of the noun to be updated exist in the dictionary tree, updating the associated fields of the corresponding characters according to the predecessor and successor relationships between the characters; If all characters of the noun to be updated do not exist in the dictionary tree, differential characters of the noun to be updated are created in the dictionary tree, where the differential characters are characters of the noun to be updated that do not exist in the dictionary tree.

6. The method according to claim 5, characterized in that Update the associated fields of the corresponding characters based on the predecessor and successor relationships between the characters, including: The insertion position is determined by binary search, and the pointer of the node where the successor character of the current character is located is inserted into the corresponding insertion position of the associated field of the node where the current character is located.

7. The method according to claim 1, characterized in that The method further comprises: Receiving a click command from the user on the window; The target character segment is deleted, and the proper noun in the window is written into the position of the target character segment.

8. A proper noun prompt system, characterized in that: include: An interception module, configured to dynamically intercept a target character segment from the edited text content, wherein the target character segment includes a plurality of characters; A query module, configured to query a pre-built dictionary tree for a plurality of proper nouns matching the target character segment; Each node in the dictionary tree represents a character in a word. There is a predecessor and successor relationship between adjacent characters of the same proper noun. The nodes include: Character field, used to store the word represented by the node; The associated field is used to store the array of successor node pointers of the node, and the associated field array is sorted by node characters to maintain order when inserting new successor nodes; The terminator field is used to store the array of term IDs ending with this node.

9. A device, characterized in that include: a memory for storing a proper noun prompt program; A processor is configured to implement the steps of the proper noun prompt method according to any one of claims 1 to 7 when executing the proper noun prompt program.

10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a proper noun prompting program, and when the proper noun prompting program is executed by a processor, the steps of the proper noun prompting method according to any one of claims 1 to 7 are implemented.