String processing method, device, electronic device and storage medium

By mapping the multi-byte characters of the target language into unique identification numbers and storing these identification numbers using nodes of preset data types in the dictionary tree, the problem of large memory overhead in the prior art is solved, and more efficient string processing is achieved.

CN114880523BActive Publication Date: 2025-05-16UBTECH ROBOTICS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210450193.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-27
Publication Date
2025-05-16
Estimated Expiration
2042-04-27

AI Technical Summary

Technical Problem

When processing multibyte characters, existing dictionary trees require multiple nodes to store information of one character, resulting in a large memory overhead.

Method used

Efficient storage of multi-byte characters is achieved by mapping each character in the pending string into a unique identification number and storing these identification numbers using nodes of preset data types in the target dictionary tree.

Benefits of technology

Reduces the number of nodes required to store each character, saves memory space, and improves the efficiency of string processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114880523B_ABST
    Figure CN114880523B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of computer technology, and provides a string processing method, device, electronic device and storage medium, including: obtaining a string to be processed; wherein the string to be processed is a string in a target language, and the number of bytes occupied by each character in the string of the target language is greater than 1; determining the unique identification number corresponding to each character in the string to be processed according to a preset target mapping relationship; wherein the target mapping relationship includes a mapping relationship between each character of the target language and the unique identification number, and the data type of the unique identification number is a preset data type; according to the unique identification number corresponding to each character of the string to be processed and a target dictionary tree, the string to be processed is processed to obtain a target processing result; wherein the data type of each node in the target dictionary tree is the preset data type. The embodiment of the present application can save storage space when performing string processing based on a dictionary tree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and in particular, relates to a string processing method, device, electronic device and storage medium. Background Art

[0002] A dictionary tree, also known as a word search tree, is a tree structure and a variant of a hash tree. A dictionary tree can be used to count, sort and save a large number of strings and is often used in search engine systems. Its advantages are: using the common prefix of the string to reduce the query time, minimizing repeated string comparisons, and having higher query efficiency than a hash tree.

[0003] At present, according to the existing character encoding rules, the number of bytes occupied by the characters corresponding to some language texts is greater than 1, and each node of the existing dictionary tree can only store one byte of data, resulting in the need to store the information of one character by multiple nodes, which has a large memory overhead. Summary of the invention

[0004] In view of this, embodiments of the present application provide a string processing method, device, electronic device and storage medium to solve the problem of how to save storage space when performing string processing based on a dictionary tree in the prior art.

[0005] A first aspect of an embodiment of the present application provides a string processing method, including:

[0006] Obtaining a character string to be processed; wherein the character string to be processed is a character string in a target language, and the number of bytes occupied by each character in the character string in the target language is greater than 1;

[0007] Determine the unique identification number corresponding to each character in the character string to be processed according to a preset target mapping relationship; wherein the target mapping relationship includes a mapping relationship between each character of the target language and the unique identification number, and the data type of the unique identification number is a preset data type;

[0008] The character string to be processed is processed according to the unique identification numbers corresponding to the respective characters of the character string to be processed and the target dictionary tree to obtain a target processing result; wherein the data type of each node in the target dictionary tree is the preset data type.

[0009] Optionally, the string to be processed includes a string to be stored, and the processing of the string to be processed according to the unique identification numbers corresponding to the respective characters of the string to be processed and the target dictionary tree to obtain a target processing result includes:

[0010] According to the unique identification numbers corresponding to the respective characters of the string to be stored, the unique identification numbers are stored in the nodes of the target dictionary tree to obtain a string storage result.

[0011] Optionally, the string to be processed includes a string to be searched, and the processing of the string to be processed according to the unique identification numbers corresponding to the respective characters of the string to be processed and the target dictionary tree to obtain a target processing result includes:

[0012] According to the unique identification numbers corresponding to the respective characters of the string to be searched, node indexing is performed level by level in the target dictionary tree to obtain a string search result.

[0013] Optionally, the character string to be processed includes a character string to be deleted, and the character string to be processed is processed according to the unique identification number corresponding to each character of the character string to be processed and the target dictionary tree to obtain a target processing result, including:

[0014] According to the unique identification numbers corresponding to the respective characters of the string to be deleted, the first target node corresponding to the string to be deleted is searched in the target dictionary tree, and the string deletion result is obtained by modifying or deleting the word mark of the first target node.

[0015] Optionally, the string to be processed includes a string to be modified, and the processing of the string to be processed according to the unique identification numbers corresponding to the respective characters of the string to be processed and the target dictionary tree to obtain a target processing result includes:

[0016] According to the unique identification numbers corresponding to the respective characters of the string to be modified, a second target node corresponding to the string to be modified is searched in the target dictionary tree, and the string modification result is obtained by modifying the information of the second target node.

[0017] Optionally, before obtaining the character string to be processed, the method further includes:

[0018] Determine the preset data type based on the number of characters contained in the target language;

[0019] The target mapping relationship is obtained by respectively setting a data type of a unique identification number of the preset data type for each character of the target language.

[0020] Optionally, the target language is any one of Chinese, Korean and Japanese, and the preset data type is a short integer.

[0021] A second aspect of an embodiment of the present application provides a string processing device, including:

[0022] An acquisition unit, configured to acquire a character string to be processed; wherein the character string to be processed is a character string in a target language, and the number of bytes occupied by each character in the character string in the target language is greater than 1;

[0023] A mapping unit, configured to determine a unique identification number corresponding to each character in the character string to be processed according to a preset target mapping relationship; wherein the target mapping relationship includes a mapping relationship between each character of the target language and the unique identification number, and the data type of the unique identification number is a preset data type;

[0024] A processing unit is used to process the string to be processed according to the unique identification numbers corresponding to the respective characters of the string to be processed and a target dictionary tree to obtain a target processing result; wherein the data type of each node in the target dictionary tree is the preset data type.

[0025] A third aspect of an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the steps of the string processing method.

[0026] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, an electronic device implements the steps of the string processing method.

[0027] A fifth aspect of the embodiments of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes the string processing method described in any one of the first aspects.

[0028] Compared with the prior art, the embodiments of the present application have the following beneficial effects: in the embodiments of the present application, a character string to be processed in the target language is obtained. Since the number of bytes occupied by each character in the character string of the target language is greater than 1, each character in the character string to be processed is first mapped to a unique identification number whose data type is a preset data type according to a preset target mapping relationship. After that, the character string to be processed is processed according to the unique identification number corresponding to each character of the character string to be processed and the target dictionary tree to obtain a target processing result. Since the data type of each node in the target dictionary tree is a preset data type, a node of the target dictionary tree can correspond to a unique identification number of a preset data type. Compared with the current method of requiring multiple nodes of the dictionary tree to jointly store information of the same character of the character string, it can save memory space when processing the character string based on the dictionary tree and improve the efficiency of character string processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.

[0030] Figure 1 This is a schematic diagram of an implementation flow of a string processing method provided in an embodiment of the present application;

[0031] Figure 2 is a schematic diagram of a target dictionary tree provided in an embodiment of the present application;

[0032] Figure 3 is a schematic diagram of another target dictionary tree provided in an embodiment of the present application;

[0033] Figure 4 is a schematic diagram of a string processing device provided in an embodiment of the present application;

[0034] Figure 5 It is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0035] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0036] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.

[0037] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0038] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0039] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0040] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0041] In addition, in the description of the present application, the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0042] In order to better illustrate the string processing method of the embodiment of the present application, some relevant concepts of the embodiment of the present application are explained below:

[0043] 1. TrieTree

[0044] In the dictionary tree, each leaf node corresponds to a hash table. The key of the hash table is the character stored in the node, and the value is a dictionary tree subtree corresponding to the node (including all child nodes of the node, that is, all suffix characters after the character). In the current dictionary tree, the data type of each node key value is char (character type), which is only one byte.

[0045] Generally, a dictionary tree has the following three basic properties:

[0046] 1) The root node (root) does not contain any characters, and each child node except the root node contains one character.

[0047] 2) From the root node to a certain node, the characters on the path are connected to form the string corresponding to the node.

[0048] 3) All child nodes of each node contain different characters.

[0049] 2. UTF-8 encoding rules

[0050] UTF-8 (8-bit Unicode) is a variable-length character encoding for Unicode. It can be used to represent any character in Unicode (Uniform Character Encoding Standard), and the first byte in its encoding is still compatible with ASCII (American Standard Code for Information Interchange), so that the original software that processes ASCII characters can continue to be used without or with only a few modifications.

[0051] Usually, only 1 byte is needed to encode English characters; 2 bytes are needed to encode characters with diacritical marks (such as characters in Latin, Greek, Arabic, etc.); 3 bytes are usually needed to encode characters in languages ​​of Chinese, Japanese, Korean, Southeast Asian, Middle Eastern and other countries; and 4 bytes are needed to encode characters in some other local languages.

[0052] When encoding characters using the UTF-8 encoding rule, characters in some languages ​​require 2 bytes to be encoded, and the data type of the key value of each node in the existing dictionary tree is a char type that can only store one byte of information. Therefore, in the character storage of some languages, multiple nodes of the dictionary tree are required to complete the storage of one character. For example, for Chinese, Japanese, and Korean languages, three nodes are required to complete the storage of one character. In this case, the storage space is greatly increased and the efficiency of string processing is reduced.

[0053] In order to solve the above-mentioned technical problems, an embodiment of the present application provides a string processing method, device, electronic device and storage medium, including: obtaining a string to be processed; wherein the string to be processed is a string in a target language, and the number of bytes occupied by each character in the string of the target language is greater than 1; according to a preset target mapping relationship, determining a unique identification number corresponding to each character in the string to be processed; wherein the target mapping relationship includes a mapping relationship between each character of the target language and the unique identification number, and the data type of the unique identification number is a preset data type; according to the unique identification numbers corresponding to each character of the string to be processed and a target dictionary tree, the string to be processed is processed to obtain a target processing result; wherein the data type of each node in the target dictionary tree is the preset data type.

[0054] When processing a character string of a target language in which the number of bytes occupied by each character is greater than 1, each character of the character string can be first converted into a unique identification number whose corresponding data type is a preset data type, and the node data type can be set to a target dictionary tree of the preset data type. Therefore, the storage of a character can be indirectly completed by storing a unique identification number in a node of the target dictionary tree. Compared with the current method of requiring multiple nodes of the dictionary tree to jointly store information of the same character of the character string, this can save memory space when processing the character string based on the dictionary tree and improve the efficiency of string processing.

[0055] Embodiment 1:

[0056] Figure 1 A flow chart of a string processing method provided by an embodiment of the present application is shown, and the details are as follows:

[0057] In S101, a character string to be processed is obtained; wherein the character string to be processed is a character string in a target language, and the number of bytes occupied by each character in the character string in the target language is greater than 1.

[0058] The target language in the embodiment of the present application includes but is not limited to Chinese, Japanese, Korean, Southeast Asian, Middle Eastern, Latin, Greek, Arabic, etc. The characters of these languages ​​are encoded by UTF-8, and the number of bytes occupied by each character is greater than 1.

[0059] The character string to be processed in the embodiment of the present application is the character string in the target language mentioned above. In one embodiment, the character string to be processed can be obtained by receiving information input by a user through an input method in the target language.

[0060] In S102, the unique identification number corresponding to each character in the character string to be processed is determined according to a preset target mapping relationship; wherein the target mapping relationship includes a mapping relationship between each character of the target language and the unique identification number, and the data type of the unique identification number is a preset data type.

[0061] In the embodiment of the present application, the target mapping relationship includes a mapping relationship between each character of the target language and a unique identification number (id) of a preset data type, and the target mapping relationship is set in advance and stored by a mapping function, a mapping table, etc. Exemplarily, the preset data type may be an integer (int), a short integer (short int), etc.

[0062] After the character string to be processed is obtained, for each character of the character string to be processed, the preset target mapping relationship is searched to determine the unique identification number corresponding to each character in the character string to be processed in sequence.

[0063] In S103, the character string to be processed is processed according to the unique identification numbers corresponding to the respective characters of the character string to be processed and the target dictionary tree to obtain a target processing result; wherein the data type of each node in the target dictionary tree is the preset data type.

[0064] In an embodiment of the present application, the target dictionary tree is a dictionary tree with customized node data types, and the data type of the nodes of the target dictionary tree is consistent with the data type of the above-mentioned unique identification number, that is, in the nodes of the target dictionary tree, the data type of the key used to store character information is a preset data type.

[0065] In one embodiment, each node in the target dictionary tree contains not only its corresponding hash table (the key of the hash table is the character information stored in the node), but also a word mark corresponding to the character of the node. The word mark can include a word end mark (for example, it can be any pre-set mark such as "True", "Yes", "1", "word end", etc.) and a non-end mark (for example, it can be any pre-set mark such as "False", "No", "0", "non-end", etc.); when the word mark corresponding to the character is a word end mark, it means that from the root node to the current node, the characters on the path are connected to form a complete word ending with the character represented by the current node; on the contrary, when the word mark corresponding to the character is a non-end mark, it means that from the root node to the current node, it is impossible to combine to form a word ending with the character of the current node. Exemplarily, in this case, the tree structure of a target dictionary tree for storing Chinese character information is as follows: Figure 2 As shown, the Chinese character in the node only indicates that the node stores the information corresponding to the Chinese character, but does not mean that the node directly stores the UTF-8 encoded characters corresponding to the Chinese character. The node actually stores the unique identification number and word tag corresponding to the Chinese character.

[0066] In another embodiment, the nodes in the target dictionary tree do not contain word markers, the leaf nodes in the target dictionary tree are word end markers, and the non-leaf nodes in the target dictionary tree represent character information. That is, the target dictionary tree in this case connects a node to a leaf node that stores word end information to indicate that the character of the node can be used as a word end, and the characters from the root node to the node can form a complete word. Exemplarily, the tree structure of the target dictionary tree for storing Chinese character information in this case is as follows: Figure 3As shown, the Chinese characters in the non-leaf nodes only indicate that the node stores the information corresponding to the Chinese characters, but does not mean that the node directly stores the UTF-8 encoded characters corresponding to the Chinese characters. The "Chinese end" mark in the leaf node is used to indicate that the characters represented by the parent node of the leaf node can be used as the end of a complete word.

[0067] After determining the unique identification number corresponding to each character in the string to be processed, based on the target dictionary tree, the string to be processed is processed by sequentially processing each unique identification number corresponding to each character of the string to be processed to obtain a target processing result. The processing of the string to be processed may include: storing, searching, deleting, and modifying the string to be processed in the target dictionary tree.

[0068] In an embodiment of the present application, when processing a character string of a target language in which the number of bytes occupied by each character is greater than 1, each character of the character string can be first converted into a unique identification number whose corresponding data type is a preset data type, and the node data type can be set to a target dictionary tree of the preset data type. Therefore, the storage of a character can be indirectly completed by storing a unique identification number in a node of the target dictionary tree. Compared with the current method of requiring multiple nodes of the dictionary tree to jointly store information of the same character of the character string, this can save memory space when processing the character string based on the dictionary tree and improve the efficiency of string processing.

[0069] Optionally, the string to be processed includes a string to be stored, and the processing of the string to be processed according to the unique identification numbers corresponding to the respective characters of the string to be processed and the target dictionary tree to obtain a target processing result includes:

[0070] According to the unique identification numbers corresponding to the respective characters of the string to be stored, the unique identification numbers are stored in the nodes of the target dictionary tree to obtain a string storage result.

[0071] The character string to be processed in the embodiment of the present application is specifically a character string to be stored, that is, the acquired character string needs to be stored at present. For the character string to be stored, after determining the unique identification number corresponding to each single character in the character string to be stored in turn, starting from the root node of the target dictionary tree, the unique identification numbers corresponding to the characters of the character string to be stored are stored in order in each node of the target dictionary tree step by step, and a word end mark is added to the node storing the unique identification number corresponding to the last character or a leaf node for indicating the word end information is connected, thereby completing the storage of the character string to be stored and obtaining the character string storage result. In one embodiment, the character string storage result may include: a result indicating that the character string to be stored is successfully stored, or a result indicating that the character string to be stored fails to be stored; when the character string to be stored is successfully stored, the character string storage result may also include the storage address of the character string to be stored, etc.

[0072] In an embodiment of the present application, by sequentially storing the unique identification number corresponding to each character of the string to be stored in each node of the target dictionary tree, the storage of a character can be indirectly completed through a node of the target dictionary tree. Compared with the current method of requiring multiple nodes of the dictionary tree to jointly store the information of the same character of the string, a large amount of storage space required for storing the string can be saved.

[0073] Optionally, the string to be processed includes a string to be searched, and the processing of the string to be processed according to the unique identification numbers corresponding to the respective characters of the string to be processed and the target dictionary tree to obtain a target processing result includes:

[0074] According to the unique identification numbers corresponding to the respective characters of the string to be searched, node indexing is performed level by level in the target dictionary tree to obtain a string search result.

[0075] The character string to be processed in the embodiment of the present application is specifically a character string to be searched, that is, it is currently necessary to search the target dictionary tree for information on whether the character string to be searched has been stored based on the acquired character string to be searched. For the character string to be searched, after determining the unique identification number corresponding to each character in the character string to be searched in turn, starting from the root node of the target dictionary tree, the nodes of the target dictionary tree are indexed in order to determine whether the unique identification numbers corresponding to the characters sequentially contained in the character string to be searched exist in the nodes of the corresponding level of the target dictionary tree. If it is determined by indexing in order and level by level that the characters passing through the path from the root node of the target dictionary tree to the end node (for example, the node whose word mark is the word end mark) can be connected to form a character string consistent with the character string to be searched, then the current character string search result is determined to be a successful search; otherwise, the current character string search result is determined to be a failed search.

[0076] For example, suppose the current target dictionary tree is as follows Figure 2 As shown, if the string to be searched is "annoyance", after determining the unique identification number corresponding to the word "annoyance" and the unique identification number corresponding to the word "annoyance", start from the root node root of the target dictionary tree and start to search whether there is a node in the first-level child node of the target dictionary tree that has stored the unique identification number corresponding to the word "annoyance"; since the second node of the first-level child node has a unique identification number corresponding to the word "annoyance", it is possible to search whether the child node of the node stores the unique identification number corresponding to the second character "annoyance" of the string to be searched. If it exists, it is determined that the string "annoyance" can be found in the target dictionary tree at present, and the current string search result is determined to be a successful search. If the string to be searched is: "annoyance", "happy", etc., since it is impossible to find the unique identification numbers corresponding to each character of these strings in the target dictionary tree in turn through the above-mentioned level-by-level indexing method, it can be determined that the string search result is a failed search.

[0077] In the embodiment of the present application, since the target dictionary tree only needs to store the information of the unique identification number corresponding to the character, it can indirectly complete the storage of a character through a node of the target dictionary tree. After obtaining the character string to be searched, each character in the character string to be searched is also converted into the corresponding unique identification number, so that the string index can be completed efficiently and accurately in the target dictionary tree, while saving character storage space, ensuring the efficiency and accuracy of string retrieval based on the dictionary tree.

[0078] Optionally, the character string to be processed includes a character string to be deleted, and the character string to be processed is processed according to the unique identification number corresponding to each character of the character string to be processed and the target dictionary tree to obtain a target processing result, including:

[0079] According to the unique identification numbers corresponding to the respective characters of the string to be deleted, the first target node corresponding to the string to be deleted is searched in the target dictionary tree, and the string deletion result is obtained by deleting or modifying the word mark of the first target node.

[0080] The string to be processed in the embodiment of the present application is specifically a string to be deleted. That is, currently, according to the obtained string to be deleted, this string to be deleted needs to be located in the target trie and deleted. Since the target trie stores a complete string by adding a word end marker to the node corresponding to the last character of the string or adding a leaf node indicating the end of the word, in the embodiment of the present application, after finding the path node consistent with the string to be deleted in the target trie through the above string search method, the last node is used as the first target node corresponding to the string, and processing this first target node can achieve the deletion of the string to be deleted. In one embodiment, if the first target node is a node carrying a word marker, the word marker of the first target node is modified from "word end marker" to "non-end marker"; for example, assume the target trie is as shown in Figure 2 shown, and the string to be deleted is "烦恼" (annoyance). After finding the node storing the unique identification number corresponding to the character "恼" in the second-level child nodes of the target trie, modifying the word marker of this node to "non-end marker" can achieve the deletion of this string, and subsequently, the string "烦恼" cannot be found in the target trie. In another embodiment, if the first target node is a leaf node indicating the end of the word, the first target node is directly deleted; for example, assume the target trie is as shown in Figure 3 shown, and the string to be deleted is "烦恼" (annoyance). After finding the leaf node storing the information "[Chinese end]" connected to the second-level child node "恼" in the target trie, deleting this leaf node can achieve the deletion of this string, and subsequently, the string "烦恼" cannot be found in the target trie.

[0081] In the embodiment of the present application, since the target trie only needs to store the information of the unique identification number corresponding to the character, that is, it can indirectly store a character through a node of the target trie. After obtaining the string to be deleted, each character in the string to be deleted is first converted into the corresponding unique identification number, so that the string to be deleted can be efficiently and accurately located in the target trie, and by modifying or deleting the word marker of the first target node, the deletion of the string to be deleted can be completed, improving the character deletion efficiency.

[0082] Optionally, the string to be processed includes a string to be modified. Processing the string to be processed according to the unique identification numbers corresponding to the respective characters of the string to be processed and the target trie to obtain a target processing result includes:

[0083] According to the unique identification numbers corresponding to the respective characters of the string to be modified, a second target node corresponding to the string to be modified is searched in the target dictionary tree, and the string modification result is obtained by modifying the information of the second target node.

[0084] The character string to be processed in the embodiment of the present application is specifically a character string to be modified, that is, it is currently necessary to modify the information stored in the target dictionary tree according to the obtained character string to be modified. For the character string to be modified, after determining the unique identification number corresponding to each character of the character string to be modified in turn, locate the second target node storing the unique identification number corresponding to the character of the character string to be modified from the target dictionary tree according to each unique identification number through the step-by-step node indexing method during the above-mentioned string search. Afterwards, determine the unique identification number corresponding to each character of the target string (that is, the string to be modified to be modified), and update the information stored in the second target node to the information of the unique identification number corresponding to the character in the target string, thereby realizing the modification of the character string to be modified and obtaining the string modification result.

[0085] In an embodiment of the present application, since the target dictionary tree only needs to store the information of the unique identification number corresponding to the character, it can indirectly complete the storage of a character through a node of the target dictionary tree. After obtaining the character string to be modified, each character in the character string to be modified is first converted into the corresponding unique identification number, so that the character string to be modified can be efficiently and accurately located in the target dictionary tree, and by modifying the second target node, the character information modification can be efficiently and accurately achieved.

[0086] Optionally, before obtaining the character string to be processed, the method further includes:

[0087] Determine the preset data type based on the number of characters contained in the target language;

[0088] The target mapping relationship is obtained by respectively setting a data type of a unique identification number of the preset data type for each character of the target language.

[0089] In the embodiment of the present application, before obtaining the character string to be processed, a target mapping relationship between each character in the target language and the unique identification number is first set and stored.

[0090] In order to ensure that each character in the target language can be uniquely corresponded to a corresponding unique identification number, the amount of information that can be represented by the preset data type corresponding to the unique identification number needs to be greater than or equal to the number of characters contained in the target language. Therefore, in an embodiment of the present application, the preset data type currently used to set the unique identification number of the character can be determined according to the number of characters contained in the target language. Exemplarily, assuming that the number of characters contained in the target language is less than 65536, and the data type of a short integer can represent a total of 65536 values ​​from -32768 to 32767, therefore the preset data type of the unique identification number of the character of the target language can be a short integer.

[0091] After determining the preset data type, each character of the target language is sequentially set with a unique identification number whose data type is the preset data type, and each character and its corresponding unique identification number are bound and stored in a mapping table to obtain a target mapping relationship.

[0092] In the embodiment of the present application, before obtaining the character string to be processed, the preset data type can be accurately set according to the number of characters in the target language, so that the target mapping relationship can be accurately set, which facilitates the subsequent accurate conversion of the characters of the character string to be processed into the corresponding unique identification number.

[0093] Optionally, the target language is any one of Chinese, Korean and Japanese, and the preset data type is a short integer.

[0094] In the embodiment of the present application, the target language can be any one of Chinese, Korean, and Japanese. Since the characters corresponding to these three languages ​​occupy 3 bytes in the UTF-8 encoding format, the string processing method of the embodiment of the present application can be used to convert each character of the string to be processed into a corresponding unique identification number, and then the string to be processed is processed based on each unique identification number. Since the number of commonly used characters contained in these three languages ​​is less than 65536, a short integer can be used as the preset data type to complete the setting of the unique identification number of each character and obtain the target mapping relationship.

[0095] Exemplarily, taking the Chinese term library as an example, all single characters in the term library are counted, and there are 6,754 in total. Sequentially assign unique identification numbers 1 to 6,754 of data type short integer to these 6,754 characters, and bind and store each character with its corresponding unique identification number to obtain the target mapping relationship. For the Chinese character '人' (person), originally under the UTF-8 encoding format, three nodes storing the three byte information of 147, 156, and 16 respectively were required to complete the storage of this Chinese character; while through the method of the embodiments of the present application, first determine the unique identification number '375' of the Chinese character '人' through the target mapping relationship, and then, in the target trie, only one node is needed to store the unique identification number 375 of data type short integer. Compared with the original method, the number of nodes is reduced by three times, so the occupied storage space can be greatly reduced.

[0096] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0097] Embodiment 2:

[0098] Figure 4 The structural schematic diagram of a string processing device provided by the embodiments of the present application is shown. For the convenience of description, only the parts related to the embodiments of the present application are shown:

[0099] The string processing device includes: an acquisition unit 41, a mapping unit 42, and a processing unit 43. Among them:

[0100] The acquisition unit 41 is used to acquire the string to be processed; wherein, the string to be processed is a string in the target language, and the number of bytes occupied by each character in the string in the target language is greater than 1.

[0101] The mapping unit 42 is used to determine the unique identification number corresponding to each character in the string to be processed according to the preset target mapping relationship; wherein, the target mapping relationship includes the mapping relationship between each character in the target language and the unique identification number, and the data type of the unique identification number is the preset data type.

[0102] The processing unit 43 is used to process the string to be processed according to the unique identification numbers respectively corresponding to the characters of the string to be processed and the target trie to obtain the target processing result; wherein, the data type of each node in the target trie is the preset data type.

[0103] Optionally, the string to be processed includes a string to be stored, and the processing unit 43 is specifically used to store the unique identification number in the node of the target dictionary tree according to the unique identification number corresponding to each character of the string to be stored, so as to obtain a string storage result.

[0104] Optionally, the string to be processed includes a string to be searched, and the processing unit 43 is specifically used to perform node indexing in the target dictionary tree step by step according to the unique identification numbers corresponding to the respective characters of the string to be searched, to obtain a string search result.

[0105] Optionally, the string to be processed includes a string to be deleted, and the processing unit 43 is specifically used to search for a first target node corresponding to the string to be deleted in the target dictionary tree according to the unique identification numbers corresponding to each character of the string to be deleted, and obtain a string deletion result by performing word tag modification or deletion processing on the first target node.

[0106] Optionally, the string to be processed includes a string to be modified, and the processing unit 43 is specifically used to search for a second target node corresponding to the string to be modified in the target dictionary tree according to the unique identification numbers corresponding to each character of the string to be modified, and obtain a string modification result by modifying the information of the second target node.

[0107] Optionally, the string processing device further includes:

[0108] The target mapping relationship setting unit is used to determine a preset data type according to the number of characters contained in the target language; set the data type of each character of the target language to a unique identification number of the preset data type to obtain the target mapping relationship.

[0109] Optionally, the target language is any one of Chinese, Korean and Japanese, and the preset data type is a short integer.

[0110] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0111] Embodiment three:

[0112] Figure 5 is a schematic diagram of an electronic device provided by an embodiment of the present application. Figure 5As shown, the electronic device 5 of this embodiment includes: a processor 50, a memory 51, and a computer program 52 stored in the memory 51 and executable on the processor 50, such as a string processing program. When the processor 50 executes the computer program 52, the steps in the above-mentioned various string processing method embodiments are implemented, such as Figure 1 Alternatively, when the processor 50 executes the computer program 52, the functions of each module / unit in the above-mentioned device embodiments are realized, for example Figure 4 The functions of acquisition units 41 to 43 are shown.

[0113] Exemplarily, the computer program 52 may be divided into one or more modules / units, which are stored in the memory 51 and executed by the processor 50 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 52 in the electronic device 5.

[0114] The electronic device 5 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic device may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will appreciate that Figure 5 It is only an example of the electronic device 5 and does not constitute a limitation of the electronic device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0115] The processor 50 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0116] The memory 51 may be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. The memory 51 may also be an external storage device of the electronic device 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 5. Further, the memory 51 may also include both an internal storage unit and an external storage device of the electronic device 5. The memory 51 is used to store the computer program and other programs and data required by the electronic device. The memory 51 may also be used to temporarily store data that has been output or is to be output.

[0117] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0118] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0119] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0120] In the embodiments provided in the present application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0121] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0122] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0123] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0124] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for processing a character string, characterized in that: include: Obtaining a character string to be processed; wherein the character string to be processed is a character string in a target language, and the number of bytes occupied by each character in the character string in the target language is greater than 1; Determine the unique identification number corresponding to each character in the character string to be processed according to a preset target mapping relationship; wherein the target mapping relationship includes a mapping relationship between each character of the target language and the unique identification number, and the data type of the unique identification number is a preset data type; According to the unique identification numbers and the target dictionary tree corresponding to the characters of the string to be processed, the string to be processed is processed to obtain a target processing result; wherein the data type of each node in the target dictionary tree is the preset data type; The target language is any one of Chinese, Korean and Japanese, and the preset data type is a short integer; The root node of the dictionary tree does not contain any characters, and each child node except the root node contains one character.

2. The character string processing method according to claim 1, characterized in that: The character string to be processed includes a character string to be stored, and the character string to be processed is processed according to the unique identification number corresponding to each character of the character string to be processed and the target dictionary tree to obtain a target processing result, including: According to the unique identification numbers corresponding to the respective characters of the string to be stored, the unique identification numbers are stored in the nodes of the target dictionary tree to obtain a string storage result.

3. The character string processing method according to claim 1, wherein: The character string to be processed includes a character string to be searched, and the character string to be processed is processed according to the unique identification number corresponding to each character of the character string to be processed and the target dictionary tree to obtain a target processing result, including: According to the unique identification numbers corresponding to the respective characters of the string to be searched, node indexing is performed level by level in the target dictionary tree to obtain a string search result.

4. The character string processing method according to claim 1, wherein: The character string to be processed includes a character string to be deleted, and the character string to be processed is processed according to the unique identification number corresponding to each character of the character string to be processed and the target dictionary tree to obtain a target processing result, including: According to the unique identification numbers corresponding to the respective characters of the string to be deleted, the first target node corresponding to the string to be deleted is searched in the target dictionary tree, and the string deletion result is obtained by modifying or deleting the word mark of the first target node.

5. The character string processing method according to claim 1, wherein: The character string to be processed includes a character string to be modified, and the character string to be processed is processed according to the unique identification number corresponding to each character of the character string to be processed and the target dictionary tree to obtain a target processing result, including: According to the unique identification numbers corresponding to the respective characters of the string to be modified, a second target node corresponding to the string to be modified is searched in the target dictionary tree, and the string modification result is obtained by modifying the information of the second target node.

6. The character string processing method according to any one of claims 1, characterized in that: Before obtaining the string to be processed, the method further includes: Determine the preset data type based on the number of characters contained in the target language; The target mapping relationship is obtained by respectively setting a data type of a unique identification number of the preset data type for each character of the target language.

7. A string processing device, characterized in that: include: An acquisition unit, configured to acquire a character string to be processed; wherein the character string to be processed is a character string in a target language, and the number of bytes occupied by each character in the character string in the target language is greater than 1; A mapping unit, configured to determine a unique identification number corresponding to each character in the character string to be processed according to a preset target mapping relationship; wherein the target mapping relationship includes a mapping relationship between each character of the target language and the unique identification number, and the data type of the unique identification number is a preset data type; A processing unit, configured to process the string to be processed according to the unique identification numbers corresponding to the respective characters of the string to be processed and a target dictionary tree, to obtain a target processing result; wherein the data type of each node in the target dictionary tree is the preset data type; The target language is any one of Chinese, Korean and Japanese, and the preset data type is a short integer; The root node of the dictionary tree does not contain any characters, and each child node except the root node contains one character.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the electronic device implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the electronic device implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for storing message theme

    CN113239307A

  • Log storage method and device, intelligent loudspeaker box and cloud server

    CN113656277A