Method and device for processing rare characters

By encoding and grouping rare characters based on standard character encoding tables and converting them into specified character encoding using mapping relationships, the difficult problem of rare characters processing in financial business systems is solved, and normal business operations and strong compatibility of non-BMP-encoded Chinese characters are achieved.

CN119358510BActive Publication Date: 2025-05-13CREDIT CENT OF THE PEOPLES BANK OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411909742.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-13
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

In the financial business system, customer information containing rare words is difficult to be recognized and processed normally during network transmission and data processing, resulting in the inability to carry out the business normally.

Method used

By character encoding of rare characters based on standard character encoding tables, grouping them according to the encoding values ​​of each character encoding, and converting the group encoding into a specified character encoding using the mapping relationship between the specified character encoding and the group encoding.

Benefits of technology

It realizes the conversion of uncommon characters from multiple character encodings to correct specified character encodings, ensuring that non-BMP-encoded Chinese characters can complete business operations normally, and has strong compatibility, without the need to develop a dedicated system or platform separately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119358510B_ABST
    Figure CN119358510B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of information processing technology, and provides a method and device for processing rare characters, the method comprising: based on a standard character encoding table, character encoding is performed on the rare characters to be processed to obtain multiple character codes; based on the encoding value of each character code, each character code is grouped to obtain a group code; based on the mapping relationship between the designated character code and the group code, the group code is converted into the designated character code. The method realizes the conversion of rare characters from multiple character codes into the correct designated character code, ensuring that non-BMP encoded Chinese characters can complete business operations normally. In addition, the present invention does not need to develop a dedicated system or platform for the designated character encoding table alone, that is, the present invention can realize the conversion of multiple character codes into the correct designated character code on the system or platform corresponding to the standard character encoding table, and has strong compatibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technology, and in particular to a method and device for processing rare characters. Background Art

[0002] With the rapid development of information technology and the deepening of economic and social digitalization, the convenience, efficiency and security of financial services have become the core pursuit of the industry. However, as the real-name system requirements become more stringent, the accuracy and completeness of customer information are crucial for financial institutions. In this context, a series of challenges faced by customers with uncommon characters in their names when handling financial services have gradually emerged.

[0003] At present, when processing customer information containing rare characters, we mainly rely on standardized character encoding systems. However, due to the particularity of rare characters in character encoding, especially those Chinese characters that do not belong to the Basic Multilingual Plane (BMP), they are often difficult to be recognized and processed normally during network transmission and data processing. As a result, these non-BMP encoded Chinese characters are often misjudged as illegal characters when passing through third-party payment, social security, securities, insurance and other financial business systems, resulting in the inability to conduct business normally. Summary of the invention

[0004] The present invention provides a method and device for processing rare characters, which are used to solve the defects in the prior art.

[0005] The present invention provides a method for processing rare characters, comprising the following steps:

[0006] Based on the standard character encoding table, character encoding is performed on the rare characters to be processed to obtain multiple character encodings;

[0007] Based on the code value of each character code, each character code is grouped to obtain a group code;

[0008] Based on the mapping relationship between the specified character code and the group code, the group code is converted into the specified character code.

[0009] According to a rare word processing method provided by the present invention, the character codes are grouped based on the code values ​​of the character codes to obtain group codes, including:

[0010] If the code value of the current character code is less than the threshold, determining whether the code value of the next character code is greater than the threshold;

[0011] If so, the current character code and the next character code are divided into the same group to obtain the group code.

[0012] According to a method for processing rare characters provided by the present invention, the method further comprises:

[0013] If the code value of the next character code is less than or equal to the threshold, the next character code is used as the current character code.

[0014] According to a rare word processing method provided by the present invention, the group code is converted into a specified character code, and then the method further includes:

[0015] Determine whether the uncommon character to be processed is a multi-code character;

[0016] If yes, then based on the multi-codeword mapping table, obtain all standard character codes corresponding to the specified character code;

[0017] All standard character codes of the specified character code are combined to obtain multiple code combinations.

[0018] According to a rare character processing method provided by the present invention, all standard character codes of the specified character code are combined to obtain multiple code combinations, including:

[0019] Cartesian product operation is performed on all standard character codes of the specified character code to obtain the multiple code combinations.

[0020] According to a rare character processing method provided by the present invention, the method further comprises: obtaining a plurality of coding combinations, and then:

[0021] Based on the query library coding area, a target coding combination is selected from various coding combinations;

[0022] Based on the target code combination, search in the query library.

[0023] The present invention also provides a device for processing rare characters, comprising the following modules:

[0024] The encoding unit is used to perform character encoding on the rare characters to be processed based on a standard character encoding table to obtain multiple character encodings;

[0025] A grouping unit, used for grouping each character code based on the code value of each character code to obtain a group code;

[0026] The conversion unit is used to convert the group code into the specified character code based on the mapping relationship between the specified character code and the group code.

[0027] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned rare character processing methods is implemented.

[0028] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for processing rare characters as described above is implemented.

[0029] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned rare character processing methods.

[0030] The method and device for processing rare characters provided by the present invention realizes the conversion of rare characters from multiple character codes into the correct designated character code (i.e., the correct non-BMP code) by encoding the rare characters based on the standard character code table, grouping according to the code value of each character code, and converting using a mapping relationship, thereby ensuring that non-BMP-coded Chinese characters can normally complete business operations. In addition, the present invention does not need to develop a dedicated system or platform for the designated character code table alone, that is, the present invention can realize the conversion of multiple character codes into the correct designated character code on the system or platform corresponding to the standard character code table, and has strong compatibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0032] Figure 1 This is one of the flow charts of the rare character processing method provided by the present invention.

[0033] Figure 2 This is the second flow chart of the rare character processing method provided by the present invention.

[0034] Figure 3 This is the third flow chart of the rare character processing method provided by the present invention.

[0035] Figure 4 It is a structural schematic diagram of the rare character processing device provided by the present invention.

[0036] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0038] At present, when processing customer information containing rare characters, we mainly rely on standardized character encoding systems. However, due to the particularity of rare characters in character encoding, especially those Chinese characters that do not belong to the Basic Multilingual Plane (BMP), they are often difficult to be recognized and processed normally during network transmission and data processing. As a result, these non-BMP encoded Chinese characters are often misjudged as illegal characters when passing through third-party payment, social security, securities, insurance and other financial business systems, resulting in the inability to conduct business normally.

[0039] In addition, in inter-bank transfer business, due to the inconsistency of character encoding standards among different financial institutions and the differences in the representation of uncommon characters in different systems, uncommon characters in account names have the phenomenon of "one word with multiple codes", making it impossible for the inter-bank transfer system to accurately match the account name, and thus unable to complete the automatic transfer operation, causing great inconvenience to customers.

[0040] Among them, the Unicode character set is an encoding scheme developed by an international organization to accommodate all characters in the world, including character sets, encoding schemes, etc. It sets a unified and unique binary encoding for each character in each language to meet cross-language and cross-platform requirements.

[0041] Through gradual iterative upgrades, the Unicode character set is divided into 17 planes, each with 65,536 code points, that is, each code point can represent a character. Among them, planes 1 to 16 are collectively referred to as non-BMP planes, and the characters represented in them are collectively referred to as non-BMP characters.

[0042] Among them, Table 1 is a list of the range and storage content of each plane code point of the Unicode character set. From Table 1, the encoding area and range of rare characters can be sorted out, and Table 2 is a list of the encoding area and range of rare characters.

[0043] Table 1

[0044]

[0045] Table 2

[0046]

[0047] Through analysis, it is found that the current mainstream application network communications are all encoded and transmitted based on the JAXB (Java Architecture for XML Binding) standard. In this process, some communication frameworks or middleware cannot correctly handle non-BMP characters due to technical limitations or version issues. That is, when these frameworks or middleware encounter non-BMP characters, they may mistakenly split them into two surrogate pairs. For example, the character "𠅤" (correctly encoded as U+20164) may be incorrectly encoded as U+D840 and U+DD64. U+D840 and U+DD64 are not valid Unicode character encodings and cannot pass the data verification of downstream application services, resulting in communication anomalies. In other words, non-BMP-encoded Chinese characters cannot complete business operations normally.

[0048] In this regard, an embodiment of the present invention provides a method for processing rare characters. Figure 1 It is one of the flow charts of the method for processing rare characters provided by the present invention, such as Figure 1 As shown, the method includes step 110 , step 120 and step 130 .

[0049] Step 110: Based on a standard character encoding table, character encoding is performed on the rare characters to be processed to obtain a plurality of character codes.

[0050] Here, the standard character encoding table can be understood as a comparison table containing characters and their corresponding encoding values. The standard character encoding table is a standard for character encoding, which is used to convert characters (such as letters, numbers, punctuation marks, and Chinese characters) into digital or binary codes for storage, processing, and transmission in computers. Among them, the standard character encoding table includes ASCII, UTF-8, UTF-16, GBK, Big5, etc.

[0051] In addition, it should be noted that the standard character encoding table is a standardized encoding table, which usually covers characters of multiple languages ​​and can be applied to most scenarios. The standard character encoding table can be compatible in different systems or platforms. Therefore, when multiple systems or platforms follow the standard character encoding table for encoding, text data can be shared and transmitted seamlessly.

[0052] That is to say, the embodiment of the present invention performs character encoding on rare characters to be processed based on the standard character encoding table, so that the obtained multiple character encodings can be seamlessly shared and transmitted between different systems or platforms, without the need to install additional corresponding software on the corresponding system or platform to obtain and transmit character encodings, which not only improves the efficiency and reliability of data transmission, but also reduces the cost of data transmission.

[0053] Among them, character encoding of the rare characters to be processed is the process of converting the rare characters to be processed into digital or binary codes. For example, if the standard character encoding table is the UTF-16 encoding table, it usually uses 16-bit or 32-bit (through surrogate pairs) character encoding to represent the rare characters to be processed.

[0054] Step 120: Group the character codes based on the code values ​​of the character codes to obtain group codes.

[0055] Specifically, the code value of each character code is used to characterize the storage order and position of each character code. For example, in the address space of a memory or storage device, relative to a certain reference point or starting point, a character code with a smaller code value is called a high-address bit code. In contrast to a high-address bit code, a character code with a larger code value is called a low-address bit code.

[0056] When the standard character encoding table is the UTF-16 encoding table, the high address bits of the character encoding represented by it are between U+D800 and U+DBFF, and the low address bits are between U+DC00 and U+DFFF.

[0057] After encoding the uncommon characters to be processed based on the standard character encoding table, the uncommon characters to be processed will be split into two surrogate pairs (i.e., character encodings with two different encoding values). For example, the character "𠅤" (the correct character encoding is U+20164) may be encoded into two surrogate pairs U+D840 and U+DD64. The encodings of U+D840 and U+DD64 are not valid Unicode character encodings and cannot pass the data verification of downstream application services, resulting in communication anomalies, etc., which makes it impossible to complete business operations normally for non-BMP encoded Chinese characters.

[0058] In this regard, the embodiment of the present invention groups each character code based on the code value of each character code to obtain a group code. For example, if the rare character to be processed includes two character codes, one character code is a high address bit, and the other character code is a low address bit, then the character code of the high address bit and the character code of the low address bit can be combined into a group to obtain a group code, and the group code corresponds to a unique specified character code.

[0059] Step 130: Based on the mapping relationship between the designated character code and the group code, convert the group code into the designated character code.

[0060] Specifically, the specified character encoding table can be understood as a character encoding table formulated according to specific needs or application scenarios. The specified character encoding table may only contain a specific language or a specific character set. The specified character encoding table is usually only applied to a specific system or platform, that is, it cannot be compatible with different systems or platforms. In other words, if the traditional method is to perform character encoding based on the specified character encoding table, it must be executed on a specific system or platform, and the portability is poor. Among them, the specified character encoding table can be a Unicode character table.

[0061] In this regard, the embodiment of the present invention converts the group code into the specified character code based on the mapping relationship between the specified character code and the group code, thereby realizing the acquisition of the specified character code on a system or platform that does not need to be compatible with the specified character code table. The mapping relationship between the specified character code table and the standard character code table is used to indicate that the characters in the standard character code table are converted into the specified character code.

[0062] For example, for the character "𠅤", the character encodings obtained based on the UTF-16 encoding table include U+D840 and U+DD64. The mapping relationship between the specified character encoding and the group encoding indicates that U+D840 and U+DD64 correspond to the specified character encoding U+20164, and then it can be obtained that the specified character encoding corresponding to "𠅤" is U+20164.

[0063] The embodiment of the present invention realizes the conversion of rare characters from multiple character codes into the correct designated character code (i.e., the correct non-BMP code) by encoding rare characters based on the standard character code table, grouping according to the code value of each character code, and converting using a mapping relationship, thereby ensuring that non-BMP-coded Chinese characters can normally complete business operations. In addition, the embodiment of the present invention does not need to develop a dedicated system or platform for the designated character code table alone, that is, the embodiment of the present invention can realize the conversion of multiple character codes into the correct designated character code on the system or platform corresponding to the standard character code table, and has strong compatibility.

[0064] In addition, it can be understood that the multiple character encodings obtained by the embodiment of the present invention based on the standard character encoding table can be seamlessly transmitted and stored between multiple systems or platforms. That is to say, the multiple character encodings obtained in the embodiment of the present invention can be transmitted to different business terminals, and different business terminals determine the designated character encoding table according to business needs, and establish a mapping relationship between their respective designated character encoding tables and the standard character encoding table. Therefore, after seamlessly receiving multiple character encodings, each business terminal can convert the character encoding into the designated character encoding required by each according to the mapping relationship established by each business terminal, without repeatedly generating the character encoding corresponding to the standard character encoding table at each business terminal, thereby further saving resources.

[0065] Based on the above embodiment, each character code is grouped based on the code value of each character code to obtain a group code, including:

[0066] If the code value of the current character code is less than the threshold, determine whether the code value of the next character code is greater than the threshold;

[0067] If so, the current character code and the next character code are divided into the same group to obtain a group code.

[0068] Specifically, the encoding value of the current character encoding is used to represent the corresponding value of the current character encoding in the standard character encoding table. In order to represent more characters, the "surrogate pair" mechanism is usually adopted to represent characters, that is, a character with a larger encoding value is represented by combining two character encodings with smaller encoding values. For example, for the word "𠅤", the character encoding obtained based on the UTF-16 encoding table includes two character encodings with smaller encoding values, U+D840 and U+DD64, and U+D840 and U+DD64 correspond to the specified character encoding U+20164 with a larger encoding value.

[0069] On this basis, the embodiment of the present invention first determines whether the encoding value of the current character code is less than the threshold value. If so, it indicates that the current character code corresponds to the high address bit code. At this time, the current character code is recorded, and it continues to be determined whether the encoding value of the next character code is greater than the threshold value. If so, it indicates that the next character code corresponds to the low address bit code. At this time, the next character code is recorded.

[0070] In this case, the current character code and the next character code are divided into the same group to obtain a group code, that is, the group code includes a high address bit code + a low address bit code, and then the group code is converted into the specified character code according to the mapping relationship between the specified character code table and the standard character code table. Among them, the threshold value can correspond to the highest code value DBFF in the high address bit code range (D800~DBFF) (converted to decimal is 56319).

[0071] It can be seen that the embodiment of the present invention can represent the specified character code with a larger code value through multiple character codes with smaller code values, and thus can represent different characters in a diversified manner. In addition, the embodiment of the present invention also takes into account the problem that the specified character code table is not compatible with different systems or platforms. Therefore, the mapping relationship between the specified character code table and the standard character code table is adopted, and after the multiple character codes are grouped, the grouped codes are converted into the specified character code, thereby avoiding the problem that the system or platform is not compatible with the specified character code table for encoding.

[0072] That is to say, the embodiment of the present invention not only ensures diversified representation of more different characters, but also takes into account the compatibility issues of the system or platform, and realizes the conversion of rare characters from multiple character encodings into the correct specified character encoding (i.e., the correct non-BMP encoding), ensuring that non-BMP encoded Chinese characters can complete business operations normally.

[0073] Based on any of the above embodiments, the method further includes:

[0074] If the code value of the next character code is less than or equal to the threshold, the next character code is used as the current character code.

[0075] Specifically, if it is determined that the encoding value of the current character code is less than the threshold, it indicates that the current character code corresponds to a high-address bit code. At this time, the current character code is recorded, and it is continued to be determined whether the encoding value of the next character code is greater than the threshold. If not, it indicates that the next character code does not correspond to a low-address bit code (that is, the character code may be a high-address bit code). At this time, the next character code is used as the current character code, and the above-mentioned current character code determination steps are repeated.

[0076] In addition, it should be noted that if the encoding value of the current character encoding is greater than or equal to the threshold, it indicates that the rare character to be processed may not be represented by the above-mentioned surrogate pair mechanism. In this case, the corresponding character encoding can be directly output.

[0077] Figure 2 FIG. 2 is a flow chart of the method for processing rare characters provided by the present invention. Figure 2 As shown, the method can be applied to the above-mentioned framework or middleware for encoding and transmission based on the JAXB standard, and solves the problem that non-BMP characters cannot be normally transmitted between upstream and downstream systems. Specifically, it includes: receiving Chinese characters input by the user, and encoding the Chinese characters based on the UTF-16 table to obtain multiple character codes. First, it is determined whether the encoding value of the current character code is less than the threshold value. If so, it indicates that the current character code corresponds to the high address bit code. At this time, the current character code is recorded, and the encoding value of the next character code is further determined to be greater than the threshold value. If so, it indicates that the next character code corresponds to the low address bit code. At this time, the next character code is recorded. In this case, the current character code and the next character code are divided into the same group to obtain a group code, that is, the group code includes the high address bit code + the low address bit code, and then according to the mapping relationship between the specified character code table and the standard character code table, the group code is converted into the specified character code. If the encoding value of the current character code is greater than or equal to the threshold, it indicates that the Chinese character may not be represented by the above-mentioned proxy pair mechanism. At this time, the corresponding character code can be directly output.

[0078] In addition, the GBK encoding specification was released in 1980, which includes 21,003 Chinese characters and 883 symbols. The Unicode encoding specification was released in 1990, which is later than the GBK encoding specification. The inclusion of new characters in Unicode requires review and upgrading, which often takes several years, but they can be temporarily encoded in the PUA area. Some input method and font manufacturers have encoded some Chinese characters in the PUA area for this purpose. Later, these Chinese characters were assigned formal codes in the Unicode character set. For example, the PUA code of the character "䶮" is U+E863, and the formal code is U+4DAE.

[0079] The above situation shows that there are multiple encodings for a Chinese character in the Unicode character set. Such Chinese characters are collectively called multi-code characters. Multi-code characters will cause the information system to store the character "䶮" encoded as U+4DAE, but the data cannot be queried using the character "䶮" U+E863.

[0080] In response to the above situation, an embodiment of the present invention adopts a multi-codeword mapping table to solve the problem. The table maintains the mapping relationship between the official code and the corresponding PUA. For example, when the access agency uses "䶮" (encoded U+4DAE) to query data, if no relevant data is matched, the multi-codeword mapping table can be used to query whether the character has multiple codes. Through this table, it can be known that the character "䶮" has multiple codes, and then another code U+E863 corresponding to "䶮" can be obtained, and the data can be queried again based on the code U+E863. In this way, the problem that multiple codewords cannot accurately query data is solved.

[0081] Specifically, the embodiment of the present invention converts the group code into a specified character code, and then further includes:

[0082] Determine whether the uncommon character to be processed is a multi-code character;

[0083] If yes, then based on the multi-codeword mapping table, obtain all standard character codes corresponding to the specified character code;

[0084] All standard character encodings of the specified character encoding are combined to obtain multiple encoding combinations.

[0085] Specifically, if the uncommon character to be processed is a multi-code word, it indicates that the corresponding uncommon character corresponds to multiple different standard character codes. Considering that in some systems, it may be possible to query based on only one standard character code in the uncommon character, and it is impossible to query based on the other standard character codes, when the embodiment of the present invention determines that the uncommon character to be processed is a multi-code word, based on the multi-code word mapping table, all standard character codes corresponding to the specified character code are obtained, and all standard character codes of the specified character code are combined to obtain multiple code combinations, and the multiple code combinations use different standard character codes to represent the specified character code.

[0086] For example, Chinese character 1 corresponds to a standard character code a, and Chinese character 2 corresponds to standard character code b and standard character code c, then the code combinations include: code a+code b, code a+code c.

[0087] In addition, the embodiment of the present invention is also based on a multi-codeword mapping table to obtain all standard character encodings corresponding to the specified character encoding. Since the multi-codeword mapping table is used to characterize the mapping relationship between the specified character encoding and the standard character encoding, there is no need to perform character encoding based on the specified character encoding table. It is possible to obtain all standard character encodings corresponding to the specified character encoding on a system or platform that does not need to be compatible with the specified character encoding table. In other words, the above embodiment can also be applied to frameworks or middleware that encode and transmit based on the JAXB standard to solve the problem that various systems cannot handle multiple codewords normally.

[0088] Based on any of the above embodiments, all standard character encodings of the specified character encoding are combined to obtain multiple encoding combinations, including:

[0089] Perform a Cartesian product operation on all standard character encodings of a specified character encoding to obtain multiple encoding combinations.

[0090] Specifically, a Cartesian operation is performed on all standard character codes of a specified character code, so that all possible code combinations can be accurately obtained. For example, Chinese character 1 corresponds to m standard character codes, and Chinese character 2 corresponds to n standard character codes. After performing a Cartesian operation on Chinese character 1 and Chinese character 2, a total of m×n code combinations can be obtained.

[0091] Based on any of the above embodiments, multiple coding combinations are obtained, and then the following is further included:

[0092] Based on the query library coding area, a target coding combination is selected from various coding combinations;

[0093] Search the query library based on the target code combination.

[0094] Specifically, the query library encoding area is used to represent the encoding range used by the characters in the query library. If the encoding combination is within the encoding range, it means that the corresponding result can be retrieved in the query library based on the encoding combination. Similarly, if the encoding combination is outside the encoding range, it means that the corresponding result cannot be retrieved in the query library based on the encoding combination.

[0095] In this regard, the embodiment of the present invention selects a target coding combination that matches the coding range from each coding combination based on the query library coding area, that is, the target coding combination is within the coding range, so that the corresponding result can be accurately retrieved from the query library based on the target coding combination.

[0096] It can be seen that the embodiment of the present invention selects a target coding combination from various coding combinations based on the query library coding area, so that the query library can be accurately searched based on the target coding combination, solving the problem that multiple codewords cannot be accurately queried.

[0097] Figure 3 FIG. 3 is a flow chart of the method for processing rare characters provided by the present invention. Figure 3 As shown, this method can be applied to the above-mentioned framework or middleware based on the JAXB standard for encoding and transmission to solve the problem that multiple codewords cannot be accurately queried, specifically including: receiving the query characters input by the user, and using String.codePointAt(offset) to convert the query characters into Unicode encoding (i.e., the specified character encoding). Read the multi-codeword mapping table in the memory, match and find the official encoding or PUA encoding (i.e., standard character encoding) corresponding to each Unicode encoding in turn, and perform Cartesian product operations on the results of each encoding match to obtain multiple encoding combinations. Then, use StringBuilder.appendCodePoint(codePoint) to convert each encoding combination into the corresponding string in turn, and output the string list to the user end.

[0098] The following is a description of the rare character processing device provided by the present invention. The rare character processing device described below and the rare character processing method described above can be referenced to each other.

[0099] Based on any of the above embodiments, Figure 4 Schematic diagram of the structure of the device for processing rare characters provided by the present invention. Figure 4 As shown, the device comprises:

[0100] An encoding unit 410 is used to perform character encoding on the rare characters to be processed based on a standard character encoding table to obtain a plurality of character codes;

[0101] A grouping unit 420, configured to group the character codes based on the code values ​​of the character codes to obtain group codes;

[0102] The conversion unit 430 is used to convert the group code into the specified character code based on the mapping relationship between the specified character code and the group code.

[0103] Based on any of the above embodiments, the character codes are grouped based on the code values ​​of the character codes to obtain group codes, including:

[0104] If the code value of the current character code is less than the threshold, determine whether the code value of the next character code is greater than the threshold;

[0105] If so, the current character code and the next character code are divided into the same group to obtain a group code.

[0106] Based on any of the above embodiments, it also includes:

[0107] If the code value of the next character code is less than or equal to the threshold, the next character code is used as the current character code.

[0108] Based on any of the above embodiments, the group code is converted into a specified character code, and then the following steps are further included:

[0109] Determine whether the uncommon character to be processed is a multi-code character;

[0110] If yes, then based on the multi-codeword mapping table, obtain all standard character codes corresponding to the specified character code;

[0111] Combine all standard character encodings of the specified character encoding to obtain multiple encoding combinations.

[0112] Based on any of the above embodiments, all standard character encodings of the specified character encoding are combined to obtain multiple encoding combinations, including:

[0113] Perform a Cartesian product operation on all standard character encodings of a specified character encoding to obtain multiple encoding combinations.

[0114] Based on any of the above embodiments, multiple coding combinations are obtained, and then the following is further included:

[0115] Based on the query library coding area, a target coding combination is selected from various coding combinations;

[0116] Search the query library based on the target code combination.

[0117] Figure 5 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the rare character processing method, which includes: based on the standard character encoding table, character encoding the rare characters to be processed to obtain multiple character codes; based on the encoding value of each character code, grouping each character code to obtain a group code; based on the mapping relationship between the specified character code and the group code, converting the group code into the specified character code.

[0118] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0119] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the rare character processing method provided by the above-mentioned methods, which includes: based on a standard character encoding table, character encoding the rare characters to be processed to obtain multiple character codes; based on the encoding value of each character code, grouping each character code to obtain a group code; based on the mapping relationship between the specified character code and the group code, converting the group code into a specified character code.

[0120] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the rare character processing method provided by the above-mentioned methods. The method includes: based on a standard character encoding table, character encoding the rare characters to be processed to obtain multiple character codes; based on the encoding value of each character code, grouping each character code to obtain a group code; based on the mapping relationship between the specified character code and the group code, converting the group code into a specified character code.

[0121] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0122] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing rare characters, characterized in that: include: Based on the standard character encoding table, character encoding is performed on the rare characters to be processed to obtain multiple character encodings; The multiple character encodings can be seamlessly shared and transmitted between different systems or platforms; Based on the code value of each character code, each character code is grouped to obtain a group code; Based on the mapping relationship between the specified character code and the group code, convert the group code into the specified character code; The step of grouping the character codes based on the code values ​​of the character codes to obtain group codes includes: If the code value of the current character code is less than the threshold, determining whether the code value of the next character code is greater than the threshold; If yes, the current character code and the next character code are grouped into the same group to obtain the group code; The step of converting the group code into a specified character code further includes: Determine whether the uncommon character to be processed is a multi-code character; If yes, then based on the multi-codeword mapping table, obtain all standard character codes corresponding to the specified character code; All standard character codes of the specified character code are combined to obtain multiple code combinations.

2. The method for processing rare characters according to claim 1, characterized in that: The method further comprises: If the code value of the next character code is less than or equal to the threshold, the next character code is used as the current character code.

3. The method for processing rare characters according to claim 1, characterized in that: The step of combining all standard character codes for the specified character code to obtain multiple code combinations includes: Cartesian product operation is performed on all standard character codes of the specified character code to obtain the multiple code combinations.

4. The method for processing rare characters according to claim 1, characterized in that: The method further comprises: obtaining a plurality of coding combinations; and then: Based on the query library coding area, a target coding combination is selected from various coding combinations; Based on the target code combination, search in the query library.

5. A device for processing rare characters, characterized in that: include: The encoding unit is used to perform character encoding on the rare characters to be processed based on a standard character encoding table to obtain multiple character encodings; The multiple character encodings can be seamlessly shared and transmitted between different systems or platforms; A grouping unit, used for grouping each character code based on the code value of each character code to obtain a group code; A conversion unit, configured to convert the group code into the specified character code based on a mapping relationship between the specified character code and the group code; The step of grouping the character codes based on the code values ​​of the character codes to obtain group codes includes: If the code value of the current character code is less than the threshold, determining whether the code value of the next character code is greater than the threshold; If yes, the current character code and the next character code are grouped into the same group to obtain the group code; The step of converting the group code into a specified character code further includes: Determine whether the uncommon character to be processed is a multi-code character; If yes, then based on the multi-codeword mapping table, obtain all standard character codes corresponding to the specified character code; All standard character codes of the specified character code are combined to obtain multiple code combinations.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for processing rare characters according to any one of claims 1 to 4 is implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for processing rare characters as claimed in any one of claims 1 to 4 is implemented.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for processing rare characters as claimed in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Method of network inquiry four syllable character and its system

    CN1719440A