Uncommon character processing method, uncommon character processing device and computer readable storage medium
By storing the configuration information of uncommon word in the local cache and quickly matching it, the problem of low processing efficiency of uncommon word in the existing technology is solved, and efficient and accurate recognition and processing of uncommon word is achieved.
Patent Information
- Application Number
- CN202510156441.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
AI Technical Summary
The processing efficiency of rare words in the prior art is low, resulting in increased data analysis complexity and unnecessary parsing redundancy. The uncommon words need to be converted into standard codes and mapped back to Chinese characters after detection, which affects the processing efficiency.
By obtaining the configuration information of uncommon characters and storing it in the local cache, we judge whether the field in the data processing request is a target field that needs uncommon characters processing, find an entry that matches the Chinese character encoding in the target field, convert the Chinese character encoding into a preset format encoding, and judge the encoding consistency under multiple mapping items to determine the uncommon characters.
It improves the efficiency and accuracy of uncommon word processing, reduces unnecessary parsing and conversion processes, and improves the speed of data processing and the system's response performance.
Smart Images

Figure CN120086358A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer information processing. Specifically, it relates to a rare Chinese character processing method, a rare Chinese character processing device, a computer-readable storage medium, and an electronic device. Background Art
[0002] Currently, there are still many rare Chinese characters not included in the Chinese character encoding standard, resulting in that electronic devices such as computers and smart phones cannot display rare Chinese characters normally like common Chinese characters in some scenarios. Therefore, it is necessary to process rare Chinese characters, that is, to identify and convert rare Chinese characters.
[0003] The identification and conversion of rare Chinese characters refer to the process of processing and managing some infrequently used or relatively rare Chinese characters. In practical applications, a character library containing rare Chinese characters is usually established to ensure that the system can correctly identify and display these characters. The maintenance of the character library is all carried out by manual character splitting and input; or the Chinese characters are decoded in sequence, the character strings within the encoding range of rare Chinese characters are extracted, converted into formal encodings that are not user-defined, and then specific Chinese characters are converted according to these formal encodings.
[0004] However, the existing rare Chinese character processing methods involve parsing a large amount of data one by one to determine which fields need to be processed, which not only increases the complexity of data parsing but also leads to unnecessary parsing redundancy. In addition, after the rare Chinese characters are detected, they need to be converted into standard codes and then mapped back to Chinese characters. The multiple conversion processes are cumbersome and affect the processing efficiency. Summary of the Invention
[0005] The main purpose of the present application is to provide a rare Chinese character processing method, a rare Chinese character processing device, a computer-readable storage medium, and an electronic device, so as to at least solve the problem of low efficiency in processing rare Chinese characters in the prior art.
[0006] To achieve the above object, according to one aspect of the present application, a rare Chinese character processing method is provided, including: obtaining configuration information of rare Chinese characters and storing the configuration information in a local cache, where the configuration information includes rare Chinese character dictionary entries, rare Chinese character fields, and interfaces participating in rare Chinese character processing; in the case of receiving a data processing request, judging whether the field in the data processing request is a target field that needs to be processed for rare Chinese characters according to the configuration information; for the target field, looking up an entry in the rare Chinese character dictionary entries in the local cache that matches the Chinese character encoding in the target field, and converting the Chinese character encoding into an encoding in a preset format; in the case where there are multiple mapping entries for the Chinese character encoding, judging whether the encoding in the preset format is consistent with the pre-stored encoding in the corresponding entry in the rare Chinese character dictionary entries. If the encoding in the preset format is consistent with the pre-stored encoding, determining the rare Chinese character corresponding to the encoding in the preset format and the rare Chinese character corresponding to the pre-stored encoding as the same rare Chinese character.
[0007] Optionally, before obtaining the configuration information of rare Chinese characters, the method further includes: for each of the rare Chinese characters, creating a rare Chinese character dictionary entry for the rare Chinese character, where the rare Chinese character dictionary entry includes the rare Chinese character and various encoding information corresponding to the rare Chinese character.
[0008] Optionally, when receiving a data processing request, according to the configuration information, determining whether a field in the data processing request is a target field that needs to be processed for rare Chinese characters, including: extracting a request field identifier from the data processing request; comparing the request field identifier with the rare Chinese character field identifiers in the local cache; if the request field identifier matches any of the rare Chinese character field identifiers, determining that the field in the data processing request is the target field.
[0009] Optionally, for the target field, looking up an entry in the rare Chinese character dictionary entry in the local cache that matches the Chinese character encoding in the target field, and converting the Chinese character encoding into an encoding in a preset format, including: parsing the Chinese character encoding in the target field to determine the encoding format of the Chinese character encoding; looking up the rare Chinese character dictionary entry in the local cache that matches the encoding format of the Chinese character encoding; according to the conversion rule in the rare Chinese character dictionary entry, converting the Chinese character encoding into the encoding in the preset format.
[0010] Optionally, after converting the Chinese character encoding into an encoding in a preset format, the method further includes: updating the mapping relationship between the converted encoding in the preset format and the corresponding rare Chinese character to the local cache; after the local cache is updated, performing cache synchronization with a remote database or server.
[0011] Optionally, when there are multiple mapping entries for the Chinese character encoding, determining whether the encoding in the preset format is consistent with the pre-stored encoding in the corresponding entry in the rare Chinese character dictionary entry. If the encoding in the preset format is consistent with the pre-stored encoding, determining that the rare Chinese character corresponding to the encoding in the preset format and the rare Chinese character corresponding to the pre-stored encoding are the same rare Chinese character, including: comparing the encoding in the preset format with the pre-stored encoding in each of the mapping entries one by one; if the encoding in the preset format matches any of the pre-stored encodings in the rare Chinese character dictionary entry, determining that the rare Chinese character corresponding to the encoding in the preset format and the rare Chinese character corresponding to the matching pre-stored encoding are the same rare Chinese character.
[0012] Optionally, the method further includes: periodically updating the configuration information in the local cache to adapt to the real-time changes of the rare Chinese character dictionary entry, the rare Chinese character field, and the interface for rare Chinese character processing.
[0013] According to another aspect of the present application, a rare character processing device is provided, including: an acquisition unit, configured to acquire configuration information of rare characters and store the configuration information in a local cache, where the configuration information includes rare character dictionary entries, rare character fields, and interfaces participating in rare character processing; a judgment unit, configured to, when receiving a data processing request, judge whether a field in the data processing request is a target field that needs to be processed for rare characters according to the configuration information; a conversion unit, configured to, for the target field, find an entry in the rare character dictionary entries in the local cache that matches the Chinese character code in the target field, and convert the Chinese character code into a code in a preset format; a determination unit, configured to, when there are multiple mapping items for the Chinese character code, judge whether the code in the preset format is consistent with the pre-stored code of the corresponding entry in the rare character dictionary entries, and if the code in the preset format is consistent with the pre-stored code, determine the rare character corresponding to the code in the preset format and the rare character corresponding to the pre-stored code as the same rare character.
[0014] According to still another aspect of the present application, a computer-readable storage medium is provided, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the rare character processing methods.
[0015] According to yet another aspect of the present application, an electronic device is provided, including: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the rare character processing methods.
[0016] Applying the technical solution of the present application, first, acquire the configuration information of rare characters and store the configuration information in a local cache, where the configuration information includes rare character dictionary entries, rare character fields, and interfaces participating in rare character processing; when receiving a data processing request, judge whether a field in the data processing request is a target field that needs to be processed for rare characters according to the configuration information; then, for the target field, find an entry in the rare character dictionary entries in the local cache that matches the Chinese character code in the target field, and convert the Chinese character code into a code in a preset format; when there are multiple mapping items for the Chinese character code, judge whether the code in the preset format is consistent with the pre-stored code of the corresponding entry in the rare character dictionary entries, and if the code in the preset format is consistent with the pre-stored code, determine the rare character corresponding to the code in the preset format and the rare character corresponding to the pre-stored code as the same rare character, thereby solving the problem of low efficiency in rare character processing in the prior art. Description of the Drawings
[0017] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0018] Figure 1 shows a hardware structure block diagram of a mobile terminal for implementing a rare Chinese character processing method provided in an embodiment of this application;
[0019] Figure 2 shows a schematic flowchart of a rare Chinese character processing method provided in an embodiment of this application;
[0020] Figure 3 shows a display diagram of a rare Chinese character dictionary item list of a rare Chinese character processing method provided in an embodiment of this application;
[0021] Figure 4 shows a schematic diagram of a dynamic configuration of a rare Chinese character processing interface and a field list of a rare Chinese character processing method provided in an embodiment of this application;
[0022] Figure 5 shows a schematic diagram of the implementation process of rare Chinese character matching of a rare Chinese character processing method provided in an embodiment of this application;
[0023] Figure 6 shows a flowchart of loading a font library to a local cache at one time for a rare Chinese character processing method provided in an embodiment of this application;
[0024] Figure 7 shows a structure block diagram of a rare Chinese character processing device provided in an embodiment of this application.
[0025] Among them, the above-mentioned drawings include the following reference numerals:
[0026] 102, a processor; 104, a memory; 106, a transmission device; 108, an input / output device. Detailed implementation manners
[0027] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0028] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0029] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so as to implement the embodiments of this application described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0030] As introduced in the background art, the existing methods for handling rare Chinese characters involve parsing a large amount of data one by one to determine which fields need to be processed, which not only increases the complexity of data parsing but also leads to unnecessary parsing redundancy. In addition, after the rare Chinese characters are detected, they need to be converted into standard codes and then mapped back to Chinese characters, and the multiple conversion processes are cumbersome, affecting the processing efficiency. To solve the problem of low efficiency in handling rare Chinese characters in the existing technology, the embodiments of this application provide a method for handling rare Chinese characters, a device for handling rare Chinese characters, a computer-readable storage medium, and an electronic device.
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0032] The method embodiments provided in the embodiments of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a method of handling rare Chinese characters according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in Figure 1 a processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in Figure 1 is only schematic and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than those shown in
[0033] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the display method of device information in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0034] In this embodiment, a rare Chinese character processing method running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0035] Figure 2 It is a schematic flowchart of the rare Chinese character processing method according to the embodiments of the present application. As Figure 2 shown, the method includes the following steps:
[0036] Step S201, obtain the configuration information of the rare Chinese character and store the above configuration information in the local cache, where the above configuration information includes rare Chinese character dictionary items, rare Chinese character fields, and interfaces participating in rare Chinese character processing;
[0037] Specifically, the rare character dictionary entry contains a detailed list of all known rare characters, including the dictionary name, dictionary value, character set, and remarks. Each rare character dictionary entry is associated with multiple possible encoding formats, such as PUA encoding, extended A encoding, etc. The establishment of such dictionary entries is based on the previous research and collection work of a large number of rare characters, ensuring that the system can identify and process rare characters under various encodings. The rare character field defines which fields in the data processing request may contain rare characters and require special processing. For example, fields such as names may contain rare characters, so these fields are marked as target fields for rare character processing in the configuration information. The interfaces involved in rare character processing list all the system interfaces that need to perform rare character processing. Since rare characters may appear in multiple business scenarios and processing links, it is necessary to clarify which interfaces need to execute the rare character recognition and conversion logic when called to avoid unnecessary resource consumption and processing delays.
[0038] The use of local cache is to solve the performance bottleneck caused by frequent database queries. By loading the rare character dictionary entry, rare character field, and interface information involved in rare character processing into the local cache at one time, the subsequent rare character processing process can directly obtain the required information from the local cache quickly without accessing the database every time, thus significantly improving the efficiency and response speed of data processing.
[0039] In summary, this step provides a basis for quick response for subsequent rare character recognition and processing by preloading key rare character configuration information into the local cache. At the same time, by restricting the processing scope to specific fields and interfaces, it further optimizes resource allocation and ensures the efficient operation of the system. This strategy is particularly applicable to scenarios with high concurrency and large data volume processing, such as bank transaction systems, large database applications, etc., providing a strong guarantee for the consistent and efficient processing of rare characters.
[0040] Step S202, in the case of receiving a data processing request, according to the above configuration information, determine whether the field in the data processing request is a target field that needs to perform rare character processing;
[0041] Specifically, first, a data processing request will be received. Taking banking business as an example, the data processing request can come from a certain business process within the bank, such as account information update, transaction record query, etc., or from an external system, such as an interface request for data exchange with customers or partners. The data processing request usually contains multiple fields, and each field represents different data information, such as customer name, address, transaction amount, etc.
[0042] Based on the rare character configuration information previously stored in the local cache, determine which fields in the data processing request are the target fields that need to be focused on and processed. The configuration information contains the identifiers of rare character fields, that is, which fields are preset to possibly contain rare characters and need special processing. For example, the customer name may contain rare characters. Therefore, the customer name is used as the target field for rare character processing. The introduction of this judgment step greatly improves the efficiency and accuracy of rare character processing. It avoids the indiscriminate rare character recognition for all fields in the request, reduces unnecessary consumption of computing resources, and also ensures that the processing logic only targets the fields that may contain rare characters, thereby enhancing the system's response speed and performance. In environments with high data volume processing such as banks, this optimization is particularly important and can significantly improve the smoothness of business processing and the customer experience.
[0043] Step S203, for the above target fields, search in the above-mentioned rare character dictionary entries in the above local cache for entries that match the Chinese character encoding in the above target fields, and convert the above Chinese character encoding into an encoding in a preset format.
[0044] Specifically, when it is determined that a certain field in the data processing request is a target field that needs to be processed for rare characters, the Chinese character encoding in this field will be further analyzed. The encoding can be in various forms, such as common encodings like GB2312, GBK, UTF-8, and special encodings for rare characters like PUA, Extended A, or Extended E. The system needs to be able to parse the text information in the field and identify the type of Chinese character encoding. This step is crucial for subsequent encoding conversion. Once the Chinese character encoding in the target field is identified, the rare character dictionary entries in the local cache are used for quick matching. The local cache contains the encoding information of all known rare characters and the mapping relationship between these encodings and the actual characters of the rare characters. Search for entries that match the Chinese character encoding in the target field in the rare character dictionary entries in the local cache. The use of the local cache greatly speeds up this matching process and avoids the inefficient operation of accessing the database every time.
[0045] When a matching dictionary entry is found, according to the conversion rules in the dictionary entry, convert the Chinese character encoding in the target field into an encoding in a preset format. Usually, this preset format is a more general and standard encoding format, such as Unicode. This conversion ensures that regardless of which encoding format the original input uses, the encoding of rare characters can be unified into a standard format, facilitating subsequent processing and compatibility. By uniformly converting all possible rare character encodings into a preset format, data consistency can be guaranteed in different data sources and different business scenarios. For those rare characters that cannot be correctly parsed in common encodings, after being converted into the preset format, it can be ensured that they can be correctly displayed in the system, avoiding problems such as garbled characters or character missing caused by inconsistent encodings.
[0046] In summary, this step involves identifying the Chinese character encoding in the target field, quickly matching it with the rare character dictionary entries in the local cache, and then converting the matched Chinese character encoding into a preset format encoding to achieve the correct recognition, processing, and display of rare characters. This process is of great significance for improving the efficiency and accuracy of processing rare characters and ensuring data consistency.
[0047] Step S204: In the case where there are multiple mapping entries for the above Chinese character encoding, determine whether the preset format encoding is consistent with the pre-stored encoding of the corresponding entry in the rare character dictionary. If the preset format encoding is consistent with the pre-stored encoding, then determine the rare character corresponding to the preset format encoding and the rare character corresponding to the pre-stored encoding as the same rare character.
[0048] Specifically, the same rare character can have multiple encoding forms. For example, a rare character may have its own encoding in different character sets such as PUA, Extended A, and Extended E. When the Chinese character encoding in the target field is recognized as corresponding to a rare character, the rare character dictionary entries in the local cache will be checked to confirm whether there are multiple mapping entries for this rare character. These mapping entries record the multiple encodings of the rare character and their equivalence relationships. In the case where multiple mapping entries are recognized, the preset format encoding (usually a preset unified encoding format such as Unicode) of the Chinese character in the target field is compared one by one with the pre-stored encoding in the rare character dictionary. The preset format encoding here refers to the encoding format used for internal storage and transmission when processing rare characters, aiming to provide a consistent encoding benchmark so that no matter in which encoding form the rare character is input, it can be uniformly processed.
[0049] Next, check whether the preset format encoding matches any of the pre-stored encodings in the rare character dictionary. If the match is successful, it means that the Chinese character encoding in the target field is equivalent to one of the encodings recorded in the dictionary entry, although they may have different forms in different encoding standards. When the preset format encoding is consistent with a certain pre-stored encoding in the dictionary entry, it can be determined that the two represent the same rare character. Even if the encoding in the target field is not exactly the same as the encoding recorded in the dictionary entry, as long as they are recognized as equivalent in the preset mapping relationship, they can be regarded as the same rare character. This determination process is crucial for ensuring the accuracy and consistency of data, especially when dealing with data exchanges across systems or encoding standards.
[0050] The realization of this process solves the recognition and processing problems caused by the diversity of rare characters, ensuring that rare characters can be accurately recognized as the same character even when there are multiple codes for one character, which is of great significance for maintaining data accuracy and improving user experience. At the same time, by unifying multiple encoding forms into a preset format encoding, the processing logic within the system is simplified, reducing the complexity of rare character processing.
[0051] In one embodiment of the present application, before obtaining the configuration information of the rare characters, the method further includes: for each of the rare characters, creating a rare character dictionary item for the rare character, the rare character dictionary item including the rare character and multiple encoding information corresponding to the rare character.
[0052] Specifically, before any rare character processing begins, a comprehensive rare character dictionary entry database needs to be established. This database contains detailed information of all known rare characters. For each rare character, the system will create a rare character dictionary entry, which includes not only the actual characters of the rare character, but also all the encoding information that the rare character may represent under different encoding standards.
[0053] The rare character dictionary item contains a detailed list of all known rare characters, including dictionary name, dictionary value, character set, and remarks. Figure 3 As shown, the dictionary name column shows the encoding identifier of the rare characters, which is usually expressed in the form of encoding standards such as Unicode or PUA. For example, "U+E863", "U+4DAE", "U+2CC56" and "U+9814" represent the encoding of specific rare characters respectively. Among them, "U+" is the prefix of the Unicode encoding, indicating that this is a rare character encoded in Unicode. The dictionary value column lists the corresponding rare characters, such as "頔". The dictionary value is the actual display character of the rare character, which is the correspondence between the encoding in the dictionary item and the actual character. The character set column identifies the encoding character set to which the rare character belongs, such as "PUA" (Private Use Area), "Extended A" and "Extended E". Different character sets usually mean different encoding ranges and usage scenarios. This information helps to choose the correct conversion strategy and encoding standard when processing rare characters. The remarks column provides additional information, such as "one character with multiple codes" or "Simplified characters cannot be displayed, and traditional characters are used for display". Notes are used to explain the special features of dictionary items, such as a Chinese character may have multiple codes (one character has multiple codes), or in some environments, rare characters can only be displayed in traditional Chinese characters (simplified Chinese characters cannot be displayed). These notes help better understand and deal with the complexity of rare characters.
[0054] In addition, the "Add", "Delete", and "Modify" buttons provide management functions for dictionary entries. Through these buttons, new rare-character dictionary entries can be added, unnecessary dictionary entries can be removed, or the encoding and character information of existing dictionary entries can be updated, realizing the dynamic maintenance of the rare-character library. The "Query" and "Reset" buttons are used to search for corresponding dictionary entries according to the input dictionary name or dictionary value, and to clear the search conditions and restore to the initial state, respectively. The query function helps users quickly locate specific rare-character information, while the reset function provides a convenient way to clear the search conditions and browse all dictionary entries again.
[0055] This process of creating rare-character dictionary entries is a prerequisite for ensuring that the rare-character processing system can correctly and efficiently identify and handle the problem of multiple codes for a single character. By pre-building a database containing rare characters and their encoding information, it is possible to quickly determine whether the rare-character encoding in the target field exists in the dictionary entries when receiving a data processing request, and then perform the necessary encoding conversion to ensure the correct display of rare characters and data consistency.
[0056] In order to quickly identify the fields that may carry rare characters in the data processing request, and thus apply the rare-character processing logic specifically, when receiving a data processing request, according to the above configuration information, it is determined whether the fields in the above data processing request are the target fields that need to be processed for rare characters, including: extracting the request field identifier from the above data processing request; comparing the above request field identifier with the rare-character field identifiers in the above local cache; if the above request field identifier matches any of the above rare-character field identifiers, it is determined that the above field in the above data processing request is the above target field.
[0057] Specifically, when receiving a data processing request, the request usually contains a series of fields, each with its specific identifier used to distinguish and identify different data types or contents. First, the request will be parsed to extract the identifier of the request field, which can be the field name, field ID, or other form of unique identifier.
[0058] The request field identifier is compared with the rare-character field identifiers in the local cache. The local cache stores a list of rare-character field identifiers. These identifiers are loaded from the database and include all preset fields that need to be processed for rare characters. The request field identifier extracted from the data processing request is compared one by one with the rare-character field identifiers in the local cache to find a match. If the request field identifier matches any of the rare-character field identifiers in the local cache, it means that the data in this field needs to be processed for rare characters. This match can be a direct field name match or a match based on the field ID or other identifiers. Once a match is successful, this field will be marked as the "target field", that is, a field that may contain rare characters and requires special processing.
[0059] By comparing the request field identifier with the rare character field identifier, it is possible to quickly and accurately determine which fields are target fields and require rare character processing. This determination process avoids the indiscriminate inspection of all fields and significantly improves the processing efficiency. Only the data marked as "target fields" will be further analyzed and processed to identify and convert the rare character codes that may be included therein.
[0060] The core purpose of this series of steps is to quickly identify the fields in the data processing request that may carry rare characters through efficient field identifier matching, so as to apply the rare character processing logic targeted. This method not only reduces the unnecessary consumption of computing resources, improves the processing speed, but also ensures the accuracy of rare character processing and the consistency of data.
[0061] Figure 4 What is shown is the schematic diagram of the dynamic configuration of the rare character processing interface and the field list. The dynamic configuration of the rare character processing interface and the field list play a key role in configuration and management in the rare character processing system. Among them, the column of interface names lists the names of all interfaces that need to perform rare character processing. In banking business, the interface name is the communication bridge between various services, such as "fun1", "fun2", etc. These interfaces may involve business scenarios such as customer information management and account transaction record processing. By listing the interface names in this column, it is possible to intuitively see which interfaces need special attention in order to perform special processing on rare characters. The column of field names corresponds to the column of interface names, and this column details the field names involved in each interface. Field names such as "custName", "ownerName", etc. specifically refer to various data types included in the interface request or response, such as customer name, account holder name, etc. The clear identification of field names helps the system to quickly locate the specific fields that may contain rare characters when processing interface requests. The column of class names contains the class names of the code for processing the corresponding fields. The class name (such as "com.test.a.Class1") is a unit used in programming to organize code and manage functions. The class names listed here point to the code modules that specifically implement the rare character processing logic. Through this column, it is possible to perform code-level tracking and maintenance on the rare character processing of specific fields to ensure the accuracy and consistency of processing. The remarks column provides additional information or explanations. For example, the field "custName" in the "fun1" interface needs to be processed specially locally, perhaps because the frequency of rare characters included in this field is relatively high, or due to business requirements, the data accuracy in this field is particularly important. The remarks column helps to understand why certain interfaces and fields are selected for special processing, as well as possible processing strategies or precautions.
[0062] "Add", "Delete", "Modify" buttons, as well as "Query" and "Reset" buttons, provide the dynamic management ability for the rare character processing interface and field configuration. These buttons allow adding new interfaces and fields in real time according to actual needs, deleting configurations that are no longer required, or modifying the details of existing configurations, such as interface names, field names, processing logics, etc. In addition, the "Query" button helps quickly retrieve the configurations of specific interfaces or fields, and the "Reset" button is used to restore the configuration to the initial state, facilitating batch or rapid configuration adjustments.
[0063] Two input boxes are respectively used to input the "interface name" and "field name". These two search boxes provide the ability to quickly locate specific interface and field configurations, helping to quickly find the objects that need attention among a large number of interfaces and fields for viewing or modifying the configurations.
[0064] By dynamically configuring the rare character processing interface and field list, the flexible response to the rare character processing requirements is realized. It not only allows dynamically adjusting which interfaces and fields need to perform rare character processing according to the business scenario, but also optimizes the processing performance through the local cache mechanism, avoiding unnecessary database access, and ensuring data consistency and processing efficiency. This dynamic configuration ability is the key for the rare character processing system to adapt to complex and changeable business environments and effectively manage and optimize the rare character processing process.
[0065] In another embodiment of the present application, for the above target field, find an entry in the above rare character dictionary items cached locally that matches the Chinese character encoding in the above target field, and convert the above Chinese character encoding into an encoding in a preset format, including: parsing the Chinese character encoding in the above target field to determine the encoding format of the above Chinese character encoding; finding the above rare character dictionary item in the above local cache that matches the encoding format of the above Chinese character encoding; according to the conversion rule in the above rare character dictionary item, converting the above Chinese character encoding into the above preset format of encoding.
[0066] Specifically, when a certain field is determined to be the target field that needs to perform rare character processing, first, the encoding format of the text data in this field will be parsed. This process involves identifying the specific encoding form of Chinese characters in the text, and the encoding may include but is not limited to PUA encoding, extended A or extended E encoding in GB18030, etc. Parsing the Chinese character encoding is the primary step of encoding conversion to ensure the understanding and processing of rare character information in the target field.
[0067] The local cache contains dictionary entries for all rare Chinese characters and their corresponding multiple encoding information. After parsing the Chinese character encoding in the target field, the encoding is compared with the dictionary entries of rare Chinese characters in the local cache to check if there is a dictionary entry that matches the current encoding format. Since the local cache has pre-loaded all the necessary information about rare Chinese characters, this lookup process can be completed quickly and efficiently, avoiding frequent access to the database and improving the system's response speed.
[0068] A dictionary entry that matches the Chinese character encoding in the target field is found in the local cache. The conversion rules within the dictionary entry are further applied to convert the Chinese character encoding in the current field into an encoding in a preset format. The encoding in the preset format may be a unified encoding used internally in the system, such as Unicode, which can ensure that rare Chinese characters can be correctly understood and processed. The conversion rules detail how to convert from one encoding format (such as a specific PUA or GB18030 encoding) to another encoding format (such as Unicode). This process is the core of rare Chinese character processing, solving the problem of rare Chinese character display caused by inconsistent encodings and improving data compatibility and consistency.
[0069] According to the conversion rules within the dictionary entry of rare Chinese characters, the Chinese character encoding in the target field is converted from its original format to a preset format. This involves mapping from a special character set encoding to Unicode encoding, ensuring that even in the case of multiple encodings for a single character, rare Chinese characters can be correctly recognized and displayed. The conversion process includes looking up the mapping table in the dictionary entry, matching specific Chinese character encodings with the preset format encoding one-to-one or one-to-many, and performing the corresponding conversion operations.
[0070] By performing the above steps, the problem of rare Chinese character encoding in the target field can be effectively handled, ensuring the consistency and correctness of rare Chinese characters under different encoding standards and business scenarios. This function not only improves data processing capabilities but also enhances the user experience. Especially in the financial field, it provides important technical support for ensuring the accurate display and processing of customer information.
[0071] To ensure that the latest rare Chinese character encoding conversion information can be quickly accessed during subsequent processing and improve processing efficiency, after converting the above Chinese character encoding into an encoding in a preset format, the above method further includes: updating the mapping relationship between the converted encoding in the preset format and the corresponding rare Chinese character to the above local cache; after the above local cache is updated, performing cache synchronization with the remote database or server.
[0072] Specifically, when the Chinese character encoding in the target field is successfully converted to the encoding in the preset format (such as Unicode), this new mapping relationship (i.e., the conversion from the original encoding to the encoding in the preset format) is updated to the local cache. The local cache is a high-speed storage area that loads rare Chinese character dictionary entries and other configuration information when the system starts, significantly reducing the system's access frequency to the remote database, thereby improving the data processing speed. Updating the mapping relationship in the local cache ensures that the latest rare Chinese character encoding conversion information can be quickly accessed during subsequent processing, improving the processing efficiency. The implementation process is as follows: by creating or modifying the entries of the corresponding rare Chinese character dictionary items in the local cache, adding the new encoding in the preset format to the encoding list of the corresponding rare Chinese characters, ensuring that the latest status of the rare Chinese characters and their encodings is always accessible in the local cache.
[0073] After the local cache is updated, it is necessary to synchronize the cache with the remote database or server to ensure that all relevant systems have a consistent data state. This synchronization process is an important guarantee for the system's robustness and data consistency. The remote database or server provides data or configuration information for other services or systems. By synchronizing the cache, it can be ensured that all systems operate based on the latest rare Chinese character dictionary items and mapping relationships, preventing errors or exceptions caused by data inconsistency. Usually, an efficient data synchronization mechanism, such as a message queue, event-driven update notification, or timed task, is used to synchronize the cache with the remote database or server. This mechanism can push the updated information in the local cache to the remote system in real-time or periodically, maintaining data consistency and integrity. The synchronization process is as follows: send an update request to the remote database or server, requesting to update or add entries in the rare Chinese character dictionary to reflect the latest encoding conversion information in the local cache. After receiving the update request, the remote system will verify the correctness of the information and then update its own database or cache, finally achieving data synchronization.
[0074] By updating the mapping relationship in the local cache and synchronizing it to the remote database or server, it can ensure the consistency and efficiency of the rare Chinese character processing logic. Even in a high-concurrency and multi-system environment, it can accurately identify and process rare Chinese characters, avoiding technical problems that may be caused by data inconsistency.
[0075] In another embodiment of the present application, when there are multiple mapping items for the above-mentioned Chinese character encoding, it is determined whether the encoding in the above-mentioned preset format is consistent with the pre-stored encoding of the corresponding entry in the above-mentioned rare character dictionary item. If the encoding in the above-mentioned preset format is consistent with the pre-stored encoding, the rare character corresponding to the encoding in the above-mentioned preset format and the rare character corresponding to the pre-stored encoding are determined to be the same rare character, including: comparing the encoding in the above-mentioned preset format with the above-mentioned pre-stored encoding in each of the above-mentioned mapping items one by one; if the encoding in the above-mentioned preset format matches any of the above-mentioned pre-stored encodings in the above-mentioned rare character dictionary item, it is determined that the rare character corresponding to the encoding in the above-mentioned preset format and the rare character corresponding to the matching pre-stored encoding are the same rare character.
[0076] Specifically, after the encoding conversion of the rare character is completed, the preset format encoding obtained by the conversion (such as the converted Unicode encoding) is compared with the pre-stored encoding of the corresponding entry in the rare character dictionary item in the local cache. The pre-stored encoding refers to the encoding form that has been recorded in the rare character dictionary item and is associated with the rare character. This comparison process is to check whether the newly converted encoding has been recorded in the dictionary item as another encoding representation of the rare character. Each pre-stored encoding is compared one by one to find out whether there is a case where the encoding is consistent. If the encoding of the preset format matches any pre-stored encoding in the rare character dictionary item, it can be determined that the two represent different encoding forms of the same rare character. For example, assuming that the dictionary item records the Chinese character "顖" with two pre-stored encodings of "U+2CC56" (PUA encoding) and "U+9814" (GB18030 encoding), if the preset format encoding after conversion happens to be "U+9814", then it is recognized that this encoding actually represents the character "顖".
[0077] When the preset format encoding matches any of the pre-stored encodings, it is determined that the uncommon character represented by the preset format encoding and the uncommon character corresponding to the pre-stored encoding are the same Chinese character. This confirmation process is based on the consistency of the encoding, which ensures that in the scenario of processing multiple codes for one character, uncommon characters can be correctly identified, avoiding glyph recognition errors caused by encoding differences. For example, in banking business, if the character "顖" encoded in "U+9814" and the character "顖" encoded in "U+2CC56" appear in two different transaction requests, by comparison, it is found that "U+9814" matches the "U+9814" encoding of the character "顖" in the dictionary item, thereby determining that the two encodings actually represent the same "顖" character, ensuring the consistency of data processing and the correctness of business logic.
[0078] Through the above steps, not only can the problem of multiple codes for one character be efficiently processed, but also the consistency and accuracy of data can be maintained in the recognition and processing of rare characters. This mechanism is an important guarantee for the system in the face of complex coding situations, ensuring that even in an environment with variable coding forms, rare characters can be accurately recognized and processed.
[0079] For the specific implementation process of rare character matching, please refer to Figure 5 . The process description is as follows:
[0080] 1) Input two characters to be compared. In this step, Chinese character information from two different data processing requests is received, and each Chinese character may have different coding forms. These characters are called name1 and name2, which are the coding representations of the rare characters to be matched and judged.
[0081] 2) Conventionally judge whether name1 and name2 are the same character. First, try to use conventional means to judge whether name1 and name2 represent the same Chinese character. In this application, the Java programming language is used, and the isEqual method is used to judge whether name1 and name2 are the same character; if so, the conclusion "name1 and name2 are the same character" is drawn, and the process ends; otherwise, proceed to the next step;
[0082] 3) Respectively judge whether there is a dictionary mapping relationship for name1 and name2 in the local cache (it is necessary to convert name1 and name2 into a unified code in advance, such as the Unicode format). If the conventional judgment fails to confirm that name1 and name2 are equal, it will further check whether there is a mapping relationship in the rare character dictionary items cached locally. The mapping relationship means that in the dictionary items, there is one or more unified codes or glyphs that can be associated with name1 and name2.
[0083] a) If there are dictionary mapping relationships for both name1 and name2, and the values value1 and value2 of their dictionary items are equal, then name1 and name2 are the same character;
[0084] b) If name1 has a dictionary item mapping relationship and name2 does not have a dictionary item mapping relationship, then name1 and name2 are not the same character;
[0085] c) If name1 does not have a dictionary item mapping relationship and name2 has a dictionary item mapping relationship, then name1 and name2 are not the same character.
[0086] Among them, value1 is the standard name or glyph of the uncommon character associated with the name1 encoding recorded in the uncommon character dictionary item. For example, if name1 is encoded as "U+E863", then value1 may be the Unicode encoding of the character "隹" (such as "U+964E") or other forms of standard representation. value2 represents the standard name or glyph of the uncommon character associated with the name2 encoding in the uncommon character dictionary item. For example, if name2 is encoded as "U+4DAE", then value2 may also be another encoding representation of the character "隹" or its standard glyph.
[0087] In another embodiment of the present application, the method further comprises: periodically updating the configuration information of the local cache to adapt to real-time changes of the rare-character dictionary items, the rare-character fields and the rare-character processing interface.
[0088] Specifically, the present application configures a periodic update mechanism that regularly checks and updates the configuration information in the local cache to ensure that the information is consistent with the latest rare character dictionary items, rare character fields, and the status of the rare character processing interface. The update of the local cache is the key to the rare character processing method to continue to maintain high efficiency and accuracy, because the encoding of rare characters, the fields to be processed, and the interface configuration may change over time, such as with the inclusion of new characters, the adjustment of business needs, or the improvement of external system compatibility.
[0089] The rare character dictionary entry contains multiple encoding methods for rare characters and their corresponding dictionary values (i.e., standard glyphs or encodings). The periodic update mechanism ensures that when the content of the dictionary item changes, such as adding a rare character or updating the encoding list of a rare character, the dictionary item information in the local cache can quickly reflect these changes. In this way, when processing new or updated rare characters, the latest version of the dictionary item information can be called immediately without frequent access to the database, which improves the real-time and accuracy of rare character processing.
[0090] Rare characters fields are fields in the system that require special attention and processing. These fields may contain rare characters in user names, account information, or other text data. As business processes are updated, some fields may be newly identified as containing rare characters, or the original rare characters fields may no longer require special processing. Periodically updating the rare characters field configuration in the local cache can ensure that the system can respond to these changes in a timely manner, process fields containing rare characters with the best strategy, avoid unnecessary processing actions, and ensure that all fields that need to be processed are properly recognized and converted.
[0091] The rare character processing interface is a logical unit dedicated to processing rare character data, distributed across multiple microservices or business processing modules. The configuration information of the interface (such as interface name, parameters, etc.) will change with the evolution of business requirements and system architecture. By periodically updating the configuration of the rare character processing interface in the local cache, it can be ensured that under the new interface configuration, the rare character processing logic can still operate efficiently, timely capture and process the data containing rare characters passed in through these interfaces, and avoid data processing delays or errors that may be caused by changes in interface configuration.
[0092] A timed task can be configured to automatically check the latest configuration information in the database and update the local cache at regular intervals (such as daily, weekly, or at specific time points). Once a change in the rare character dictionary items, fields, or interface configuration is detected, the update of the local cache is immediately triggered to achieve real-time synchronization of the configuration information. The update of the local cache can also be manually triggered when data or configuration changes.
[0093] Through the periodic update mechanism, it can better adapt to the rapidly changing business requirements and external environment, ensuring the accuracy and timeliness of data processing, which is particularly important for data processing in the financial field, ensuring business continuity and consistency of the customer experience. This mechanism provides a dynamic optimization framework for rare character processing.
[0094] The application of this application to the overall process of banking business is as follows:
[0095] 1) The user configures rare character fields and interfaces involved in rare character processing on the interface, and these information will be stored in the database after the user's operation is successful;
[0096] 2) The backend service starts to read these user configuration information from the database and cache them in the local dictionary items: for example The uncode encoding of the character can have three values: E05D, F429, 39D1;
[0097] 3) When a transaction request triggers the rare character processing interface (according to the interface configuration information), a rare character judgment process will be carried out, and the Chinese character encoding in the request will be converted into a unified unicode encoding. For example There are two transaction requests, one of which uses the encoding E05D,
[0098] 4) The other uses the encoding 39D1, and the user's Chinese character encoding saved in the business system is F429. Therefore, these two transaction requests are still considered to be the requested This Chinese character is used to continue the transaction. It should be noted here that the Chinese character information carried by the two requests is not necessarily the two string information of E05D and 39D1. It is necessary to use a unified encoding for special processing and convert it into E05D and 39D1. That is, the two Chinese characters are respectively encoded in hexadecimal.
[0099] In order to enable those skilled in the art to more clearly understand the technical solution of this application, the following will take the banking business background as an example to describe in detail the implementation process of the rare Chinese character processing method of this application.
[0100] Suppose there is a customer named "Tan Feng" in the bank's customer information database, where the character "Tan" is a rare Chinese character, which has different representations in different coding systems, such as the PUA coding "U+E863", the GB18030 coding "U+4DAE", and some special coding. When the bank system performs customer information entry, transaction record processing, or customer service, it needs to process such rare Chinese characters with multiple encodings to ensure data consistency and accuracy.
[0101] Construct rare Chinese character dictionary entries: The IT department maintains a rare Chinese character dictionary in the system, which contains multiple coding mapping relationships of a rare Chinese character "Tan" (the dictionary value is "Tan"). For example, the character "Tan" is represented as "U+E863" in the PUA coding system and "U+4DAE" in the GB18030 coding. The flowchart for loading the font library into the local cache at one time is as Figure 6 shown. When the system starts, it will load all the encodings of the rare Chinese characters in the dictionary entries into the local cache at one time for subsequent quick reading and use. The specific process is as follows: The microservice program starts; during the program startup process, all the rare Chinese character dictionary information and interface field configuration information in the database are read; the read dictionary information and interface field configuration information are stored in the local cache.
[0102] Dynamic configuration field processing: The IT department of the bank dynamically configures the fields that need to be processed for rare Chinese characters, such as the "customer name" field in the "account opening interface". This means that when the system processes a customer account opening request, it will pay special attention to the "customer name" field and check whether it contains rare Chinese characters.
[0103] Transaction request processing: When the bank system receives a customer account opening request, the "Tan Feng" in the request may use the PUA coding "U+E863". The system first reads the rare Chinese character dictionary entries in the local cache for quick matching.
[0104] Uncommon character matching check: The system first tries a conventional judgment, that is, comparing the "镡" in "镡锋" with the dictionary item in the local cache. Since the character "镡" has multiple encodings, the system will check whether "U+E863" is mapped to "镡" in the dictionary item. If there is a mapping, it is confirmed that the PUA encoding of the character "镡" is consistent with the mapping relationship in the dictionary, and the system will determine that the "镡" in this request and the known "镡" are different encoding representations of the same rare character.
[0105] Periodic update of configuration information: In order to cope with the real-time changes in rare-character dictionary items, rare-character fields, and rare-character processing interfaces, the bank system is configured with a scheduled task to periodically (for example, automatically update every night) check the latest configuration information in the database and update the local cache. For example, if another encoding of the character "镡" is added to the database or the rare-character processing strategy of the "customer name" field is adjusted, these changes will be automatically reflected in the local cache during the next update without manual intervention.
[0106] Implementation effect:
[0107] Improved efficiency: By loading the dictionary of rare characters into the local cache once, the banking system avoids redundant access to the database for each transaction request, greatly reducing data processing time and system response time, and improving the overall business processing speed.
[0108] Accuracy guarantee: Utilizing the efficient search of the data dictionary, the system can accurately identify and process rare characters with multiple codes, avoiding data errors or customer service issues caused by encoding differences.
[0109] Enhanced flexibility: The ability to dynamically configure rare-character fields and processing interfaces enables the banking system to quickly adapt to changes in business needs. Rare-character processing strategies can be adjusted without modifying core code or restarting services, improving the maintainability and flexibility of the system.
[0110] Real-time guarantee: The mechanism of periodically updating the local cache configuration information ensures that the rare word processing strategy in the system can reflect the latest database changes in real time, ensuring the real-time and accuracy of data processing.
[0111] Through this specific example, we can see the implementation process of this application in the context of banking and the significant advantages it brings, including improving data processing efficiency, ensuring data accuracy, enhancing system flexibility, and ensuring real-time performance. These advantages play an important role in the stable operation of the financial system and improving customer satisfaction.
[0112] Of course, the solution of the present application can also be widely applied in multiple fields and backgrounds. For example, in the education industry, such as online education platforms, e-book systems, academic paper databases, etc., various literature materials need to be processed, which may contain rare Chinese characters in ancient or professional terms. This solution helps to improve the character recognition and processing capabilities of these platforms; in the medical and health industry, in medical records and health archives, rare Chinese characters that may be included in patient names and disease names need to be accurately recognized and processed to maintain the reliability of medical data; in media and publishing, such as online news platforms, e-publications, digital libraries, etc., Chinese content containing rare Chinese characters needs to be processed and displayed to ensure the accuracy of information dissemination and the reading experience; in addition, it can also be applied to fields such as mobile applications and smart devices, e-commerce, file management, and historical research.
[0113] The embodiment of the present application also provides a rare Chinese character processing device. It should be noted that the rare Chinese character processing device of the embodiment of the present application can be used to execute the rare Chinese character processing method provided by the embodiment of the present application. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0114] The following introduces the rare Chinese character processing device provided by the embodiment of the present application.
[0115] Figure 7 is a structural block diagram of the rare Chinese character processing device according to the embodiment of the present application. As Figure 7 shown, the device includes: an acquisition unit 10, a judgment unit 20, a conversion unit 30, and a determination unit 40.
[0116] The acquisition unit 10 is used to acquire the configuration information of rare Chinese characters and store the above configuration information in the local cache, where the above configuration information includes rare Chinese character dictionary entries, rare Chinese character fields, and interfaces participating in rare Chinese character processing;
[0117] The judgment unit 20 is used to, when receiving a data processing request, judge whether the field in the data processing request is a target field that needs to be processed for rare Chinese characters according to the above configuration information;
[0118] The conversion unit 30 is used to, for the above target field, find an entry in the above rare Chinese character dictionary entries in the local cache that matches the Chinese character encoding in the above target field, and convert the above Chinese character encoding into a preset format encoding;
[0119] A determination unit 40, configured to determine, when there are multiple mapping items for the above-mentioned Chinese character encoding, whether the encoding in the above-mentioned preset format is consistent with the pre-stored encoding of the corresponding entry in the above-mentioned rare Chinese character dictionary entry. If the encoding in the above-mentioned preset format is consistent with the above-mentioned pre-stored encoding, the rare Chinese character corresponding to the encoding in the above-mentioned preset format and the rare Chinese character corresponding to the above-mentioned pre-stored encoding are determined to be the same rare Chinese character.
[0120] In an embodiment of the present application, before obtaining the configuration information of the rare Chinese character, the above-mentioned device further includes:
[0121] A creation unit, configured to create, for each of the above-mentioned rare Chinese characters, a rare Chinese character dictionary entry for the above-mentioned rare Chinese character, where the rare Chinese character dictionary entry includes the above-mentioned rare Chinese character and various encoding information corresponding to the above-mentioned rare Chinese character.
[0122] In order to quickly identify the fields in the data processing request that may carry rare Chinese characters, so as to apply the rare Chinese character processing logic in a targeted manner, the above-mentioned judgment unit includes:
[0123] An extraction module, configured to extract a request field identifier from the above-mentioned data processing request;
[0124] A comparison module, configured to compare the above-mentioned request field identifier with the rare Chinese character field identifiers in the above-mentioned local cache;
[0125] A first determination module, configured to determine, if the above-mentioned request field identifier matches any of the above-mentioned rare Chinese character field identifiers, that the above-mentioned field in the above-mentioned data processing request is the above-mentioned target field.
[0126] In another embodiment of the present application, the above-mentioned conversion unit includes:
[0127] A second determination module, configured to parse the Chinese character encoding in the above-mentioned target field and determine the encoding format of the above-mentioned Chinese character encoding;
[0128] A search module, configured to search in the above-mentioned local cache for the above-mentioned rare Chinese character dictionary entry that matches the encoding format of the above-mentioned Chinese character encoding;
[0129] A conversion module, configured to convert the above-mentioned Chinese character encoding into the encoding in the above-mentioned preset format according to the conversion rule in the above-mentioned rare Chinese character dictionary entry.
[0130] After converting the above-mentioned Chinese character encoding into the encoding in the preset format, the above-mentioned device further includes:
[0131] An update unit, configured to update the mapping relationship between the converted encoding in the above-mentioned preset format and the corresponding rare Chinese character to the above-mentioned local cache;
[0132] A synchronization unit, configured to perform cache synchronization with a remote database or server after the above-mentioned local cache is updated.
[0133] In another embodiment of the present application, the above-mentioned determination unit includes:
[0134] A one-by-one comparison module, configured to compare the above-mentioned encoded data in the preset format with the above-mentioned pre-stored encoded data in each of the above-mentioned mapping items one by one;
[0135] A third determination module, configured to determine that the rare Chinese character corresponding to the above-mentioned encoded data in the preset format and the rare Chinese character corresponding to the matched above-mentioned pre-stored encoded data are the same rare Chinese character if the above-mentioned encoded data in the preset format matches any of the above-mentioned pre-stored encoded data in the rare Chinese character dictionary item.
[0136] In yet another embodiment of the present application, the above-mentioned device further includes:
[0137] A periodic update unit, configured to periodically update the above-mentioned configuration information cached locally to adapt to the real-time changes of the above-mentioned rare Chinese character dictionary item, the above-mentioned rare Chinese character field, and the interface for rare Chinese character processing.
[0138] The above-mentioned rare Chinese character processing device includes a processor and a memory. The above-mentioned acquisition unit, judgment unit, conversion unit, determination unit, etc. are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above-mentioned program units stored in the memory. The above-mentioned modules are all located in the same processor; or, the above-mentioned each module is located in different processors in any combined form.
[0139] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0140] An embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned rare Chinese character processing method.
[0141] Specifically, the rare Chinese character processing method includes:
[0142] Step S201, obtain the configuration information of the rare Chinese character and store the above-mentioned configuration information in the local cache, where the above-mentioned configuration information includes the rare Chinese character dictionary item, the rare Chinese character field, and the interface participating in the rare Chinese character processing;
[0143] Step S202, when receiving a data processing request, judge whether the field in the above-mentioned data processing request is a target field that needs to perform rare Chinese character processing according to the above-mentioned configuration information;
[0144] Step S203: For the above target field, search for an entry in the above rarely-used Chinese character dictionary entries cached locally that matches the Chinese character encoding in the above target field, and convert the above Chinese character encoding into an encoding in a preset format.
[0145] Step S204: In the case where there are multiple mapping entries for the above Chinese character encoding, determine whether the encoding in the preset format is consistent with the pre-stored encoding of the corresponding entry in the above rarely-used Chinese character dictionary entry. If the encoding in the preset format is consistent with the pre-stored encoding, then determine the rarely-used Chinese character corresponding to the encoding in the preset format and the rarely-used Chinese character corresponding to the pre-stored encoding as the same rarely-used Chinese character.
[0146] An embodiment of the present invention provides an electronic device, including a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, at least the following steps are implemented:
[0147] Step S201: Obtain the configuration information of rarely-used Chinese characters, and store the above configuration information in the local cache, where the above configuration information includes rarely-used Chinese character dictionary entries, rarely-used Chinese character fields, and interfaces participating in rarely-used Chinese character processing.
[0148] Step S202: In the case of receiving a data processing request, according to the above configuration information, determine whether the field in the above data processing request is a target field that needs to perform rarely-used Chinese character processing.
[0149] Step S203: For the above target field, search for an entry in the above rarely-used Chinese character dictionary entries cached locally that matches the Chinese character encoding in the above target field, and convert the above Chinese character encoding into an encoding in a preset format.
[0150] Step S204: In the case where there are multiple mapping entries for the above Chinese character encoding, determine whether the encoding in the preset format is consistent with the pre-stored encoding of the corresponding entry in the above rarely-used Chinese character dictionary entry. If the encoding in the preset format is consistent with the pre-stored encoding, then determine the rarely-used Chinese character corresponding to the encoding in the preset format and the rarely-used Chinese character corresponding to the pre-stored encoding as the same rarely-used Chinese character.
[0151] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps:
[0152] Step S201: Obtain the configuration information of rarely-used Chinese characters, and store the above configuration information in the local cache, where the above configuration information includes rarely-used Chinese character dictionary entries, rarely-used Chinese character fields, and interfaces participating in rarely-used Chinese character processing.
[0153] Step S202: In the case of receiving a data processing request, according to the above configuration information, determine whether the field in the above data processing request is a target field that needs to perform rarely-used Chinese character processing.
[0154] Step S203: For the above target field, search for an entry in the above rare Chinese character dictionary entries cached locally that matches the Chinese character encoding in the above target field, and convert the above Chinese character encoding into an encoding in a preset format.
[0155] Step S204: In the case where there are multiple mapping entries for the above Chinese character encoding, determine whether the encoding in the preset format is consistent with the pre-stored encoding of the corresponding entry in the above rare Chinese character dictionary entry. If the encoding in the preset format is consistent with the pre-stored encoding, then determine the rare Chinese character corresponding to the encoding in the preset format and the rare Chinese character corresponding to the pre-stored encoding as the same rare Chinese character.
[0156] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program code executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0157] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0158] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0159] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements in the process Figure 1 one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.
[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0161] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0162] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0163] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0164] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0165] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for processing rare characters, characterized in that: include: Acquire configuration information of rare characters, and store the configuration information in a local cache, wherein the configuration information includes rare character dictionary items, rare character fields, and interfaces involved in rare character processing; When receiving a data processing request, determining, based on the configuration information, whether a field in the data processing request is a target field that needs to be processed for rare characters; For the target field, searching the locally cached rare character dictionary for an entry that matches the Chinese character encoding in the target field, and converting the Chinese character encoding into an encoding in a preset format; In the case where there are multiple mapping items for the Chinese character encoding, determine whether the encoding in the preset format is consistent with the pre-stored encoding of the corresponding entry in the rare character dictionary item. If the encoding in the preset format is consistent with the pre-stored encoding, the rare character corresponding to the encoding in the preset format and the rare character corresponding to the pre-stored encoding are determined to be the same rare character.
2. The method according to claim 1, characterized in that Before obtaining the configuration information of the rare characters, the method further includes: For each of the uncommon characters, an uncommon character dictionary item of the uncommon character is created, wherein the uncommon character dictionary item includes the uncommon character and a plurality of encoding information corresponding to the uncommon character.
3. The method according to claim 1, characterized in that In the case of receiving a data processing request, judging whether a field in the data processing request is a target field that needs to be processed for rare characters according to the configuration information includes: extracting a request field identifier from the data processing request; Compare the request field identifier with the rare word field identifier in the local cache; If the request field identifier matches any of the rare-character field identifiers, the field in the data processing request is determined to be the target field.
4. The method according to claim 1, characterized in that: For the target field, searching the locally cached rare character dictionary for an entry that matches the Chinese character encoding in the target field, and converting the Chinese character encoding into an encoding in a preset format, including: Parsing the Chinese character code in the target field to determine the encoding format of the Chinese character code; Searching the local cache for the rare character dictionary entry that matches the encoding format of the Chinese character encoding; According to the conversion rules in the rare character dictionary item, the Chinese character code is converted into the code in the preset format.
5. The method according to claim 1, characterized in that After converting the Chinese character code into a code in a preset format, the method further includes: Update the mapping relationship between the converted encoding in the preset format and the corresponding uncommon characters into the local cache; After the local cache is updated, cache synchronization is performed with a remote database or server.
6. The method according to claim 1, characterized in that In the case where there are multiple mapping items for the Chinese character code, judging whether the code in the preset format is consistent with the pre-stored code of the corresponding entry in the rare character dictionary item, and if the code in the preset format is consistent with the pre-stored code, determining the rare character corresponding to the code in the preset format and the rare character corresponding to the pre-stored code as the same rare character, including: Compare the code in the preset format with the pre-stored code in each of the mapping items one by one; If the code in the preset format matches any of the pre-stored codes in the rare-character dictionary item, it is determined that the rare character corresponding to the code in the preset format and the rare character corresponding to the matching pre-stored code are the same rare character.
7. The method according to claim 1, characterized in that The method further comprises: The configuration information of the local cache is updated periodically to adapt to the real-time changes of the rare-character dictionary items, the rare-character fields and the rare-character processing interface.
8. A device for processing rare characters, characterized in that: include: An acquisition unit, used for acquiring configuration information of rare characters and storing the configuration information in a local cache, wherein the configuration information includes rare character dictionary items, rare character fields and interfaces involved in rare character processing; A judging unit, configured to, upon receiving a data processing request, judge whether a field in the data processing request is a target field that needs to be processed for rare characters according to the configuration information; A conversion unit, configured to search for an entry matching the Chinese character encoding in the target field from the rare character dictionary items in the local cache for the target field, and convert the Chinese character encoding into an encoding in a preset format; A determination unit is used to determine whether the encoding in the preset format is consistent with the pre-stored encoding of the corresponding entry in the rare character dictionary item when there are multiple mapping items in the Chinese character encoding. If the encoding in the preset format is consistent with the pre-stored encoding, the rare character corresponding to the encoding in the preset format and the rare character corresponding to the pre-stored encoding are determined to be the same rare character.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the rare-character processing method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing the rare character processing method described in any one of claims 1 to 7.