A standard processing method and system based on rare characters

By preprocessing the input text data and standardizing dictionary table marks, and combining multimodal features to identify uncommon words, the problem of insufficient recognition accuracy and efficiency of uncommon words in the existing technology is solved, and the accurate recognition and extraction of uncommon words is achieved.

CN119416742BActive Publication Date: 2025-05-16BANK OF SHANGHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510025748.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-16
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The prior art is difficult to dynamically adapt to the changes of rare words, and there are many limitations in the recognition accuracy, processing efficiency and versatility of rare words.

Method used

By obtaining input text data for preprocessing, a standardized dictionary table is established for marking suspected uncommon characters, and comprehensively identifying uncommon characters based on multimodal character features, transforming identified uncommon characters and forming a list output of uncommon characters.

Benefits of technology

It improves the accuracy and flexibility of recognition of rare words, realizes the accurate recognition and extraction of rare words, reduces the recognition processing volume, and improves the recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416742B_ABST
    Figure CN119416742B_ABST
Patent Text Reader

Abstract

The present invention discloses a standardized processing method and system based on rare characters, which relates to the field of character recognition processing technology, including obtaining input text data for preprocessing and unifying the text data format, establishing a standardized dictionary table for suspected rare characters marking; extracting multimodal text features based on suspected rare characters marking to comprehensively identify rare characters, converting recognized rare characters, and outputting unrecognized rare characters in a list; displaying rare character recognition results and storing the recognition results. The present invention obtains user input text data for preprocessing and marks suspected rare characters, thereby reducing the rare character recognition processing volume and improving recognition efficiency. At the same time, by extracting rare characters multimodal feature vectors for comprehensive recognition of rare characters, the accuracy and flexibility of rare character recognition are greatly improved, and accurate recognition and extraction of rare characters are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of character recognition processing, and in particular to a standardization processing method and system based on rare characters. Background Art

[0002] In recent years, with the rapid development of information technology and the widespread popularity of word processing applications, word processing technology has become one of the important research directions in the field of natural language processing (NLP). However, in Chinese text processing, the recognition and processing of rare characters has always been a technical bottleneck that is difficult to break through. Due to their low frequency of occurrence, complex character shapes, and ambiguous semantics, rare characters are prone to cause a series of problems during text input, transmission, storage, and display, such as character encoding errors, information loss, and text content parsing failure. In the existing technology, some methods attempt to address the problem of rare character processing by building a character library, using optical character recognition (OCR) technology, or matching based on semantic context. However, these methods usually rely on fixed character libraries or models, and it is difficult to dynamically adapt to the changes of rare characters. There are also many limitations in the recognition accuracy, processing efficiency, and versatility of rare characters. Summary of the invention

[0003] In view of the above existing problems, the present invention is proposed.

[0004] Therefore, the present invention provides a standardized processing method and system based on rare characters, which solves the problem that the existing methods are difficult to dynamically adapt to the changes of rare characters and have many limitations in the recognition accuracy, processing efficiency and versatility of rare characters.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] In a first aspect, the present invention provides a standardization processing method based on rare characters, which comprises:

[0007] Obtain input text data for preprocessing and unify the text data format, and establish a standardized dictionary table to mark suspected rare characters;

[0008] Extract multimodal text features based on suspected rare characters, comprehensively identify rare characters, transform the recognized rare characters, and output the unrecognized rare characters in a list;

[0009] Display the recognition results of rare characters and store them.

[0010] As a preferred solution of the standardized processing method based on uncommon characters described in the present invention, wherein: the acquisition of input text data for preprocessing and unifying the text data format refers to acquiring text data containing uncommon characters input by a user, converting the text data into a standard character format, using a string concatenation method to form continuous text content from the text data, and using regular expressions to remove redundant content in the continuous text content, using a word segmentation algorithm based on bidirectional maximum matching to perform word segmentation on the continuous text content to extract characters and output the word segmentation results.

[0011] As a preferred solution of the standardized processing method based on uncommon characters described in the present invention, wherein: the establishment of a standardized dictionary table to mark suspected uncommon characters refers to establishing a standardized dictionary table by querying national language and character norms and commonly used Chinese character standards, extracting character images and character pinyin for each character in the word segmentation result, matching them one by one in the standardized dictionary table, marking successfully matched characters as ordinary characters, and simultaneously marking unmatched characters as suspected uncommon characters and adding them to the suspected uncommon character set.

[0012] As a preferred solution of the standardization processing method based on rare characters described in the present invention, wherein: the extraction of multimodal text features based on suspected rare characters marks to comprehensively identify rare characters refers to obtaining a set of suspected rare characters, decomposing each rare character in the set of suspected rare characters into specific components using a parser based on configuration rules, and defining the specific components as leaf nodes to construct a binary tree structure, extracting the leaf nodes in the binary tree structure :

[0013] ;in is the pinyin of the component, is the component structure information;

[0014] The morphological feature extraction model is constructed and trained through convolutional neural network. Input the trained convolutional neural network to obtain the morphological features of rare characters ;

[0015] Define the context window size of rare characters as , extract context characters from continuous text content Forming context character set ;

[0016] ;

[0017] in Indicates the first characters;

[0018] Use the word vector model to convert context characters into word vectors , using the character vector model to convert the context characters into character vectors , and simultaneously use the pinyin generation tool to convert the context characters into pinyin to form a pinyin embedding vector , respectively for the word vectors in the context character set , character vector And the pinyin embedding vector The corresponding context features are generated by averaging: ;

[0019] in Context features, including word-level context features , character-level contextual features and phonetic context features , k is the sequence number;

[0020] All context features are weighted and fused to generate a comprehensive context feature vector ;

[0021] Generate pinyin embedding vectors for rare characters based on pinyin generation tools , build an associative learning network Encoder and train it to embed the pinyin into the vector Morphological features of rare characters Input the trained associative learning network Encoder to obtain visual and speech associative features ;

[0022] The morphological characteristics , comprehensive context feature vector and visual-phonological association features Perform L2 normalization and weighted fusion to generate a comprehensive embedding vector , use TreeLSTM to integrate the embedding vector Optimize to form the final feature vector ;

[0023] Generate the final feature vector of dictionary characters based on each dictionary character in the standardized dictionary table , respectively calculate the final feature vector of rare characters Final feature vector with each dictionary character If the cosine similarity is higher than the set threshold, the dictionary character with the highest cosine similarity is used as the recognition result of the rare character and the rare character is marked as a recognized character. If the cosine similarity is not higher than the set threshold, the rare character is marked as an unrecognized character.

[0024] As a preferred solution of the standardization processing method based on uncommon characters described in the present invention, wherein: the conversion of recognized uncommon characters and outputting a list of unrecognized uncommon characters refers to directly replacing the recognized uncommon characters with standardized characters according to the uncommon character recognition results, using a unified Unicode encoding representation, adding emphasis marks to unrecognized characters and using a pinyin library to generate pinyin information that is attached to the emphasis marks of unrecognized characters, and outputting the unrecognized characters in a list.

[0025] As a preferred scheme of the standardization processing method based on uncommon characters described in the present invention, wherein: the display of uncommon character recognition results refers to replacing the recognized uncommon characters with standardized characters and then replacing them into continuous text content for display to the user, and highlighting all unrecognized characters in the unrecognized character list in the continuous text content to prompt the user's attention, and exporting the replaced continuous text content according to user needs.

[0026] As a preferred solution of the standardized processing method based on uncommon characters described in the present invention, wherein: the storing of the recognition results refers to adding the uncommon character recognition results into the standardized dictionary table for content update, synchronously generating update records, storing the updated standardized dictionary table and update records in the database, and the database synchronously uploads the stored data to the cloud for synchronous backup storage each time the stored data is updated.

[0027] In a second aspect, the present invention provides a standardization processing system based on rare characters, comprising:

[0028] A text processing module is used to obtain user input text data for preprocessing and establish a standardized dictionary table to mark suspected rare characters;

[0029] A recognition module, for comprehensively identifying rare characters based on the multimodal text features used to extract rare characters and for converting and marking the rare characters;

[0030] The display and storage module is used to replace the uncommon character recognition results and store them in the database and the cloud.

[0031] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the standardization processing method based on rare characters as described in the first aspect of the present invention is implemented.

[0032] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the standardization processing method based on rare characters as described in the first aspect of the present invention is implemented.

[0033] The beneficial effects of the present invention are as follows: the present invention obtains user input text data for preprocessing and marks suspected rare characters, thereby reducing the processing volume of rare character recognition and improving recognition efficiency. At the same time, the present invention extracts multimodal feature vectors of rare characters for comprehensive recognition of rare characters, greatly improving the accuracy and flexibility of rare character recognition, and realizing accurate recognition and extraction of rare characters. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0035] Figure 1 This is a flow chart of the standardization processing method based on rare characters in Example 1.

[0036] Figure 2 This is a structural diagram of the standardization processing system based on rare characters in Example 1. DETAILED DESCRIPTION

[0037] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0038] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0039] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0040] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides a standardization processing method based on rare characters, comprising the following steps:

[0041] S1, obtaining input text data for preprocessing and unifying the text data format, and establishing a standardized dictionary table to mark suspected rare characters;

[0042] Specifically, obtaining input text data for preprocessing and unifying the text data format refers to obtaining text data input by a user that contains uncommon characters, wherein the text data includes structured text (such as JSON, XML), unstructured text (such as TXT), document files (such as Word, PDF) and image files (converted to text through OCR), converting the text data into a standard character format, using a string concatenation method to form continuous text content from the text data, and using regular expressions to remove redundant content (such as blank characters, special symbols) in the continuous text content, using a word segmentation algorithm based on bidirectional maximum matching to perform word segmentation on the continuous text content, extracting characters and outputting a word segmentation result, wherein the word segmentation result includes each character in the continuous text content.

[0043] The standardized character format ensures the consistency of processing of diverse text data and provides high-quality input data for subsequent steps. In the prior art, the diversity of data formats may lead to character recognition failure or require additional adaptation and conversion. The present invention significantly reduces the cost of data cleaning and correction through a standardized preprocessing strategy, improves the overall efficiency of the system, and through efficient cleaning of regular expressions, all redundant content can be eliminated so that the text content remains highly compact, thereby ensuring the accuracy of word segmentation. The text after redundant content cleaning is more concise and coherent, reducing the processing complexity of the word segmentation algorithm and improving the performance of subsequent uncommon character recognition and marking. This optimization not only improves the processing speed, but also reduces errors caused by redundant characters. The bidirectional maximum matching algorithm takes into account the bidirectional dependency of the context when segmenting words, thereby reducing the occurrence of erroneous segmentation.

[0044] Furthermore, establishing a standardized dictionary table to mark suspected rare characters means establishing a standardized dictionary table by querying national language and character specifications and commonly used Chinese character standards, dynamically synchronizing with the latest language and character standards issued by the country, and regularly expanding the dictionary table content through administrator review. For each character in the word segmentation result, the character image and character pinyin are extracted and matched one by one in the standardized dictionary table. Matching can be performed by comparing the structure and pinyin similarity, and successfully matched characters are marked as ordinary characters, and unmatched characters are simultaneously marked as suspected rare characters and added to the suspected rare characters set.

[0045] By establishing a standardized dictionary table based on national language and character standards and commonly used Chinese character standards, dynamically synchronizing the latest language and character standards, ensuring that the dictionary table can be consistent with national standards, this dynamic update mechanism avoids the problem that static dictionary tables in the prior art cannot adapt to the development and changes of language and characters. Character images and character pinyins are extracted one by one for each character in the word segmentation result for matching. This is a fine-grained marking method. Compared with the traditional word-level matching method (such as fuzzy recognition based only on dictionary rules), the present invention performs matching operations at the character level, which can accurately locate suspected rare characters and avoid omissions or mis-marking. In the one-by-one matching, the characters that are successfully matched are marked as ordinary characters, and the characters that are not successfully matched are marked as suspected rare characters and added to the suspected rare character set. This design of clearly distinguishing marks not only makes the classification of ordinary characters and rare characters clearer, but also provides a clear goal for the subsequent further processing of rare characters (such as dynamically expanding the dictionary table or manual verification). Through clear marking rules and the creation of a suspected rare character set, the accuracy of marking and the pertinence of processing are ensured, thereby solving the error problem in marking and classification.

[0046] S2, extracting multimodal text features based on suspected rare characters marks to comprehensively identify rare characters, converting the recognized rare characters, and outputting the unrecognized rare characters in a list;

[0047] Specifically, extracting multimodal text features based on suspected rare characters marks to comprehensively identify rare characters refers to obtaining a set of suspected rare characters, decomposing each rare character in the set of suspected rare characters into specific components using a parser based on configuration rules, defining the specific components as leaf nodes to construct a binary tree structure, and extracting the leaf nodes in the binary tree structure. ;

[0048] ;in is the pinyin of the component, is the component structure information,(such as radical,stroke sequence);

[0049] The phonetic and structural information of Chinese characters plays a key role in the recognition of rare characters. In particular, for characters with similar shapes or similar sounds, these characters can be effectively distinguished by combining phonetic and structural information for multimodal feature extraction. The parser based on the configuration rules can decompose rare characters into specific components, and a binary tree structure is constructed to define the components as leaf nodes. The uniqueness of this design lies in its ability to capture the morphological feature hierarchy of rare characters, thereby transforming the problem of processing complex character shapes into the problem of recognizing simple components.

[0050] The morphological feature extraction model is constructed and trained through convolutional neural network. Input the trained convolutional neural network to obtain the morphological features of rare characters ;

[0051] Define the context window size of rare characters as , extract context characters from continuous text content Forming context character set :

[0052] ;

[0053] in Indicates the first characters;

[0054] The context window captures the semantic, phonetic and character-level information of the context before and after the rare characters, which helps to associate and identify the semantics of the rare characters.

[0055] Use the word vector model to convert context characters into word vectors , using the character vector model to convert the context characters into character vectors , and simultaneously use the pinyin generation tool to convert the context characters into pinyin to form a pinyin embedding vector , respectively for the word vectors in the context character set , character vector And the pinyin embedding vector The corresponding context features are generated by averaging:

[0056] ;in Context features, including word-level context features , character-level contextual features and phonetic context features , For serial number:

[0057] All context features are weighted and fused to generate a comprehensive context feature vector ;

[0058] By combining multimodal information (semantics, characters, and pinyin), the association between rare characters and their context can be fully captured. By weighted fusion of multimodal context features, the semantic consistency of rare characters in the context is enhanced, significantly improving the recognition accuracy in complex text scenarios.

[0059] Generate pinyin embedding vectors for rare characters based on pinyin generation tools , build an associative learning network Encoder and train it to embed the pinyin into the vector Morphological features of rare characters Input the trained associative learning network Encoder to obtain visual and speech associative features ;

[0060] In the prior art, most methods process the morphological or phonetic features of rare characters separately, but fail to effectively combine the two. The present invention uses an associative learning network encoder to jointly model the morphological features and phonetic embedding vectors, effectively solving the problem of difficulty in distinguishing characters with similar shapes and similar sounds;

[0061] The morphological characteristics , comprehensive context feature vector and visual-phonological association features Perform L2 normalization and weighted fusion to generate a comprehensive embedding vector , use TreeLSTM to integrate the embedding vector Optimize to form the final feature vector ;

[0062] TreeLSTM can capture hierarchical feature associations in a binary tree structure. By capturing the hierarchical features of rare characters, TreeLSTM optimizes the comprehensive feature representation and improves recognition accuracy.

[0063] Generate the final feature vector of dictionary characters based on each dictionary character in the standardized dictionary table , respectively calculate the final feature vector of rare characters Final feature vector with each dictionary character If the cosine similarity is higher than the set threshold, the dictionary character with the highest cosine similarity is used as the recognition result of the rare character and the rare character is marked as a recognized character. If the cosine similarity is not higher than the set threshold, the rare character is marked as an unrecognized character.

[0064] Furthermore, converting the recognized rare characters and outputting the unrecognized rare characters in a list means directly replacing the recognized rare characters with standardized characters according to the rare character recognition results, using a unified Unicode encoding representation, adding emphasis marks to the unrecognized characters and using a pinyin library to generate pinyin information that is attached to the emphasis marks of the unrecognized characters, and outputting the unrecognized characters in a list.

[0065] By replacing recognized rare characters with standardized characters and using unified Unicode encoding to represent them, the consistency of recognized rare characters in storage, transmission and display is ensured; in current technology, the processing of recognized rare characters is often limited to specific scenarios and fails to achieve standardized processing, which leads to data loss or display anomalies when used across platforms or transmitted across multiple devices. The present invention solves the problem of inconsistent character encoding through a standardized strategy of Unicode encoding, significantly improves the versatility and stability of rare character processing, and designs additional strategies for emphasis marks and pinyin information for unrecognized characters. The emphasis marks are for quickly locating unrecognized characters in subsequent processing, while the pinyin information provides a phonetic basis for manual verification and algorithm optimization. By adding emphasis marks and pinyin information, unrecognized characters can be marked significantly and phonetic information is retained, thereby avoiding the problem of information loss and providing operability for subsequent processing.

[0066] S3, displaying the recognition results of rare characters and storing the recognition results;

[0067] Specifically, displaying the recognition results of rare characters means replacing the recognized rare characters with standardized characters and then replacing them in the continuous text content for display to the user, and highlighting all unrecognized characters in the unrecognized character list in the continuous text content to prompt the user's attention, and exporting the replaced continuous text content according to user needs.

[0068] By replacing the recognized rare characters with standardized characters and bringing them into the original continuous text for display, not only the display problem of rare characters is solved, but also the integrity and semantic coherence of the original text are maintained, allowing users to intuitively view the recognition results in the context, improving the convenience and accuracy of text processing. Highlighting can quickly attract the user's attention by visually highlighting the unrecognized characters, allowing them to quickly locate the problem characters in a large amount of text, solving the problem of difficulty in locating unrecognized characters and providing great convenience for the user's further processing. The continuous text export function after replacement provides users with an efficient subsequent processing tool, significantly improving the efficiency of the text processing process.

[0069] Furthermore, storing the recognition results means adding the uncommon character recognition results into a standardized dictionary table for content update, synchronously generating an update record, storing the updated standardized dictionary table and the update record in a database, and the database synchronously uploading the stored data to the cloud for synchronous backup storage each time the stored data is updated.

[0070] This embodiment also provides a standardization processing system based on rare characters, including:

[0071] A text processing module is used to obtain user input text data for preprocessing and establish a standardized dictionary table to mark suspected rare characters;

[0072] A recognition module, for comprehensively identifying rare characters based on the multimodal text features used to extract rare characters and for converting and marking the rare characters;

[0073] The display and storage module is used to replace the uncommon character recognition results and store them in the database and the cloud.

[0074] This embodiment also provides a computer device, which is suitable for the standardized processing method based on uncommon characters, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the standardized processing method based on uncommon characters proposed in the above embodiment.

[0075] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0076] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the standardized processing method based on rare characters proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0077] In summary, the present invention obtains user input text data for preprocessing and marks suspected rare characters, thereby reducing the amount of rare character recognition processing and improving recognition efficiency. At the same time, it extracts multimodal feature vectors of rare characters for comprehensive recognition of rare characters, greatly improving the accuracy and flexibility of rare character recognition, and achieving accurate recognition and extraction of rare characters.

[0078] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A standardization processing method based on rare characters, characterized by: include, Obtain input text data for preprocessing and unify the text data format, and establish a standardized dictionary table to mark suspected rare characters; Extract multimodal text features based on suspected rare characters, comprehensively identify rare characters, transform the recognized rare characters, and output the unrecognized rare characters in a list; Display the recognition results of rare characters and store them; The multimodal text features include morphological features F s , contextual features F c and visual-phonological association features F p ; The morphological feature F s , comprehensive context feature vector F c and visual-phonological association features F p Perform L2 normalization and weighted fusion to generate a comprehensive embedding vector F, and use TreeLSTM to optimize the comprehensive embedding vector F to form the final feature vector F h ; Generate the final feature vector F of the dictionary character based on each dictionary character in the standardized dictionary table z , calculate the final feature vector F of rare characters respectively h The final feature vector F with each dictionary character z If the cosine similarity is higher than the set threshold, the dictionary character with the highest cosine similarity is used as the recognition result of the rare character and the rare character is marked as a recognized character. If the cosine similarity is not higher than the set threshold, the rare character is marked as an unrecognized character.

2. The standardization processing method based on rare characters as claimed in claim 1, characterized in that: The obtaining of input text data for preprocessing and unifying the text data format refers to obtaining text data input by a user containing uncommon characters, converting the text data into a standard character format, using a string concatenation method to form continuous text content from the text data, and using a regular expression to remove redundant content in the continuous text content, using a word segmentation algorithm based on bidirectional maximum matching to perform word segmentation on the continuous text content to extract characters and output the word segmentation result.

3. The standardization processing method based on rare characters as claimed in claim 2, characterized in that: The establishment of a standardized dictionary table to mark suspected rare characters refers to establishing a standardized dictionary table by querying national language and writing standards and commonly used Chinese character standards, extracting character images and character pinyin for each character in the word segmentation result, matching them one by one in the standardized dictionary table, marking successfully matched characters as ordinary characters, and simultaneously marking unmatched characters as suspected rare characters and adding them to a suspected rare character set.

4. The standardization processing method based on rare characters as claimed in claim 3 is characterized in that: The method of extracting multimodal text features based on suspected rare characters marks to comprehensively identify rare characters refers to obtaining a set of suspected rare characters, decomposing each rare character in the set of suspected rare characters into specific components using a parser based on configuration rules, defining the specific components as leaf nodes to construct a binary tree structure, and extracting leaf nodes B in the binary tree structure: B={v p ,v s }; where v p is the pinyin of the component, v s is the component structure information; The morphological feature extraction model is constructed and trained through the convolutional neural network, and the leaf node B is input into the trained convolutional neural network to obtain the morphological feature F of the uncommon characters. s ; Define the context window size of rare characters as T, and extract context characters t from the continuous text content i+j Form the context character set C(t i ): C(t i )={t i-T ,t i-T+1 ,…,t i+T }; where t i+j Represents the jth character in the context window; Use the word vector model to convert the context characters into word vectors w1(t i+j ), using the character vector model to convert the context character into a character vector w2(t i+j ), and simultaneously use the pinyin generation tool to convert the context characters into pinyin to form the pinyin embedding vector w3(t i+j ), respectively for the word vector w1(t i+j ), character vector w2(t i+j ) and the pinyin embedding vector w3(t i+j ) is averaged to generate the corresponding context features: where h k is the context feature, including word-level context feature h1, character-level context feature h2 and pinyin context feature h3, and k is the sequence number; All context features are weighted and fused to generate a comprehensive context feature vector F c ; Generate the pinyin embedding vector w3(t i ), build an associative learning network Encoder and train it, embed the pinyin into the vector w3(t i ) and the morphological features of rare characters F s Input the trained associative learning network Encoder to obtain the visual speech associative feature F p ; The morphological feature F s , comprehensive context feature vector F c and visual-phonological association features F p Perform L2 normalization and weighted fusion to generate a comprehensive embedding vector F, and use TreeLSTM to optimize the comprehensive embedding vector F to form the final feature vector F h ; Generate the final feature vector F of the dictionary character based on each dictionary character in the standardized dictionary table z , calculate the final feature vector F of rare characters respectively h The final feature vector F with each dictionary character z If the cosine similarity is higher than the set threshold, the dictionary character with the highest cosine similarity is used as the recognition result of the rare character and the rare character is marked as a recognized character. If the cosine similarity is not higher than the set threshold, the rare character is marked as an unrecognized character.

5. The standardization processing method based on rare characters as claimed in claim 4, characterized in that: The converting of recognized rare characters and outputting a list of unrecognized rare characters refers to directly replacing the recognized rare characters with standardized characters according to the rare character recognition results, using unified Unicode encoding, adding emphasis marks to unrecognized characters, using a pinyin library to generate pinyin information and appending it to the emphasis marks of unrecognized characters, and outputting the unrecognized characters in a list.

6. The standardization processing method based on rare characters as claimed in claim 5, characterized in that: The display of the rare character recognition results refers to replacing the recognized rare characters with standardized characters and then replacing them in the continuous text content for display to the user, and highlighting all unrecognized characters in the unrecognized character list in the continuous text content to prompt the user's attention, and exporting the replaced continuous text content according to user needs.

7. The standardization processing method based on rare characters according to claim 6, characterized in that: Storing the recognition results refers to adding the uncommon character recognition results to the standardized dictionary table for content update, synchronously generating update records, storing the updated standardized dictionary table and update records in the database, and the database synchronously uploading the stored data to the cloud for synchronous backup storage each time the stored data is updated.

8. A standardization processing system based on rare characters, based on the standardization processing method based on rare characters according to any one of claims 1 to 7, characterized in that: include, A text processing module is used to obtain user input text data for preprocessing and establish a standardized dictionary table to mark suspected rare characters; A recognition module, for comprehensively identifying rare characters based on the multimodal text features used to extract rare characters and for converting and marking the rare characters; The display and storage module is used to replace the uncommon character recognition results and store them in the database and the cloud.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the standardization processing method based on rare characters described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the standardization processing method based on rare characters described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and device for inputting characters

    CN101615084A

  • Method and device for prompting information of uncommon characters

    CN103425257A