Typo Recognition Method and Device
By determining the number of elements based on the matching degree between image features and element prototypes in handwritten Chinese characters, and verifying the prediction results in combination with the number of elements independent of language information, the problem of low typo recognition accuracy caused by model bias in the prior art is solved, and higher recognition accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510388523.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing technology has model bias in the recognition of handwritten Chinese characters, resulting in low typo recognition accuracy, mainly due to the over-reliance of the model or the language information learned during the training process.
By determining the number of various element prototypes contained in the target text based on the matching degree between the image characteristics of the target text and the element prototypes, the number of various element prototypes contained in the target text is determined, and combined with the number of various element prototypes obtained independently of the language information, the element prototypes obtained based on image characteristics prediction are verified and corrected to ensure that the final element sequence is accurate.
The accuracy of typo recognition results is improved, the recognition error caused by the model's excessive dependence on language information is avoided, and the accuracy and reliability of the recognition results are ensured.
Smart Images

Figure CN119888763B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a misspelling recognition method and device. Background Art
[0002] A handwritten misspelling is an incorrect glyph that appears during the handwritten process. This kind of error may be manifested as the confusion of strokes, structures, or glyphs, making the originally intended character inconsistent with the finally presented glyph, which may lead to the misunderstanding or confusion of information. Therefore, it is necessary to recognize handwritten misspellings.
[0003] Taking handwritten Chinese characters as an example, currently, the Chinese characters are mainly decomposed into radicals through a model to obtain a radical sequence, and whether the Chinese character is a misspelling is recognized through the radical sequence. However, when the model performs radical decomposition reasoning, it tends to utilize the learned language information and output the radical sequence of the correct character corresponding to the misspelling, and thus it is easy to misjudge the misspelling as a correct character. Summary of the Invention
[0004] The present invention provides a misspelling recognition method and device to solve the defects existing in the prior art.
[0005] The present invention provides a misspelling recognition method, including the following steps:
[0006] Based on the matching degree between the image features of the target text and each element prototype, determine the quantity of various element prototypes included in the target text, where the element prototype refers to the basic unit that constitutes a text;
[0007] Based on the image features and the quantity of various element prototypes, decompose the target text into an element sequence, where the element sequence refers to the sequence of basic units that constitute the target text;
[0008] Based on the element sequence, determine the misspelling recognition result of the target text.
[0009] According to the misspelling recognition method provided by the present invention, the step of determining the quantity of various element prototypes included in the target text based on the matching degree between the image features of the target text and each element prototype includes:
[0010] Based on the matching degree between the image features and each element prototype, generate a heat map corresponding to each element prototype, and each heat map is used to characterize the probability of the existence of the corresponding element prototype in the target text;
[0011] Based on the heat maps corresponding to all element prototypes, obtain the quantity of various element prototypes included in the target text.
[0012] A misspelling recognition method provided by the present invention, obtaining the quantity of various element prototypes included in the target text based on the heat maps corresponding to all element prototypes, includes:
[0013] Group the heat maps corresponding to all element prototypes, and independently perform convolution operations on the heat maps in each group;
[0014] Based on the heat maps after convolution of each group, determine the quantity of various element prototypes included in the target text.
[0015] A misspelling recognition method provided by the present invention, the grouping of the heat maps corresponding to all element prototypes and independently performing convolution operations on the heat maps in each group includes:
[0016] Separate the heat maps corresponding to each element prototype into individual groups, and independently perform convolution operations on the heat maps in each group.
[0017] A misspelling recognition method provided by the present invention, decomposing the target text into an element sequence of the target text based on the image features and the quantity of various element prototypes, includes:
[0018] Based on the image features, predict the initial element sequence of the target text;
[0019] Based on the quantity of various element prototypes, correct the initial element sequence to obtain the element sequence of the target text.
[0020] A misspelling recognition method provided by the present invention, correcting the initial element sequence based on the quantity of various element prototypes to obtain the element sequence of the target text, includes:
[0021] Based on the quantity of various element prototypes, correct the probabilities of each element prototype in the initial element sequence;
[0022] Based on the corrected probabilities of each element prototype, obtain the element sequence of the target text.
[0023] A misspelling recognition method provided by the present invention, the image features of the target text are obtained by encoding the target text image, and the initial element sequence is obtained by decoding the image features.
[0024] The present invention also provides a misspelling recognition device, including the following modules:
[0025] A determination unit, configured to determine the quantity of various element prototypes included in the target text based on the matching degree between the image features of the target text and each element prototype, where the element prototype refers to the basic unit that constitutes a text;
[0026] A decomposition unit, configured to decompose the target text into element sequences based on the image features and the quantities of various element prototypes, where the element sequences refer to the sequences of basic units that constitute the target text;
[0027] An identification unit, configured to determine the misspelling identification result of the target text based on the element sequences.
[0028] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the misspelling identification method described in any one of the above is implemented.
[0029] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the misspelling identification method described in any one of the above is implemented.
[0030] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the misspelling identification method described in any one of the above is implemented.
[0031] For the misspelling identification method and device provided by the present invention, since the quantities of various element prototypes included in the target text are based on the results of the matching degree between the image features and the element prototypes, the quantities of various element prototypes included in the target text are independent of the language information. That is to say, the determination of the quantities of various element prototypes is not interfered by the language information, so that the problem of relatively low misspelling identification accuracy that may be caused by the over-reliance or learning of language information by the model during the training process in the related art can be avoided. By combining the quantities of various element prototypes obtained independently of the language information to verify and correct the element prototypes predicted based on the image features, it is ensured that the finally obtained element sequences are accurate, and thus the misspelling identification result can be accurately obtained based on the element sequences, improving the accuracy of the misspelling identification result. Description of the Drawings
[0032] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 is a schematic diagram of the radical decomposition of the model in the related art.
[0034] Figure 2 is a schematic flowchart of the misspelling identification method provided by the present invention.
[0035] Figure 3 It is a schematic flowchart of another misspelling recognition method provided by the present invention.
[0036] Figure 4 It is a schematic diagram of the working principle of the counter provided by the present invention.
[0037] Figure 5 It is a schematic structural diagram of the misspelling recognition device provided by the present invention.
[0038] Figure 6 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0039] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0040] Handwritten Chinese character correction mainly includes two major stages: "evaluation" and "correction". When a user submits a handwritten Chinese character sample, in the "evaluation" stage, the model needs to determine whether the sample is a misspelling; if it is determined to be a misspelling, then in the "correction" stage, the model needs to predict and give the correct Chinese character corresponding to the misspelling, so as to achieve the purpose of correction.
[0041] Currently, handwritten Chinese character correction models generally adopt a method based on radical modeling. The overall process is as follows: First, the model decomposes the handwritten Chinese character input by the user into a radical sequence. In the "evaluation" stage, based on the decomposed radical sequence, the model will query whether it belongs to the set of radical sequences of correct Chinese characters. If it does not belong, it is determined to be a misspelling. Finally, in combination with the decomposed radical sequence, the model will predict and give the correct Chinese character corresponding to the misspelling.
[0042] Among them, Figure 1 is a schematic diagram of model radical decomposition in the related art, such as Figure 1As shown in the figure, when decomposing the radical sequence, the above model often has "bias", that is, when a wrong character is input, although the image is clear, the model accuracy is not high, and it often outputs the radical sequence of the correct Chinese character corresponding to the wrong character. The reason is that wrong characters are unpredictable and infinite, so the model may not have seen some wrong characters during training, but only the correct Chinese characters corresponding to them. The model will "memorize" the radical sequences of these correct Chinese characters, such as "⿱⿰屰月土" and "⿱大⿱冖牛", that is, the model has learned the language information contained in these radical sequences. However, wrong characters are contrary to language information. Therefore, in the reasoning stage, if a wrong character that has not been seen during training is input, the model will tend to use language information to output the radical sequence of the correct Chinese character corresponding to the wrong character. Figure 1 Taking the wrong character “塑” in the video as an example, the model has seen the correct writing of “塑” during training, so it learned the radical sequence of “⿱⿰屰月土”; in the testing phase, when encountering the wrong character “塑”, the model still outputs “⿱⿰屰月土”, even though it is visually clear that the character has “王” instead of “土” at the bottom.
[0043] In this regard, in order to solve the "bias" problem of the above-mentioned model, the present invention provides a method for identifying wrong characters, which can be applied to scenarios such as intelligent dictation and intelligent marking. In addition, the method is not limited to the identification of wrong characters in Chinese characters, but can also be expanded to the identification of wrong characters in other languages that have a similar structural system (such as radical structure) to Chinese characters. For example, Japanese and Korean have a radical structure similar to that of Chinese characters, that is, the present invention can also be used for the identification of wrong characters in Japanese and Korean. In order to facilitate the understanding of the technical solution of the present invention, the following embodiments are all explained by taking the application to Chinese characters as an example.
[0044] in, Figure 2 It is a flow chart of the method for identifying wrong characters provided by the present invention, such as Figure 2 As shown, the method includes step 210, step 220 and step 230.
[0045] Step 210: Based on the matching degree between the image features of the target text and each element prototype, the number of each element prototype contained in the target text is determined, where the element prototype refers to the basic unit constituting the text.
[0046] Here, the target text can be understood as the handwritten Chinese characters to be recognized. The image features of the target text refer to a series of characteristics or attributes of the target text that are visually recognizable and used to distinguish and recognize the text. The image features of the target text can be extracted from a text image containing the target text, such as encoding the text image to obtain the image features of the target text.
[0047] In addition, an element prototype refers to the basic unit that constitutes a character. This element prototype can be a learnable prototype in a dictionary, which is learned from a large number of character samples through machine learning or deep learning algorithms and can represent common structures or components in characters. For Chinese characters, the element prototype can be understood as a radical prototype.
[0048] The degree of matching between the image features of the target character and each element prototype is used to characterize the similarity between the target character and each element prototype. The higher the degree of matching, the closer the target character is to the corresponding element prototype in terms of visual features, that is, the target character may contain that element prototype.
[0049] Since different element prototypes have different roles in character composition and the target character may contain multiple different element prototypes, the number of various element prototypes contained in the target character can be determined based on the degree of matching between the image features of the target character and each element prototype. For example, if the degree of matching between the image features of the target character and a certain radical prototype is greater than a preset threshold, it can be considered that the target character contains the corresponding radical prototype.
[0050] Here, the target character refers to a single handwritten Chinese character. If misspelled character recognition needs to be performed on multiple handwritten Chinese characters, the image containing multiple handwritten Chinese characters can be segmented, and each sub-image obtained only includes one handwritten Chinese character, that is, the handwritten Chinese character in each sub-image can be regarded as the handwritten Chinese character in step 210 above.
[0051] Since the number of various element prototypes contained in the target character is a statistical result based on the matching degree between the image features and the element prototypes, the number of various element prototypes contained in the target character is independent of the specific character content. Furthermore, the number of various element prototypes contained in the target character is independent of language information, that is to say, the determination of the number of various element prototypes is not interfered by language information, thus avoiding the problem of misspelled character recognition accuracy that may be caused by the model relying too much on or learning language information during the training process in related technologies. That is to say, step 210 directly determines the number of elements based on the image features and the element prototype matching degree, which can more objectively and accurately reflect the actual element prototype composition of the target character and improve the accuracy and reliability of misspelled character recognition.
[0052] Step 220: Decompose the target character based on the image features and the number of various element prototypes to obtain an element sequence of the target character. The element sequence refers to the sequence of basic units that constitute the target character.
[0053] Specifically, element decomposition refers to decomposing the target text into multiple basic units, that is, decomposing the target text into multiple element prototypes. Element sequence refers to the sequence of basic units that constitute the target text. For handwritten Chinese characters, element sequence refers to the sequence of radicals that constitute the handwritten Chinese characters. For example, for the character "塑", its corresponding element sequence is "⿱⿰屰月土", "⿱" represents the combination of the upper and lower structures, that is, "屰月" and "土" are combined as the upper and lower structures, and "⿰" represents the combination of the left and right structures, that is, "屰" and "月" are combined as the left and right structures.
[0054] If the target character is a wrong character, then there must be a radical error in the radical prototype that constitutes the target character. Figure 1 As shown, the correct radical of the character "塑" includes "屰月土", but the radical of the wrong character includes "屰月王", that is, in the wrong character, "土" is mistakenly written as "王".
[0055] When decomposing target text elements based on image features, the language information learned by the model is usually used to decompose the target text elements. When the model encounters a wrong word that has not been seen during training, it will introduce the learned language information, that is, the model has a "bias" and is prone to identifying the wrong word as the correct word. Figure 1 When the wrong character "塑" is shown, it is easy to identify the wrong radical "王" as the correct radical "土".
[0056] In order to avoid the "bias" problem caused by the above model, the embodiment of the present invention introduces the number of various element prototypes obtained independently of language information when performing element decomposition. This number is used to represent the number of element prototypes actually contained in the target text. Based on the number of various element prototypes, the element prototypes predicted based on the image features can be corrected, thereby accurately obtaining the element sequence of the target text. Figure 1 When the image features of the wrong character "塑" are used for element decomposition, the element sequence of the target text is predicted to be "⿱⿰屰月土", but the number of element prototypes "土" determined based on step 210 is 0. It can be determined that when the element decomposition is performed based on the image features, the wrong character "塑" is mistakenly identified as the correct character, and then its elements are mistakenly decomposed into "⿱⿰屰月土".
[0057] It can be seen that the number of various element prototypes is introduced to solve the "bias" problem that may be introduced by the model during the recognition process, that is, the model may tend to recognize the wrong characters that have not been seen as the correct characters seen during the training process. By combining the number of various element prototypes obtained independently of language information, the element prototypes predicted based on image features can be verified and corrected to ensure that the final element sequence is accurate.
[0058] Step 230: Determine the wrong character recognition result of the target text based on the element sequence.
[0059] Specifically, the element sequence is a sequence of basic units that constitute the target text. This sequence is used to represent the element prototypes contained in the target text and the combination methods of each element prototype (the combination methods include upper-lower structure combination, left-right structure combination, etc.). If any element prototype is incorrect and / or the combination method of the element prototypes is incorrect, it indicates that the target text is written incorrectly, that is, the misspelling recognition result is that the target text is a misspelled word.
[0060] Based on this, the element sequence of the target text can be compared with the element sequences corresponding to each correct text stored in the sequence library. If the element sequence of the target text matches the element sequence corresponding to any correct text in the sequence library, it indicates that the target text is written correctly, that is, the misspelling recognition result is that the target text is a correct word; if the element sequence of the target text does not match the element sequences corresponding to all correct texts in the sequence library, it indicates that the target text is written incorrectly, that is, the misspelling recognition result is that the target text is a misspelled word. Here, the match means that the element prototypes contained in the target text are the same as those contained in the correct text, and the combination method of each element prototype in the target text is the same as the combination method of each element prototype in the correct text.
[0061] In addition, the misspelling recognition result can also be the probability that the target text is a misspelled word or a correct word. Exemplarily, the similarity between the element sequence of the target text and the element sequences corresponding to each correct text stored in the sequence library can be determined, and the misspelling recognition result can be determined based on the maximum similarity. For example, if the maximum similarity is greater than the threshold, it indicates that the target text is probably a correct word. At this time, the value corresponding to the maximum similarity can be used as the probability that the target text is a correct word; if the maximum similarity is less than or equal to the threshold, it indicates that the target text is probably a misspelled word. At this time, the value corresponding to (1 - maximum similarity) can be used as the probability that the target text is a misspelled word.
[0062] In the misspelling recognition method provided by the embodiments of the present invention, since the number of various element prototypes contained in the target text is based on the result of the matching degree between the image feature and the element prototype, the number of various element prototypes contained in the target text is independent of the language information. That is to say, the determination of the number of various element prototypes is not interfered by the language information, thereby avoiding the problem that the misspelling recognition accuracy may be low due to the over-reliance or learning of language information by the model during the training process in the related art. By combining the number of various element prototypes obtained independently of the language information to verify and correct the element prototypes predicted based on the image feature, it is ensured that the finally obtained element sequence is accurate, and then the misspelling recognition result can be accurately obtained based on the element sequence, improving the accuracy of the misspelling recognition result.
[0063] Based on the above embodiments, determining the number of various element prototypes contained in the target text based on the matching degree between the image feature of the target text and each element prototype includes:
[0064] Generate a heat map corresponding to each element prototype based on the matching degree between the image features and each element prototype, and each heat map is used to characterize the probability of the existence of the corresponding element prototype in the target text;
[0065] Based on the heat maps corresponding to all element prototypes, obtain the quantity of various element prototypes included in the target text.
[0066] Specifically, considering that in a text image with a complex background, the target text may be confused with background elements, resulting in a large amount of noise in the extracted image features, or the illumination change may cause some regions of the target text to be too dark or too bright, affecting the quality of image feature extraction. If directly identifying the element prototypes included in the target text in the text image based on the image features containing a large amount of noise or the image features with poor quality, it may interfere with the recognition accuracy, and further affect the quantity of various element prototypes included in the recognized target text.
[0067] Furthermore, considering that the heat map provides rich spatial distribution information by showing the probability of the existence of a specific element prototype in each pixel or image region, and this spatial distribution information helps to distinguish the target text from background elements and reduce the interference of background elements on the recognition of element prototypes. In addition, when there is an illumination change in the text image, some regions of the target text may become too bright or too dark, resulting in difficulty in image feature extraction. However, the heat map can reduce the impact of illumination change on recognition to a certain extent by calculating the feature matching degree of the local region. Even if the illumination conditions in some regions are not good, the heat map can still infer the feature information of the illumination change region based on the features of other regions.
[0068] Based on this, when determining the quantity of various element prototypes included in the target text in the embodiments of the present invention, first generate a heat map corresponding to each element prototype based on the matching degree between the image features and each element prototype. This heat map can characterize the probability of the existence of the corresponding element prototype in the target text. For example, a high-brightness (or warm-color tone) region indicates that there is a high probability of the existence of the corresponding element prototype in this region, while a low-brightness (or cold-color tone) region indicates that there is a low probability of the existence of the corresponding element prototype in this region. Exemplarily, the matching degree between the two can be determined based on the image features and the feature vectors corresponding to each element prototype, and a heat map of the corresponding element prototype is generated based on the matching degree between the two. Among them, the value range of the heat map is [0, 1], and the larger the value, the higher the probability of the existence of the corresponding element prototype in the target text.
[0069] Since the heatmaps of each element prototype are used to characterize the probability of the corresponding element prototype existing in the target text, it is then possible to determine whether the corresponding element prototype exists in the target text based on the heatmaps of each element prototype. For example, a corresponding threshold can be set. When the value of the heatmap is greater than the threshold, it is considered that the corresponding element prototype exists in the target text; otherwise, the corresponding element prototype does not exist.
[0070] After identifying whether each element prototype exists in the target text, count the number of various element prototypes that exist to obtain the number of various element prototypes included in the target text.
[0071] Based on any of the above embodiments, based on the heatmaps corresponding to all element prototypes, obtain the number of various element prototypes included in the target text, including:
[0072] Group the heatmaps corresponding to all element prototypes, and independently perform convolution operations on the heatmaps in each group;
[0073] Based on the heatmaps after convolution of each group, determine the number of various element prototypes included in the target text.
[0074] Specifically, considering that the heatmap generation algorithm may produce inaccurate results when dealing with complex data or boundary conditions, and these results are manifested as noise in the heatmap, and these noises may affect the subsequent determination of the number of element prototypes. Therefore, the embodiments of the present invention introduce a convolution operation to smooth the heatmap, remove these noises, enhance the spatial continuity of the heatmap, and make the representation of the heatmap more accurate and reliable.
[0075] In addition, considering that if all heatmaps are convolved in one channel, the convolution kernel will act on the heatmaps of all element prototypes at the same time, extracting the common or unique language features of each of them. Since the convolution kernel is reused throughout the input data, it can learn the common features between different element prototypes. These features may reflect a certain structure or rule in the language, and then the convolution kernel can learn the language information between different element prototypes. In this way, there will be the same problem of the above-mentioned model "bias".
[0076] Based on this, the embodiments of the present invention group the heatmaps corresponding to all element prototypes and independently perform convolution operations on the heatmaps in each group, so that the heatmaps in each group remain independent during the convolution process. That is, when the heatmaps in a certain group are convolved, they will not learn the language information of the element prototypes in other groups, ensuring that the feature extraction of the heatmaps in each group is independent of the language information of other groups. Furthermore, based on the heatmaps after convolution of each group, it is possible to more accurately analyze and determine the number of various element prototypes included in the target text, avoiding the "bias" caused by the model learning the common features between different element prototypes, and realizing the accurate statistics of the number of element prototypes in the target text.
[0077] For example, when grouping the heat maps corresponding to all element prototypes, the heat maps corresponding to element prototypes with similar or related semantic categories can be grouped together, so that when grouping the same group of heat maps, the mutual interference of feature information in different groups can be avoided. For example, "木" is usually related to trees and plants, and "氵" is usually related to water and liquids. Since "木" and "氵" have a certain correlation in semantic categories, the heat maps corresponding to "木" and "氵" can be divided into the same group.
[0078] Based on any of the above embodiments, the heat maps corresponding to all element prototypes are grouped, and a convolution operation is performed independently on the heat maps in each group, including:
[0079] The heat maps corresponding to each element prototype are grouped separately, and the convolution operation is performed independently on the heat maps in each group.
[0080] Specifically, the heatmaps corresponding to each element prototype are grouped separately, which means creating a separate channel for each heatmap and performing convolution operations on the corresponding heatmaps in the separate channels. In this way, the heatmaps in each channel will not affect each other, and the convolution kernel will not learn the language information in different element prototypes, ensuring that the feature extraction of each group of heatmaps is independent of the feature information of other groups, avoiding the "bias" caused by learning the common features between different element prototypes, and ensuring that the convolution operation can focus on the unique features of each element prototype, so as to more accurately count the number of various element prototypes contained in the target text.
[0081] Based on any of the above embodiments, the target text is decomposed into elements based on the image features and the number of prototypes of various elements to obtain an element sequence of the target text, including:
[0082] Based on the image features, the initial element sequence of the target text is predicted;
[0083] Based on the number of prototypes of each type of element, the initial element sequence is modified to obtain the element sequence of the target text.
[0084] Specifically, when the target text is decomposed into elements based on image features, the language information learned by the model is usually used to decompose the target text into elements. When the model encounters a wrong character that has not been seen during the training process, it will introduce the learned language information, that is, the model has a "bias" and is prone to identifying the wrong character as the correct character. In other words, the initial element sequence predicted based on image features may contain incorrect element prototype recognition results.
[0085] To avoid the "bias" problem brought about by the above model, in the embodiment of the present invention, when performing element decomposition, the number of various element prototypes independent of language information acquisition is introduced. This number is used to represent the number of element prototypes actually included in the target text. Based on the number of various element prototypes, the initial element sequence predicted based on the image features can be corrected, and then the element sequence of the target text can be accurately obtained.
[0086] Exemplarily, the initial element sequence may include the prediction probabilities of the existence of various element prototypes in the target text. If the number of any of the obtained element prototypes is large, it indicates that the probability of the corresponding element prototype existing in the target text is high. At this time, the prediction probability of the corresponding element prototype in the initial element sequence can be increased. Similarly, if the number of any element prototype is small, it indicates that the probability of the corresponding element prototype existing in the target text is low. At this time, the prediction probability of the corresponding element prototype in the initial element sequence can be decreased, thereby realizing the correction of the initial element sequence and obtaining the accurate element sequence of the target text.
[0087] Based on any of the above embodiments, correcting the initial element sequence based on the number of various element prototypes to obtain the element sequence of the target text includes:
[0088] Based on the number of various element prototypes, correcting the probabilities of the various element prototypes in the initial element sequence;
[0089] Based on the corrected probabilities of the various element prototypes, obtaining the element sequence of the target text.
[0090] Specifically, the probabilities of the various element prototypes in the initial element sequence are used to reflect the possibility of the target text containing the various element prototypes. Since the initial element sequence is obtained by predicting based on image features, and language information learned by the model during the training phase is incorporated in this prediction process, this may lead to misjudging misspelled words as correct words. In other words, the initial element sequence generated by predicting based on image features may contain incorrect element prototype recognition results, that is, there may be errors in the probability values of the various element prototypes in the initial element sequence.
[0091] To address this problem, the embodiment of the present invention introduces a number of various element prototypes independent of the language information acquisition process. This number is used to represent the number of element prototypes actually included in the target text. By using this number, the probability values of the various element prototypes in the initial element sequence can be corrected. Based on the corrected probability values of the various element prototypes, the element sequence of the target text can be accurately determined.
[0092] Specifically, if the number of a certain element prototype is large, it means that the probability of the appearance of this element prototype in the target text is high. At this time, the prediction probability of the corresponding element prototype in the initial element sequence can be appropriately increased. On the contrary, if the number of a certain element prototype is small, it indicates that the probability of the appearance of this element prototype in the target text is low. At this time, the prediction probability of the corresponding element prototype in the initial element sequence can be appropriately decreased. In this way, the initial element sequence can be corrected to obtain an accurate target text element sequence.
[0093] Based on any of the above embodiments, the image features of the target text are obtained by encoding the target text image, and the initial element sequence is obtained by decoding the image features.
[0094] Specifically, the target text image refers to an image containing the target text. Encoding refers to the process of converting the target text image into a format or data structure that can represent the key information of the image, and this format or data structure can be understood as "image features". The purpose of encoding is to remove redundant information from the target text image and extract features useful for misspelling recognition for subsequent processing. Decoding refers to the process of converting the encoded image features into an initial element sequence.
[0095] Exemplarily, the target text image can be input into an encoder, and the encoder encodes the target text image to obtain the image features of the target text. The image features of the target text are input into a decoder, and the decoder decodes the image features to obtain an initial element sequence. In addition, the decoder can also correct the initial element sequence based on the number of various element prototypes to obtain the element sequence of the target text.
[0096] Based on any of the above embodiments, Figure 3 is a schematic flowchart of another misspelling recognition method provided by the present invention. As Figure 3 shown, this method performs misspelling recognition on the target text through a misspelling recognition model. The misspelling recognition model includes an encoder, a counter, and a decoder. Among them, Figure 4 is a schematic diagram of the working principle of the counter provided by the present invention. There is a dictionary inside the counter, and different radical prototypes (Learnable Prototype) are stored in this dictionary.
[0097] Among them, the encoder is used to encode the input target text image to obtain the image features F of the target text, and input the image features F of the target text into the counter. The counter first performs a linear transformation (Linear) on the image features F, then performs a dot product (Dot Product) with the radical prototype (Learnable Prototype), and after passing through an activation function (such as the Sigmoid activation function), obtains the heat map of the corresponding radical prototype. As Figure 4As shown, for the radicals existing in the image, such as "female" and "strength", the maximum value of the heatmap is nearly 1, while for the non-existing radical "again", the heatmap is almost all close to 0. Therefore, by taking the global maximum value (Global Max Pooling) of the heatmap, a probability vector can be obtained. Each value of this probability vector represents the probability of the corresponding radical existing in the target text. For example, Figure 4 in Figure 4 , 0.9 in the probability vector indicates that the probability of "female" existing in the target text is 0.9.
[0098] To obtain the quantity of each radical, perform a group convolution on the heatmap, that is, use a separate channel for each heatmap to perform group convolution, avoiding the mutual influence of different heatmaps, which may cause the counter to learn the language knowledge in the radical sequence. Finally, use global average pooling (Global Average Pooling) to obtain a count vector. Among them, each value in the count vector represents the quantity of each radical in the target text. For example, Figure 4 in Figure 4 , "0.0" in the count vector indicates that there are 0 "again" in the target text, that is, "again" does not exist in the target text.
[0099] After that, the count vector is linearly transformed (Linear) and then input to the decoder. The decoder gradually obtains the radicals at each position based on the image feature F, obtains the radical sequence of the target text, and finally determines the misspelling recognition result of the target text based on the radical sequence of the target text. The radical at each specific position is obtained based on the following steps: Obtain the predicted probability P of the radical at the current position based on the image feature F, multiply the predicted probability P element-wise with Tanh(v) (v is the count vector) to obtain the pseudo-probability Q corresponding to each radical, and take the radical corresponding to the maximum pseudo-probability as the radical at the current position. That is, for the radical with a low predicted existence probability by the counter, the decoder also reduces its predicted probability; on the contrary, for the radical with a high predicted existence probability by the counter, the decoder also increases its predicted probability.
[0100] Among them, the counter is different from the decoder. The counter does not output step by step, but independently outputs the quantities of all radicals, thus avoiding the interference of language information on the counter, and then being able to assist the decoder to accurately decompose the radicals of the target text and improve the misspelling recognition accuracy of the target text.
[0101] The misspelling recognition device provided by the present invention will be described below. The misspelling recognition device described below can be mutually corresponded and referred to with the misspelling recognition method described above.
[0102] Based on any of the above embodiments, Figure 5 is a schematic structural diagram of the misspelling recognition device provided by the present invention. As Figure 5 shown, the device includes:
[0103] A determination unit 510, configured to determine the quantity of various element prototypes included in the target text based on the matching degree between the image features of the target text and each element prototype, where the element prototype refers to the basic unit that constitutes the text.
[0104] A decomposition unit 520, configured to decompose the target text into an element sequence based on the image features and the quantity of various element prototypes, where the element sequence refers to the basic unit sequence that constitutes the target text.
[0105] An identification unit 530, configured to determine the misspelling identification result of the target text based on the element sequence.
[0106] Based on any of the above embodiments, determining the quantity of various element prototypes included in the target text based on the matching degree between the image features of the target text and each element prototype includes:
[0107] Generating a heat map corresponding to each element prototype based on the matching degree between the image features and each element prototype, where each heat map is used to represent the probability of the corresponding element prototype existing in the target text.
[0108] Obtaining the quantity of various element prototypes included in the target text based on the heat maps corresponding to all element prototypes.
[0109] Based on any of the above embodiments, obtaining the quantity of various element prototypes included in the target text based on the heat maps corresponding to all element prototypes includes:
[0110] Grouping the heat maps corresponding to all element prototypes, and independently performing a convolution operation on the heat maps in each group.
[0111] Determining the quantity of various element prototypes included in the target text based on the heat maps after convolution of each group.
[0112] Based on any of the above embodiments, grouping the heat maps corresponding to all element prototypes, and independently performing a convolution operation on the heat maps in each group includes:
[0113] Grouping the heat maps corresponding to each element prototype separately, and independently performing a convolution operation on the heat maps in each group.
[0114] Based on any of the above embodiments, decomposing the target text into an element sequence based on the image features and the quantity of various element prototypes includes:
[0115] Predicting an initial element sequence of the target text based on the image features.
[0116] Correcting the initial element sequence based on the quantity of various element prototypes to obtain the element sequence of the target text.
[0117] Based on any of the above embodiments, the initial element sequence is corrected based on the quantities of various element prototypes to obtain the element sequence of the target text, including:
[0118] Based on the quantities of various element prototypes, correct the probabilities of the element prototypes in the initial element sequence;
[0119] Based on the corrected probabilities of the element prototypes, obtain the element sequence of the target text.
[0120] Based on any of the above embodiments, the image features of the target text are obtained by encoding the target text image, and the initial element sequence is obtained by decoding the image features.
[0121] Figure 6 is a schematic structural diagram of an electronic device provided by the present invention. As Figure 6 shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute the misspelling recognition method, which includes: determining the quantities of various element prototypes included in the target text based on the matching degrees between the image features of the target text and the element prototypes, where the element prototype refers to the basic unit that constitutes the text; decomposing the target text into elements based on the image features and the quantities of various element prototypes to obtain the element sequence of the target text, where the element sequence refers to the basic unit sequence that constitutes the target text; and determining the misspelling recognition result of the target text based on the element sequence.
[0122] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0123] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the typo recognition method provided by each of the above methods. The method includes: determining the quantity of various element prototypes included in the target text based on the matching degree between the image features of the target text and each element prototype, where the element prototype refers to the basic unit that constitutes the text; decomposing the target text into elements based on the image features and the quantity of various element prototypes to obtain an element sequence of the target text, where the element sequence refers to a sequence of basic units that constitute the target text; and determining the typo recognition result of the target text based on the element sequence.
[0124] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the typo recognition method provided by each of the above methods. The method includes: determining the quantity of various element prototypes included in the target text based on the matching degree between the image features of the target text and each element prototype, where the element prototype refers to the basic unit that constitutes the text; decomposing the target text into elements based on the image features and the quantity of various element prototypes to obtain an element sequence of the target text, where the element sequence refers to a sequence of basic units that constitute the target text; and determining the typo recognition result of the target text based on the element sequence.
[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for identifying wrong characters, characterized in that: include: Based on the matching degree between the image features of the target text and each element prototype, the number of each element prototype contained in the target text is determined, where the element prototype refers to the basic unit constituting the text; Based on the image features and the number of prototypes of various elements, the target text is decomposed into elements to obtain an element sequence of the target text, where the element sequence refers to a basic unit sequence constituting the target text; Based on the element sequence, determining a wrong character recognition result of the target text; The step of performing element decomposition on the target text based on the image features and the number of each type of element prototypes to obtain an element sequence of the target text includes: Based on the image features, predicting an initial element sequence of the target text; Based on the number of each type of element prototypes, the initial element sequence is modified to obtain the element sequence of the target text.
2. The method for identifying wrong characters according to claim 1, characterized in that: The method of determining the number of each type of element prototypes contained in the target text based on the matching degree between the image features of the target text and each element prototype includes: Based on the matching degree between the image features and each element prototype, a heat map corresponding to each element prototype is generated, and each heat map is used to represent the probability of the corresponding element prototype existing in the target text; Based on the heat maps corresponding to all element prototypes, the number of each type of element prototype contained in the target text is obtained.
3. The method for identifying wrong characters according to claim 2, characterized in that: The heat map corresponding to all element prototypes is used to obtain the number of each type of element prototype contained in the target text, including: Group the heatmaps corresponding to all element prototypes and perform convolution operations on the heatmaps in each group independently; Based on each group of heat maps after convolution, the number of each type of element prototypes contained in the target text is determined.
4. The method for identifying wrong characters according to claim 3, characterized in that: The heat maps corresponding to all element prototypes are grouped, and the heat maps in each group are independently convolved, including: The heat maps corresponding to each element prototype are grouped separately, and the convolution operation is performed independently on the heat maps in each group.
5. The method for identifying wrong characters according to claim 1, characterized in that: The step of correcting the initial element sequence based on the number of each type of element prototypes to obtain the element sequence of the target text includes: Based on the number of each element prototype, the probability of each element prototype in the initial element sequence is modified; Based on the corrected probabilities of the prototypes of each element, the element sequence of the target text is obtained.
6. The method for identifying wrong characters according to claim 1, characterized in that: The image feature of the target text is obtained by encoding the target text image, and the initial element sequence is obtained by decoding the image feature.
7. A wrong character recognition device, characterized in that: include: A determination unit, configured to determine the number of each type of element prototypes contained in the target text based on the matching degree between the image features of the target text and each element prototype, wherein the element prototype refers to a basic unit constituting the text; A decomposition unit, configured to decompose the target text into elements based on the image features and the number of each type of element prototypes to obtain an element sequence of the target text, wherein the element sequence refers to a basic unit sequence constituting the target text; A recognition unit, used for determining a wrong character recognition result of the target text based on the element sequence; The step of performing element decomposition on the target text based on the image features and the number of each type of element prototypes to obtain an element sequence of the target text includes: Based on the image features, predicting an initial element sequence of the target text; Based on the number of each type of element prototypes, the initial element sequence is modified to obtain the element sequence of the target text.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for identifying wrong characters according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying wrong characters according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for identifying wrong characters according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Method and device for correcting wrongly written characters, electronic equipment and storage medium
CN112668312A
Method and system for detecting and correcting Chinese characters and computing equipment
CN114387603A