Method, device and medium for generating misprint image

By extracting stroke information from standard printed images to generate printed misspelling images and performing style transfer processing, the problem of high cost of acquiring misspelling images is solved, enabling low-cost generation of handwritten misspelling images and improving the accuracy of misspelling recognition technology.

CN116597445BActive Publication Date: 2026-04-07深圳市星桐科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

The high cost of acquiring images of misspelled words in existing technologies leads to a scarcity of such images, which has become a bottleneck in the development of misspelled word recognition technology.

Method used

By acquiring the standard printed image of the target character, extracting the stroke information to be erased, performing the erasure operation to generate a printed misspelling image, and then generating a handwritten misspelling image with a preset handwriting style through style transfer processing.

Benefits of technology

It enables low-cost and rapid automatic generation of handwritten misspelled word images, greatly expanding the number of misspelled word images, reducing acquisition costs, and promoting the development and accuracy of misspelled word recognition technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116597445B_ABST
    Figure CN116597445B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, device, and medium for generating misspelled character images. The method includes: acquiring a standard printed image of a target character; acquiring stroke information of a target stroke to be erased in the standard printed image of the target character; performing an erasure operation based on the stroke information of the target stroke to obtain a printed misspelled character image corresponding to the target character; and performing style transfer processing on the printed misspelled character image to obtain a handwritten misspelled character image with a preset handwriting style. This disclosure can conveniently and quickly automatically generate handwritten misspelled character images, greatly reducing the cost of acquiring misspelled character images.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of artificial intelligence, and in particular, to a method and apparatus for generating a wrong character image, an electronic device, and a computer readable medium. BACKGROUND

[0002] In scenarios such as intelligent correction, a neural network model can automatically identify wrong characters in user homework or test papers, and further correct the wrong characters. However, before this, the model needs to be trained using wrong character images. However, in related technologies, the cost of obtaining wrong character images is high, such as the need to screen wrong characters from a large amount of written content, which requires meticulous and tedious work, and the use of manually written wrong characters is uncontrollable, which also requires a great amount of manual cost. In summary, in related technologies, the high-cost wrong character image acquisition method leads to a lack of wrong character images, and the lack of wrong character images has become a major reason restricting the development of wrong character recognition technology. SUMMARY

[0003] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method and apparatus for generating a wrong character image, an electronic device, and a computer readable medium.

[0004] According to an aspect of the present disclosure, a method for generating a wrong character image is provided, comprising: obtaining a printed standard image of a target single character; obtaining stroke information of a target stroke to be removed in the printed standard image of the target single character; performing a removal operation based on the stroke information of the target stroke to obtain a printed wrong character image corresponding to the target single character; and performing style transfer processing on the printed wrong character image to obtain a handwritten wrong character image having a preset handwriting style.

[0005] According to another aspect of the present disclosure, a device for generating a wrong character image is provided, comprising: an image obtaining module configured to obtain a printed standard image of a target single character; a target stroke obtaining module configured to obtain stroke information of a target stroke to be removed in the printed standard image of the target single character; a stroke removal module configured to perform a removal operation based on the stroke information of the target stroke to obtain a printed wrong character image corresponding to the target single character; and a style transfer module configured to perform style transfer processing on the printed wrong character image to obtain a handwritten wrong character image having a preset handwriting style.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform the method for generating a wrong character image.

[0007] According to another aspect of this disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program for performing the method for generating the misspelled image.

[0008] The technical solution provided in this embodiment can perform an erasure operation based on the stroke information of the target stroke to be erased in the standard printed image of the target character, thereby obtaining a printed misspelling image corresponding to the target character. Furthermore, by performing style transfer processing on the printed misspelling image, a handwritten misspelling image with a preset handwriting style can be obtained. Through this method, handwritten misspelling images can be generated automatically and quickly, greatly reducing the cost of obtaining misspelling images.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0011] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a method for generating a misspelled image according to an embodiment of this disclosure;

[0013] Figure 2 A schematic diagram of a standard printed image provided in an embodiment of this disclosure;

[0014] Figure 3 A schematic diagram of a printed misspelling image provided in an embodiment of this disclosure;

[0015] Figure 4 A schematic diagram of a handwritten misspelling image provided in an embodiment of this disclosure;

[0016] Figure 5 A schematic diagram of the structure of a device for generating misspelled images provided in an embodiment of this disclosure;

[0017] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0020] The term "comprising" and its variations as used in this disclosure are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0022] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0023] Figure 1 This is a flowchart illustrating a method for generating a misspelled image according to an embodiment of the present disclosure. This method can be executed by a misspelled image generating device, which can be implemented in software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, the method mainly includes the following steps S102 to S108:

[0024] Step S102: Obtain the standard printed image of the target character. For ease of understanding, please refer to... Figure 2A schematic diagram of a standard printed image is shown, which shows the standard printed image of the character "技" and is exemplified in bold. In practical applications, printed fonts include various fonts, such as Song typeface, regular script, official script, boldface, etc. Exemplarily, the required font corresponding to the printed font can be selected from the specified national character standard fonts according to requirements. The standard printed image can be a binary image as shown in Figure 2 , which is more convenient for subsequent analysis and processing.

[0025] Step S104: Obtain the stroke information of the target stroke to be erased in the standard printed image of the target single character.

[0026] In practical applications, the target stroke can be determined according to requirements. Specifically, the target stroke can be determined based on the number of target strokes to be erased. In some implementation examples, when the number of target strokes is one, the target stroke is any stroke in the target single character. In practical applications, the target stroke to be erased can be specified according to requirements. In order to fully improve the utilization rate of the standard printed image, each stroke of the target single character can also be used as the target stroke. For example, if the total number of strokes of the target single character is N, then each stroke can be used as the target stroke to be erased in sequence according to the stroke order of the target single character. Therefore, at most N printed misspelled character images can be obtained based on one standard printed image.

[0027] In some other implementation examples, when the total number of strokes of the target single character is not less than a preset threshold and the number of target strokes is greater than one, the multiple target strokes do not form the radical of the target single character. In practical applications, if multiple target strokes need to be erased, the embodiments of the present disclosure ensure the effectiveness of the finally obtained printed misspelled character image by restricting the total number of strokes of the target single character. Further, the embodiments of the present disclosure fully consider that if multiple target strokes are erased simultaneously and the multiple target strokes form the radical of the target single character, the remaining strokes are likely to form a correct single character after the radical is erased. For example, taking the number of target strokes as two strokes, for characters with radicals consisting of two strokes such as "佣" and "冰", after erasing the radical, they are converted into another correct character such as "用" and "水". Therefore, the embodiments of the present disclosure can effectively improve the success probability of the finally obtained misspelled character image by restricting that the multiple target strokes do not form the radical of the target single character. In practical applications, a radical set can be constructed in advance. The radical set can include the stroke order and stroke type of the radical, such as {亻:(1 - left-falling stroke, 2 - vertical stroke), 冫:(1 - dot, 2 - rising stroke)...}, etc. Based on the stroke information of the target stroke and the preset radical set, it is possible to accurately and effectively determine whether the erased stroke belongs to the radical, so as to reliably select multiple target strokes that do not form the radical.

[0028] In some specific implementation examples, the number of target strokes can be flexibly set according to requirements. For example, in order to reduce the impact of erased strokes on the font structure and reduce the occurrence probability of ineffective erasure, the number of target strokes can be set to at most two strokes. In addition, the preset threshold corresponding to the total number of strokes of the target single character can be flexibly set according to requirements. For example, it can be set to five strokes. That is, when the total number of strokes of the target single character is not less than five strokes, two strokes can be selected from the target single character as the target strokes, and the two selected strokes cannot form the radical of the target single character. For a target single character, two target strokes to be erased can be specified according to requirements. In order to fully improve the utilization rate of the standard printed image, the strokes of the target single character can also be combined in pairs, and various combinations of target strokes can be obtained. It can be understood that if the total number of strokes of the target single character is larger, the number of combinations of selecting two strokes as the target strokes will increase exponentially. Therefore, it can also be restricted that each target single character generates at most m printed misspelled character images. If there are a total of n combinations of target strokes, m combinations can be selected from the n combinations to generate the corresponding printed misspelled character images, which can be randomly selected or selected according to preset conditions, and no restrictions are imposed here.

[0029] In practical applications, the stroke information includes the pixel point information of the target stroke in the standard printed image. Specifically, the pixel point information includes the coordinate positions of the pixel points of the target stroke in the standard printed image. The pixel points of the target stroke can be all the pixel points of the target stroke or multiple key pixel points, and no restrictions are imposed here. In addition, the stroke information can also include the stroke order information (i.e., the writing order information) of the target stroke in the target single character, and the content included in the stroke information can be flexibly set according to requirements. In practical applications, the stroke information of each stroke of the target single character can be preset. The stroke information can be represented in the following format: target single character - stroke order number - pixel point coordinates. Taking "技" as an example, the stroke information corresponding to its first stroke can be recorded as "技—1—横:[(x1,y1),(x2,y2)...(xn,yn)]", the stroke information corresponding to its second stroke can be recorded as "技—2—竖勾:[(x1,y1),(x2,y2)...(xn,yn)]", and the stroke information corresponding to its seventh stroke can be recorded as "技—7—捺:[(x1,y1),(x2,y2)...(xn,yn)]", where (x,y) are the horizontal and vertical coordinates of the pixel coordinates. The stroke information can be manually marked or automatically marked using a model, and no restrictions are imposed here. The above stroke information is mainly used for subsequent erasure operations, so it can also be called mask information. In practical applications, the stroke information of each stroke can be预先获取并按序记录,以便后续针对目标单字的笔画进行抹除时,直接从记录结果中查找该笔画对应的笔画信息即可。

[0030] Step S106, perform an erasing operation based on the stroke information of the target stroke to obtain a printed misspelled character image corresponding to the target single character. The embodiments of the present disclosure do not limit the erasing method, and any method capable of erasing the target stroke can be used. For ease of understanding, reference can be made to Figure 3 a schematic diagram of a printed misspelled character image shown in the figure. This schematic diagram is a printed misspelled character image obtained by erasing the third stroke of the target single character "ji". In practical applications, when performing an erasing operation based on the stroke information of the target stroke to obtain an image after erasing the target stroke, in order to prevent the target single character from being converted into another correct single character after erasing the target stroke, the single character in the image after erasing the target stroke can be further discriminated. Only when it is confirmed that the single character is a misspelled character, the printed misspelled character image corresponding to the target single character is confirmed to be generated.

[0031] Step S108, perform style transfer processing on the printed misspelled character image to obtain a handwritten misspelled character image with a preset handwritten style.

[0032] In practical applications, the style transfer processing of the printed misspelled character image can be performed through a preset generative adversarial network. The generative adversarial network can be, for example, CGAN (Conditional Generative Adversarial Nets) or CycleGAN (Cycle Generative Adversarial Nets), etc., which are not limited here. Specifically, the generative adversarial network can be obtained through pre-training. For example, the initial network is trained using printed standard image samples and handwritten image samples with a preset handwritten style until a generative adversarial network capable of converting printed images into handwritten images is obtained. For ease of understanding, reference can be made to Figure 4 a schematic diagram of a handwritten misspelled character image shown in the figure, showing the conversion of a printed misspelled character image with the third stroke missing corresponding to "ji" into a handwritten misspelled character image with the third stroke missing.

[0033] Through the above method, the handwritten misspelled character image can be automatically generated conveniently and quickly, greatly reducing the acquisition cost of the misspelled character image.

[0034] In some embodiments, in order to further improve the convenience and reliability of obtaining the printed misspelled character image, the step of performing an erasing operation based on the stroke information of the target stroke to obtain a printed misspelled character image corresponding to the target single character can be performed according to the following steps A and steps B:

[0035] Step A, perform an erasing operation based on the stroke information of the target stroke to obtain a single character image after erasing the target stroke.

[0036] In some specific implementation examples, the stroke information includes the pixel point information of the target stroke in the standard printed image, specifically including the position coordinates of the pixel points. When step A is specifically executed, based on the pixel point information of the target stroke, the color value corresponding to the target pixel point in the target stroke can be set to the target color value; among them, the target color value is determined based on the base color value of the standard printed image, and the target pixel point does not belong to the strokes other than the target stroke in the target single character. For example, assuming that the background color of the standard printed image is white, the color values of the target pixel values can be set to white, that is, set to 255. By pre-obtaining the pixel point information of the target stroke and then directly adjusting the color value of the target pixel point of the target stroke to the background color based on this, it is more convenient and fast. In addition, the embodiments of the present disclosure process the target pixel points that only belong to the target stroke, and the target pixel points do not belong to other strokes. For example, the pixel points that intersect with other strokes in the target stroke do not belong to the target pixel points, avoiding erasing the intersection part of the target stroke and other strokes. Through the above method, the reliability of the erasing operation can be fully guaranteed.

[0037] Step B, when the single character in the single character image after erasing the target stroke does not belong to the correct single character, the single character image after erasing the target stroke is used as the printed misspelled character image corresponding to the target single character.

[0038] The embodiments of the present disclosure take into account that there may still be a situation where the single character image after erasing the target stroke still becomes another correct single character. For example, when the first stroke of the character "王" is erased, it becomes the character "土", and when the third stroke is erased, it becomes the character "三". Therefore, it will further determine whether the single character in the single character image after erasing the target stroke belongs to the correct single character. Only when it is confirmed that the single character in the single character image after erasing the target stroke does not belong to the correct single character, will the single character image after erasing the target stroke be used as the printed misspelled character image corresponding to the target single character.

[0039] In practical applications, the confidence levels of the single-character image after the target stroke is erased can be obtained for each single-character category through a pre-set single-character classification model. Then, when the maximum value among the confidence levels is not greater than a pre-set first confidence threshold, it is confirmed that the single character in the single-character image after the target stroke is erased does not belong to the correct single character. Specifically, during implementation, a single-character classification model can be trained using font image samples (all of which are correct single characters). The number of single-character categories corresponding to the single-character classification model is the number of single-character types included in the font image samples. For example, "one" and "two" are considered two categories, but the regular script form of "one" and the bold form of "one" still count as one category. The structure of the single-character classification model is not limited in the embodiments of the present disclosure. Exemplarily, the single-character classification model includes a CNN layer, a linear layer, and a softmax layer. In practical applications, the single-character image to be processed is input into the single-character classification model, and the single-character classification model can output the confidence levels of the image to be processed for each single-character category and determine the single character corresponding to the image to be processed based on the highest confidence level. Therefore, the confidence levels of the single-character image after the target stroke is erased can be obtained for each single-character category through the single-character classification model. If the highest confidence level is not greater than the first confidence threshold, it indicates that the character in the single-character image after the target stroke is erased does not belong to a correct single character, that is, the character in the single-character image after the target stroke is erased is a wrong character. On the contrary, if the highest confidence level is greater than the first confidence threshold, it indicates that the character in the single-character image after the target stroke is erased belongs to a correct single character, that is, the character in the single-character image after the target stroke is erased is another correct character. Through the above method, the situation where the target single character is converted into another correct character after the stroke is erased can be reliably excluded, fully ensuring the effectiveness of the printed wrong-character image.

[0040] On the basis of obtaining the printed wrong-character image, in order to further ensure the effect of the finally obtained handwritten wrong-character image, in some embodiments, the steps of performing style transfer processing on the printed wrong-character image include: performing style transfer processing on the printed wrong-character image when the printed wrong-character image meets the pre-set conditions. That is, the printed wrong-character image is further screened through the pre-set conditions, and only the printed wrong-character image that meets the pre-set conditions is subjected to style transfer processing to obtain the corresponding handwritten wrong-character image.

[0041] In some specific implementation examples, the preset conditions include: the maximum value of the confidence levels corresponding to the printed misspelled character image obtained through a preset single-character classification model for each single-character category is not less than a preset second confidence threshold. The single-character classification model can refer to the foregoing related content and will not be elaborated herein. By setting the second confidence threshold, it can be ensured as much as possible that the printed misspelled character image has a certain similarity to the correct single character, that is, has a certain degree of confusion with the correct single character, so as to be more challenging for subsequent actual misspelled character discrimination. The obtained handwritten misspelled character image can effectively become difficult samples required by network models such as misspelled character recognition models that need to be trained using misspelled character images, and can provide greater gain during model training.

[0042] Furthermore, the method provided in the embodiments of the present disclosure further includes: obtaining the application scenario type of the handwritten misspelled character image; and determining the label corresponding to the handwritten misspelled character image based on the application scenario type. Exemplarily, in the case where the application scenario type is a stroke writing sequence recognition scenario, the label corresponding to the handwritten misspelled character image includes the writing sequence of the remaining strokes of the target single character after erasing the target stroke; for example, Figure 4 the label corresponding to the misspelled character with the third stroke missing as shown is "horizontal - vertical hook - horizontal - vertical - horizontal left-falling stroke - right-falling stroke". In the case where the application scenario type is a misspelled character recognition scenario, the label corresponding to the handwritten misspelled character image includes the target single character. For example, Figure 4 the label corresponding to the misspelled character with the third stroke missing as shown is "ji". Among them, the above application scenario type can be determined based on the type of the model that needs to be trained using the handwritten misspelled character image. The embodiments of the present disclosure can further carry corresponding labels for the handwritten misspelled character image for model training with the corresponding labels. For example, if training a model that requires stroke writing sequence recognition ability, the handwritten misspelled character image carrying the writing sequence information of the remaining strokes of the target single character after erasing the target stroke can be used to train the model. If training a model that requires misspelled character recognition ability, the handwritten misspelled character image carrying the target single character (that is, the original correct single character corresponding to the misspelled character) can be used to train the model. Through the above method, misspelled character negative examples can be effectively provided for the model to avoid model overfitting. For example, if using image samples all of which are correct single characters to train a model that requires stroke writing sequence recognition ability, it may cause the model to only learn the simple mapping relationship from the image to the character sequence, and it is difficult to ensure that the model learns the true stroke order and stroke features. By introducing the handwritten misspelled character image with labels provided in the embodiments of the present disclosure, it helps to further improve the recognition accuracy and generalization ability of such models. [[ID=⑧]] [[ID=⑨]]

[0043] Furthermore, it should be noted that the method for generating misspelled characters provided in this embodiment first generates a printed misspelled character image based on a standard printed image, and then converts it into a handwritten misspelled character image, rather than generating a handwritten image first and then erasing the strokes. The main reason for this is that in practice, it has been found that most available handwritten text is written by adults, and some strokes are connected or simplified, which, while aesthetically pleasing, are not conducive to stroke segmentation and erasure. In contrast, the stroke structure of a standard printed font is clear, making it more advantageous in annotating stroke information and in the final font style conversion.

[0044] In summary, the method for generating misspelled characters provided in this embodiment of the present disclosure does not require the high cost of obtaining misspelled characters manually. Instead, it can automatically generate a large number of handwritten misspelled character images at low cost. This is not only convenient and fast, but also greatly expands the number of handwritten misspelled character images, improves the problem of the scarcity of existing misspelled character images, and can further promote the development of misspelled character recognition technology and improve the accuracy of network models in recognizing misspelled characters.

[0045] Corresponding to the aforementioned method for generating misspelled images, this disclosure also provides an apparatus for generating misspelled images. Figure 5 This is a schematic diagram of a device for generating misspelled images according to an embodiment of the present disclosure. This device can be implemented by software and / or hardware, and is generally integrated into an electronic device. Figure 5 As shown, the typo image generation device 500 includes:

[0046] Image acquisition module 502 is used to acquire the standard printed image of the target character;

[0047] The target stroke acquisition module 504 is used to acquire the stroke information of the target stroke to be erased in the standard printed image of the target character.

[0048] The stroke erasure module 506 is used to perform an erasure operation based on the stroke information of the target stroke to obtain the printed text image of the target character with the error.

[0049] Style transfer module 508 is used to perform style transfer processing on printed misspelling images to obtain handwritten misspelling images with a preset handwriting style.

[0050] The above-mentioned device can conveniently and quickly generate images of handwritten misspelled words automatically, greatly reducing the cost of obtaining misspelled word images.

[0051] In some implementations, when the number of target strokes is one, the target stroke is any stroke in the target character; when the total number of strokes in the target character is not less than a preset threshold and the number of target strokes is greater than one, multiple target strokes do not constitute the radical of the target character.

[0052] In some implementations, the stroke erasure module 506 is specifically used to: perform an erasure operation based on the stroke information of the target stroke to obtain a single character image after the target stroke is erased; if the single character in the single character image after the target stroke is erased is not a correct single character, use the single character image after the target stroke is erased as the printed text error image corresponding to the target single character.

[0053] In some implementations, the stroke information includes the pixel information of the target stroke in the standard image of the printed text. The stroke erasure module 506 is specifically used to: set the color value corresponding to the target pixel in the target stroke to the target color value based on the pixel information of the target stroke; wherein, the target color value is determined based on the background color value of the standard image of the printed text, and the target pixel does not belong to any stroke in the target character other than the target stroke.

[0054] In some embodiments, the above-mentioned device further includes a single-character confirmation module, which is used to obtain the confidence level of the single-character image after the target strokes have been erased in each single-character category through a preset single-character classification model; if the maximum value of the confidence level is not greater than a preset first confidence level threshold, it is confirmed that the single character in the single-character image after the target strokes have been erased does not belong to the correct single character.

[0055] In some implementations, the style transfer module 508 is specifically used to perform style transfer processing on the printed text misspelling image when the misspelling image meets preset conditions.

[0056] In some implementations, the preset conditions include: the maximum confidence score of the printed misspelled image obtained by the preset single-character classification model in each single-character category is not less than the preset second confidence threshold.

[0057] In some implementations, the style transfer module 508 is specifically used to perform style transfer processing on printed misspelling images using a pre-trained generative adversarial network.

[0058] In some embodiments, the above-mentioned apparatus further includes: a label determination module, used to obtain the application scenario type of the handwritten misspelling image; and to determine the label corresponding to the handwritten misspelling image based on the application scenario type.

[0059] In some implementations, when the application scenario is a stroke writing sequence recognition scenario, the label corresponding to the handwritten misspelling image includes the writing sequence of the remaining strokes of the target character after the target stroke is erased; when the application scenario is a misspelling recognition scenario, the label corresponding to the handwritten misspelling image includes the target character.

[0060] The device for generating misspelled images provided in this disclosure can execute the method for generating misspelled images provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0061] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device embodiments can be referred to the corresponding process in the method embodiments, and will not be repeated here.

[0062] Exemplary embodiments of this disclosure also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of this disclosure.

[0063] Exemplary embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to embodiments of this disclosure.

[0064] Exemplary embodiments of this disclosure also provide a computer program product, including a computer program, wherein, when executed by a processor of a computer, the computer program is used to cause the computer to perform a method according to an embodiment of this disclosure.

[0065] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0066] Furthermore, embodiments of this disclosure can also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the method for generating misspelled images provided in embodiments of this disclosure. The computer-readable storage medium can be any combination of one or more readable media. A readable medium can be a readable signal medium or a readable storage medium. A readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0067] refer to Figure 6 The present invention describes a structural block diagram of an electronic device 600 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0068] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0069] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, output unit 607, storage unit 608, and communication unit 609. Input unit 606 can be any type of device capable of inputting information to electronic device 600. Input unit 606 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 607 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 608 may include, but is not limited to, disks and optical discs. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0070] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above. For example, in some embodiments, the method for generating misspelled images can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. In some embodiments, the computing unit 601 can be configured to perform the method for generating misspelled images by any other suitable means (e.g., by means of firmware).

[0071] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0072] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0073] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0074] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0075] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0076] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0077] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0078] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating a misspelled image, comprising: Obtain the standard printed image of the target word; Obtain the stroke information of the target stroke to be erased in the standard printed image of the target character; wherein, the number of the target strokes is greater than or equal to one; when the number of the target strokes is one, the target stroke is any stroke in the target character; when the total number of strokes of the target character is not less than a preset threshold and the number of the target strokes is greater than one, multiple target strokes do not constitute the radical of the target character. An erasure operation is performed based on the stroke information of the target stroke to obtain the printed misspelled image corresponding to the target character; The printed misspelling image is subjected to style transfer processing to obtain a handwritten misspelling image with a preset handwriting style.

2. The method for generating a misspelled image as described in claim 1, wherein, The step of performing an erasure operation based on the stroke information of the target stroke to obtain the printed misspelled image corresponding to the target character includes: An erasure operation is performed based on the stroke information of the target stroke to obtain a single character image after the target stroke is erased. If a character in the single-character image after the target strokes have been erased is not a correct character, the single-character image after the target strokes have been erased shall be used as the printed misspelling image corresponding to the target character.

3. The method for generating a misspelled image as described in claim 1 or 2, wherein, The stroke information includes the pixel information of the target stroke in the standard printed image, and the step of performing an erasure operation based on the stroke information of the target stroke includes: Based on the pixel information of the target stroke, the color value corresponding to the target pixel in the target stroke is set as the target color value; wherein, the target color value is determined based on the background color value of the standard printed image, and the target pixel does not belong to any stroke in the target character other than the target stroke.

4. The method for generating a misspelled image as described in claim 2, wherein, The method further includes: The confidence level of the target stroke-removed single character image under each single character category is obtained by using a preset single character classification model; If the maximum value of the confidence level is not greater than the preset first confidence level threshold, it is confirmed that the single character in the single character image after the target stroke is erased is not a correct single character.

5. The method for generating a misspelled image as described in claim 1, wherein, The step of performing style transfer processing on the printed text misspelling image includes: If the printed text misspelling image meets the preset conditions, style transfer processing is performed on the printed text misspelling image.

6. The method for generating a misspelled image as described in claim 5, wherein, The preset conditions include: the maximum confidence score of the printed misspelled image obtained by the preset single-character classification model in each single-character category is not less than the preset second confidence score threshold.

7. The method for generating a misspelled image as described in claim 1 or 5, wherein, The step of performing style transfer processing on the printed text misspelling image includes: The printed text error image is style-transfer processed using a pre-defined generative adversarial network.

8. The method for generating a misspelled image as described in claim 1, wherein, The method further includes: The application scenario type for obtaining the handwritten misspelling image; Based on the application scenario type, the label corresponding to the handwritten misspelling image is determined.

9. The method for generating a misspelled image as described in claim 8, wherein, In the case where the application scenario is a stroke writing sequence recognition scenario, the label corresponding to the handwritten misspelling image includes the writing sequence of the remaining strokes of the target character after erasing the target stroke; In the case where the application scenario is a misspelling recognition scenario, the label corresponding to the handwritten misspelling image includes the target character.

10. An apparatus for generating images of misspelled words, comprising: The image acquisition module is used to acquire the standard printed image of the target character; The target stroke acquisition module is used to acquire the stroke information of the target stroke to be erased in the standard printed image of the target character; wherein, the number of the target strokes is greater than or equal to one; when the number of the target strokes is one, the target stroke is any stroke in the target character; when the total number of strokes of the target character is not less than a preset threshold and the number of the target strokes is greater than one, multiple target strokes do not constitute the radical of the target character. The stroke erasure module is used to perform an erasure operation based on the stroke information of the target stroke to obtain the printed text image of the target character with the misspelling. The style transfer module is used to perform style transfer processing on the printed misspelling image to obtain a handwritten misspelling image with a preset handwriting style.

11. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method for generating a misspelled image according to any one of claims 1-9.

12. A computer-readable storage medium, wherein, The storage medium stores a computer program for executing the method for generating a misspelled image according to any one of claims 1-9.

Citation Information

Patent Citations

  • Erroneous-character input method and system

    CN102103571A

  • Character generation model training method and device, character generation method and device and equipment

    CN113792849A

  • Error character recognition method and device, electronic equipment and storage medium

    CN115294581A