Image processing method and device and storage medium
By detecting and traversing text lines in the ticket image, identifying overlapping areas and removing background text, extracting effective text of text overlapping areas in the ticket image, the problems of inaccurate text prediction and difficulty in processing multiple overlapping in the prior art are solved, and higher text recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510101480.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to accurately predict text in the text overlapping area in the ticket image, especially in the interference between layer 1 and layer 2, resulting in insufficient prediction of bottom plate text and machine text. Furthermore, when there are more than two overlapping texts, the prior art cannot effectively separate the text.
By obtaining the original image, detecting text lines and forming a collection of text lines, traversing the text lines collection, identifying that the text lines overlap with other text lines, the local image blocks are intercepted from the original image, the background text is removed, and the target local image block is generated, thereby extracting effective text.
This method can effectively avoid interference from background text, improve the accuracy of text recognition, and can handle the situation where two or more texts overlap.
Smart Images

Figure CN120013982A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image processing method, device and storage medium. Background Art
[0002] At present, for images with overlapping text areas such as bills (such as various invoices, receipts, bills, etc.), it is necessary to extract unstructured data from them and then convert them into structured data for further processing and analysis. However, due to the overlap in the image, the conventional OCR algorithm cannot extract valid text in the text overlapping area, so a text separation solution is urgently needed to extract the valid text in the text overlapping area in the image.
[0003] The related technology discloses a text separation scheme for text overlapping areas in bill images based on multiple layers. The principle is to use two layers to separate the text in the text overlapping areas in the bill image, where layer 1 is responsible for predicting the background text in the text overlapping areas in the bill image, and layer 2 is responsible for predicting the machine-printed text in the text overlapping areas in the bill image, thereby achieving text separation in the text overlapping areas in the bill image.
[0004] However, in the process of predicting the bottom text in layer 1, the machine-printed text interferes, resulting in inaccurate prediction of the bottom text. Correspondingly, in the process of predicting the machine-printed text in layer 2, the bottom text interferes, resulting in inaccurate prediction of the machine-printed text. In addition, when there are more than two overlapping texts in the text overlapping area of the bill image, the two layers obviously cannot achieve text separation in the text overlapping area of the bill image. Summary of the invention
[0005] In order to solve the above-mentioned problem that in the process of predicting the bottom text in layer 1, the machine-printed text is interfered, resulting in inaccurate prediction of the bottom text, and correspondingly in the process of predicting the machine-printed text in layer 2, the bottom text is interfered, resulting in inaccurate prediction of the machine-printed text. In addition, when there are more than two overlapping texts in the text overlapping area of the bill image, the two layers obviously cannot achieve the technical problem of text separation in the text overlapping area of the bill image, the embodiment of the present application provides an image processing method, device, electronic device and storage medium. The specific technical solution is as follows:
[0006] In a first aspect of an embodiment of the present application, an image processing method is first provided, the method comprising:
[0007] Acquire an original image, detect text lines in the original image, and store them in a text line set;
[0008] Traversing the text lines in the text line set;
[0009] When the text line overlaps with any other text line in the text line set, intercepting a local image block corresponding to the text line from the original image, removing background text in the local image block, and obtaining a target local image block;
[0010] The valid text corresponding to the text line is obtained according to the target local image block.
[0011] In an optional implementation, after traversing the text lines in the text line set, the method further includes:
[0012] Obtaining a first minimum bounding rectangular frame of the text line, and obtaining a second minimum bounding rectangular frame of any other text line in the text line set;
[0013] Determining, according to the first minimum bounding rectangular frame and the second minimum bounding rectangular frame, whether the text line overlaps with any other text line in the text line set;
[0014] Wherein, determining, according to the first minimum bounding rectangular frame and the second minimum bounding rectangular frame, whether the text line overlaps with any other text line in the text line set comprises:
[0015] Determine an intersection-and-union ratio of the first minimum bounding rectangular frame and the second minimum bounding rectangular frame;
[0016] When the IoU is greater than a preset IoU threshold, it is determined that the text line overlaps with any other text line in the text line set.
[0017] In an optional implementation, removing the background text in the local image block to obtain the target local image block includes:
[0018] Inputting the local image block into a pre-trained generator to generate a target local image block with background text removed;
[0019] The background text is incomplete text in the local image block and / or text in the local image block whose area ratio is less than a preset ratio threshold.
[0020] In an optional embodiment, the pre-trained generator is obtained by:
[0021] Acquire a sample image and a sample standard image, wherein the sample image has text overlap, and the sample standard image is an image in which background text is removed;
[0022] Inputting the sample image into a generator, generating a predicted image with background text removed, and determining a first loss between the predicted image and the sample standard image;
[0023] The generator is trained according to the first loss, and when the first loss converges, the training is stopped to obtain a pre-trained generator.
[0024] In an optional embodiment, the method further comprises:
[0025] Inputting the predicted image into a first discriminator to obtain a first discrimination result, and determining a second loss between the first discrimination result and a first sample discrimination result;
[0026] The first discriminator is trained according to the second loss, and when the second loss converges, the training is stopped to obtain a pre-trained first discriminator.
[0027] In an optional embodiment, the method further comprises:
[0028] Inputting the predicted image into a second discriminator to obtain a second discriminant result for each local area in the predicted image;
[0029] Determine, according to the second discrimination result and the second sample discrimination result, a third loss corresponding to each local area in the predicted image, and obtain a loss sum of the third losses;
[0030] The second discriminator is trained according to the sum of losses, and when the sum of losses converges, the training is stopped to obtain a pre-trained second discriminator.
[0031] In an optional implementation, the acquiring of the sample image and the sample standard image includes:
[0032] Acquire a background image, fill a first text in the background image to obtain a first text image, wherein no text exists in the background image; fill a second text near the first text to obtain a second text image, wherein the first text and the second text overlap; perform enhancement processing on the second text image to obtain a sample image, and perform enhancement processing on the first text image to obtain a sample standard image;
[0033] Alternatively, a text image is obtained, pixels are filled at the edge of the text image to obtain an extended text image, wherein there is no text overlap in the text image; a third text is filled near the text in the extended text image to obtain a third text image, wherein there is text overlap between the third text and the text in the extended text image; the third text image is cropped to obtain a cropped text image, wherein the size of the cropped text image is the same as the size of the text image; the cropped text image is enhanced to obtain a sample image, and the text image is enhanced to obtain a sample standard image;
[0034] Alternatively, a fourth text image and a fifth text image are obtained, wherein there is no text overlap between the fourth text image and the fifth text image; a local text image block is captured from the fourth text image, and the local text image block is processed and pasted to the fifth text image to obtain a sixth text image with text overlap; the sixth text image is enhanced to obtain a sample image, and the fifth text image is enhanced to obtain a sample standard image.
[0035] In an optional implementation, determining a first loss between the predicted image and the sample standard image includes:
[0036] Determining an absolute value loss and a perceptual loss between the predicted image and the sample standard image;
[0037] A weighted sum is performed on the absolute value loss and the perceptual loss to obtain a first loss between the predicted image and the sample standard image.
[0038] In an optional embodiment, the generator supervises the training using a first absolute value loss between the predicted image and the sample standard image and a loss function of a gradient between the predicted image and the sample standard image;
[0039] Determining a first loss between the predicted image and the sample standard image includes:
[0040] determining a first absolute value loss between the predicted image and the sample standard image, and an image gradient between the predicted image and the sample standard image;
[0041] Determine a second absolute value loss corresponding to the image gradient, and perform weighted summation on the first absolute value loss and the second absolute value loss to obtain a first loss.
[0042] In a second aspect of the embodiments of the present application, an image processing device is further provided, the device comprising:
[0043] An image acquisition module, used to acquire an original image, detect text lines in the original image, and store them in a text line set;
[0044] A text line traversal module, used for traversing the text lines in the text line set;
[0045] A text removal module, configured to intercept a local image block corresponding to the text line from the original image when the text line overlaps with any other text line in the text line set, remove background text in the local image block, and obtain a target local image block;
[0046] A text recognition module is used to obtain the valid text corresponding to the text line according to the target local image block.
[0047] In an optional implementation, the device further includes: an overlap detection module, specifically configured to:
[0048] Obtaining a first minimum bounding rectangular frame of the text line, and obtaining a second minimum bounding rectangular frame of any other text line in the text line set;
[0049] Determining, according to the first minimum bounding rectangular frame and the second minimum bounding rectangular frame, whether the text line overlaps with any other text line in the text line set;
[0050] Wherein, determining, according to the first minimum bounding rectangular frame and the second minimum bounding rectangular frame, whether the text line overlaps with any other text line in the text line set comprises:
[0051] Determine an intersection-and-union ratio of the first minimum bounding rectangular frame and the second minimum bounding rectangular frame;
[0052] When the IoU is greater than a preset IoU threshold, it is determined that the text line overlaps with any other text line in the text line set.
[0053] In an optional implementation, the text removal module is specifically used to:
[0054] Inputting the local image block into a pre-trained generator to generate a target local image block with background text removed;
[0055] The background text is incomplete text in the local image block and / or text in the local image block whose area ratio is less than a preset ratio threshold.
[0056] In an optional embodiment, the device further comprises:
[0057] A sample acquisition module, used to acquire a sample image and a sample standard image, wherein the sample image has overlapping texts, and the sample standard image is an image in which background texts are removed;
[0058] An image generation module, used for inputting the sample image into a generator to generate a predicted image with background text removed;
[0059] A first loss determination module, configured to determine a first loss between the predicted image and the sample standard image;
[0060] The first training module is used to train the generator according to the first loss, and stop the training when the first loss converges to obtain a pre-trained generator.
[0061] In an optional embodiment, the device further comprises:
[0062] a second loss determination module, configured to input the predicted image into the first discriminator, obtain a first discrimination result, and determine a second loss between the first discrimination result and a first sample discrimination result;
[0063] The second training module is used to train the first discriminator according to the second loss, and stop the training when the second loss converges to obtain a pre-trained first discriminator.
[0064] In an optional embodiment, the device comprises:
[0065] An image discrimination module, used for inputting the predicted image into a second discriminator to obtain a second discrimination result for each local area in the predicted image;
[0066] a loss sum determination module, configured to determine a third loss corresponding to each local area in the predicted image according to the second discrimination result and the second sample discrimination result, and obtain a loss sum of the third losses;
[0067] The third training module is used to train the second discriminator according to the sum of losses, and stop the training when the sum of losses converges to obtain a pre-trained second discriminator.
[0068] In an optional implementation, the sample acquisition module is specifically used to:
[0069] Acquire a background image, fill a first text in the background image to obtain a first text image, wherein no text exists in the background image; fill a second text near the first text to obtain a second text image, wherein the first text and the second text overlap; perform enhancement processing on the second text image to obtain a sample image, and perform enhancement processing on the first text image to obtain a sample standard image;
[0070] Alternatively, a text image is obtained, pixels are filled at the edge of the text image to obtain an extended text image, wherein there is no text overlap in the text image; a third text is filled near the text in the extended text image to obtain a third text image, wherein there is text overlap between the third text and the text in the extended text image; the third text image is cropped to obtain a cropped text image, wherein the size of the cropped text image is the same as the size of the text image; the cropped text image is enhanced to obtain a sample image, and the text image is enhanced to obtain a sample standard image;
[0071] Alternatively, a fourth text image and a fifth text image are obtained, wherein there is no text overlap between the fourth text image and the fifth text image; a local text image block is captured from the fourth text image, and the local text image block is processed and pasted to the fifth text image to obtain a sixth text image with text overlap; the sixth text image is enhanced to obtain a sample image, and the fifth text image is enhanced to obtain a sample standard image.
[0072] In an optional implementation, the first loss determination module is specifically configured to:
[0073] Determining an absolute value loss and a perceptual loss between the predicted image and the sample standard image;
[0074] A weighted sum is performed on the absolute value loss and the perceptual loss to obtain a first loss between the predicted image and the sample standard image.
[0075] In an optional embodiment, the generator supervises the training using a first absolute value loss between the predicted image and the sample standard image and a loss function of a gradient between the predicted image and the sample standard image; the first loss determination module is specifically used to:
[0076] determining a first absolute value loss between the predicted image and the sample standard image, and an image gradient between the predicted image and the sample standard image;
[0077] Determine a second absolute value loss corresponding to the image gradient, and perform weighted summation on the first absolute value loss and the second absolute value loss to obtain a first loss.
[0078] In a third aspect of the embodiments of the present application, there is further provided an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0079] Memory, used to store computer programs;
[0080] The processor is used to implement any image processing method described in the first aspect when executing the program stored in the memory.
[0081] In a fourth aspect of the embodiments of the present application, a storage medium is further provided, wherein instructions are stored in the storage medium, and when the storage medium is run on a computer, the computer executes the image processing method described in any one of the first aspects above.
[0082] In a fifth aspect of the embodiments of the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any of the above-mentioned image processing methods.
[0083] The technical solution provided by the embodiment of the present application obtains an original image, detects text lines in the original image, stores them in a text line set, traverses the text lines in the text line set, and when a text line overlaps with any other text line in the text line set, cuts out a local image block corresponding to the text line from the original image, removes background text in the local image block, obtains a target local image block, and obtains valid text corresponding to the text line based on the target local image block.
[0084] By detecting the text lines in the original image and traversing them, for the traversed text lines, if they overlap with any other text lines, the background text in the local image block corresponding to the text line is removed, and then the valid text corresponding to the text line is obtained based on this. By removing the background text in the local image block corresponding to the text line and obtaining the valid text corresponding to the text line based on the target local image block, the interference of the background text can be avoided, the recognition accuracy of the valid text is improved, and the situation of two texts overlapping can be dealt with. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0086] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0087] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0088] Figure 1 A schematic diagram of an implementation flow of an image processing method shown in an embodiment of the present application;
[0089] Figure 2 A schematic diagram showing a comparison of a local image block before and after background text is removed, shown in an embodiment of the present application;
[0090] Figure 3 It is a schematic diagram of an implementation flow of another image processing method shown in an embodiment of the present application;
[0091] Figure 4 A schematic diagram of an implementation flow of a network training method shown in an embodiment of the present application;
[0092] Figure 5 A schematic diagram of an implementation flow of a method for generating a sample training pair shown in an embodiment of the present application;
[0093] Figure 6 A schematic diagram of a comparison between a sample image and a sample standard image shown in an embodiment of the present application;
[0094] Figure 7 It is a schematic diagram of an implementation flow of another method for generating sample training pairs shown in an embodiment of the present application;
[0095] Figure 8 A schematic diagram showing a comparison between another sample image and a sample standard image shown in an embodiment of the present application;
[0096] Fig. 9 It is a schematic diagram of an implementation flow of another method for generating sample training pairs shown in an embodiment of the present application;
[0097] Fig.10 A schematic diagram showing a comparison between another sample image and a sample standard image shown in an embodiment of the present application;
[0098] Fig.11 A schematic diagram of the structure of an image processing device shown in an embodiment of the present application;
[0099] Fig.12 It is a schematic diagram of the structure of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0100] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0101] The disclosure below provides many different embodiments or examples to realize the different structures of the present application. In order to simplify the disclosure of the present application, the parts and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present application. In addition, the present application can repeat reference numbers and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0102] like Figure 1 FIG. 1 is a schematic diagram of an implementation flow of an image processing method provided in an embodiment of the present application. The method is applied to an electronic device and may specifically include the following steps:
[0103] S101, obtaining an original image, detecting text lines in the original image, and storing them in a text line set.
[0104] In the embodiment of the present application, an original image is obtained, wherein the original image may be a bill image, such as various invoices, receipts, bills, tickets, etc., or other images with overlapping text areas, which is not limited in the embodiment of the present application.
[0105] For the original image, multiple text lines are detected and stored in a text line set. The existing text line detection algorithm can be used to perform text line detection on the original image, so that multiple text lines in the original image can be detected.
[0106] S102, traversing text lines in the text line set.
[0107] In an embodiment of the present application, the text lines in the text line set are traversed, wherein the text lines in the text line set can be traversed according to the position sequence of the text lines in the original image.
[0108] S103, when the text line overlaps with any other text line in the text line set, intercepting a local image block corresponding to the text line from the original image, removing background text in the local image block, and obtaining a target local image block.
[0109] In an embodiment of the present application, for a traversed text line, if the text line overlaps with any other text line in a text line set, a local image block corresponding to the text line is captured from the original image, and the background text in the local image block corresponding to the text line is removed to obtain a target local image block.
[0110] It should be noted that background text usually refers to incomplete text in a local image block and / or text whose area share in a local image block is less than a preset share threshold, and the rest of the text is foreground text, which is not limited in the embodiments of the present application.
[0111] For the traversed text line, if the text line overlaps with any other text line in the text line set, for example Figure 2 The screenshot above has text overlap. The local image blocks corresponding to the text lines are captured from the original image. The local image blocks are as follows: Figure 2As shown in the screenshot above, the background text in the local image block corresponding to the text line is removed to obtain the target local image block. The target local image block is as follows: Figure 2 As shown in the screenshot below.
[0112] In addition, in an embodiment of the present application, for a traversed text line, when there is no overlap between the text line and any other text lines in a text line set, optical character recognition is performed on the local image block corresponding to the text line to obtain valid text corresponding to the text line.
[0113] S104, obtaining valid text corresponding to the text line according to the target local image block.
[0114] In the embodiment of the present application, the background text in the local image block is removed in the above steps, and the target local image block is obtained, and the valid text corresponding to the text line can be obtained according to the target local image block. Among them, optical character recognition can be performed on the target local image block to obtain the valid text corresponding to the text line.
[0115] It should be noted that for local image blocks, after background text processing, there may be some residues, that is, some background text may remain in the target local image block, but it does not affect optical character recognition. Therefore, optical character recognition can be directly performed on the target local image block to obtain the valid text corresponding to the text line.
[0116] Through the above description of the technical solution provided in the embodiment of the present application, an original image is obtained, and text lines in the original image are detected and stored in a text line set. The text lines in the text line set are traversed, and when a text line overlaps with any other text line in the text line set, a local image block corresponding to the text line is intercepted from the original image, and the background text in the local image block is removed to obtain a target local image block, and valid text corresponding to the text line is obtained based on the target local image block.
[0117] By detecting the text lines in the original image and traversing them, for the traversed text lines, if they overlap with any other text lines, the background text in the local image block corresponding to the text line is removed, and the valid text corresponding to the text line is obtained according to the target local image block. By removing the background text in the local image block corresponding to the text line and obtaining the valid text corresponding to the text line according to the target local image block, the interference of the background text can be avoided, the recognition accuracy of the valid text is improved, and the situation where two texts overlap can be dealt with.
[0118] In addition, in an embodiment of the present application, a pre-trained generative adversarial network is provided, which includes a pre-trained generator, a pre-trained first discriminator, and a pre-trained second discriminator. The pre-trained generator is responsible for generating a local image block with background text removed, the pre-trained first discriminator is responsible for judging whether there is text overlap in the local image block generated by the pre-trained generator as a whole, and the pre-trained second discriminator is responsible for judging whether there is text overlap in each local area of the local image block generated by the pre-trained generator.
[0119] In the application stage of the pre-trained generative adversarial network, only the pre-trained generator can be used, so Figure 3 FIG. 1 is a schematic diagram of an implementation flow of another image processing method provided in an embodiment of the present application. The method is applied to an electronic device and may specifically include the following steps:
[0120] S301, obtaining an original image, detecting text lines in the original image, and storing them in a text line set.
[0121] In the embodiment of the present application, this step is similar to the above-mentioned step S101, and the embodiment of the present application will not be described one by one here.
[0122] S302, traverse the text lines in the text line set.
[0123] In an embodiment of the present application, the text lines in the text line set are traversed, wherein the text lines in the text line set can be traversed according to the position sequence of the text lines in the original image.
[0124] In addition, for the traversed text line, it is necessary to determine whether there is text overlap between it and any other text line in the text line set. If not, optical character recognition can be directly performed on the local image block corresponding to the text line to obtain the valid text corresponding to the text line. Otherwise, subsequent steps need to be executed.
[0125] Based on this, for the traversed text line, the first minimum bounding rectangular box of the text line is obtained, and the second minimum bounding rectangular box of any other text line in the text line set is obtained, and based on the first minimum bounding rectangular box and the second minimum bounding rectangular box, it is determined whether the text line overlaps with any other text line in the text line set.
[0126] Specifically, an intersection-and-union ratio of the first minimum bounding rectangular frame and the second minimum bounding rectangular frame is determined, and when the intersection-and-union ratio is greater than a preset intersection-and-union ratio threshold, it is determined that the text line overlaps with any other text line in the text line set.
[0127] For example, obtain the first minimum enclosing rectangular box of text line 1, and obtain the second minimum enclosing rectangular box of any other text line in the text line set, determine the intersection and union ratio of the first minimum enclosing rectangular box and the second minimum enclosing rectangular box, and when the intersection and union ratio is greater than 0.1, determine that text line 1 overlaps with any other text line in the text line set.
[0128] S303, when a text line overlaps with any other text line in the text line set, a local image block corresponding to the text line is cut out from the original image, and the local image block is input into a pre-trained generator to generate a target local image block for removing background text.
[0129] In an embodiment of the present application, for a traversed text line, if the text line overlaps with any other text line in a text line set, a local image block corresponding to the text line that overlaps with any other text line in the text line set is captured from the original image and then input into a pre-trained generator to generate a target local image block for removing background text.
[0130] It should be noted that, in one embodiment of the present application, the pre-trained generator can be a convolutional neural network similar to U-Net, which is composed of an encoder and a decoder. The encoder gradually reduces the spatial resolution of the input local image block and extracts high-level features. The decoder gradually restores the spatial resolution and restores the extracted high-level features to the size of the local image block.
[0131] In addition, in one embodiment of the present application, for the pre-trained generator, in order to improve the global perception ability of the pre-trained generator, a self-attention mechanism can also be used. For example, Swin-transformer can be used as a pre-trained generator. Of course, other models using a self-attention mechanism can also be used as a pre-trained generator. This embodiment of the present application is not limited to this.
[0132] In addition, in an embodiment of the present application, for a traversed text line, when there is no overlap between the text line and any other text lines in a text line set, optical character recognition is performed on the local image block corresponding to the text line to obtain valid text corresponding to the text line.
[0133] S304, obtaining valid text corresponding to the text line according to the target local image block.
[0134] In the embodiment of the present application, this step is similar to the above-mentioned step S104, and the embodiment of the present application will not be described one by one here.
[0135] By detecting the text lines in the original image and traversing them, for the traversed text lines, if they overlap with any other text lines, the local image block corresponding to the text line is input into the pre-trained generator to generate the target local image block for removing the background text, and then the valid text corresponding to the text line is obtained based on this. In this way, the background text in the local image block corresponding to the text line is removed, and the valid text corresponding to the text line is obtained according to the target local image block. This can avoid the interference of background text, improve the recognition accuracy of the valid text, and can deal with the situation where two texts overlap.
[0136] In addition, the above-mentioned pre-trained generative adversarial network needs to be trained in advance. Figure 4 FIG. 1 is a schematic diagram of an implementation flow of a network training method provided in an embodiment of the present application. The method is applied to an electronic device and may specifically include the following steps:
[0137] S401, obtaining a sample image and a sample standard image, wherein there is text overlap in the sample image, and the sample standard image is an image with background text removed.
[0138] In an embodiment of the present application, a sample image and a sample standard image are obtained, there is text overlap in the sample image, and the sample standard image is an image with background text removed, which means that there is no text overlap in the sample standard image.
[0139] Among them, for the sample image and the sample standard image, the embodiment of the present application provides three generation methods, and each generation method is described below.
[0140] For the first generation method, Figure 5 FIG. 1 is a schematic diagram of an implementation flow of a method for generating a sample training pair provided in an embodiment of the present application. The method is applied to an electronic device and may specifically include the following steps:
[0141] S501, obtaining a background image, filling the background image with a first text, and obtaining a first text image, wherein no text exists in the background image.
[0142] In the embodiment of the present application, a background image is obtained, and there is no text in the background image, which means that any image without text can be used as the background image. Among them, an image without text can be randomly obtained as the background image, and the embodiment of the present application does not limit this.
[0143] For the background image, the first text is filled in the background image to obtain the first text image. The font color, font size, font transparency, font type, text position, text content, etc. of the filled first text can be random, and the first text can be filled in a random position of the background image, which is not limited in the embodiment of the present application.
[0144] S502, filling a second text near the first text to obtain a second text image, wherein the first text and the second text overlap.
[0145] In an embodiment of the present application, a second text is filled near a first text to obtain a second text image, wherein there is text overlap between the first text and the second text, which means that the texts filled twice in the background image overlap.
[0146] It should be noted that the font color, font size, font transparency, font type, text position, text content, etc. of the filled second text can all be random, and the embodiments of the present application do not limit this.
[0147] In addition, the vicinity of the first text usually refers to the edge of the first text, so the second text is filled on the edge of the first text, so that the filled second text overlaps with the first text, and the overlapping part is a small part, so the first text serves as the foreground text and the second text serves as the background text.
[0148] S503, performing enhancement processing on the second text image to obtain a sample image, and performing enhancement processing on the first text image to obtain a sample standard image.
[0149] In the embodiment of the present application, the second text image is enhanced to obtain a sample image, and the first text image is enhanced to obtain a sample standard image. The same enhancement processing method is used for the second text image and the first text image, and the sample image and the sample standard image obtained in this way can be as follows: Figure 6 As shown, Figure 6 The upper screenshot shown is a sample image, and the lower screenshot is a sample standard image.
[0150] It should be noted that the enhancement processing may specifically include one or more of the following methods: adding noise (such as shadows, moiré, compression noise, blur), randomly bending images, perspective transformation, cropping, etc. The embodiments of the present application do not limit this, nor do they limit the processing order of adding noise, randomly bending images, perspective transformation, cropping, etc.
[0151] In one embodiment of the present application, the second text image is enhanced to obtain a sample image, including sequentially performing noise addition, random image bending, perspective transformation and cropping operations on the second text image to obtain a sample image, and similarly, sequentially performing noise addition, random image bending, perspective transformation and cropping operations on the first text image to obtain a sample standard image.
[0152] For the second generation method, Figure 7 FIG. 1 is a schematic diagram of an implementation flow of another method for generating a sample training pair provided in an embodiment of the present application. The method is applied to an electronic device and may specifically include the following steps:
[0153] S701, acquiring a text image, and filling pixels at the edge of the text image to obtain an expanded text image, wherein there is no text overlap in the text image.
[0154] In the embodiment of the present application, a text image is obtained, and there is no text overlap in the text image. The text image without text overlap can be obtained randomly, and the embodiment of the present application does not limit this.
[0155] For the text image, pixels are filled at the edge of the text image to obtain an expanded text image. The edge of the text image may be filled with pixels of, for example, white color, which is not limited in the embodiment of the present application.
[0156] S702: Fill third text near the text in the expanded text image to obtain a third text image, wherein the third text overlaps with the text in the expanded text image.
[0157] In an embodiment of the present application, a third text is filled near the text in the expanded text image to obtain a third text image, wherein the third text overlaps with the text in the expanded text image, meaning that after the text image is pixel expanded, the filled text overlaps with the original text in the text image.
[0158] It should be noted that, the vicinity of the text in the expanded text image usually refers to the edge of the text in the expanded text image. In this way, the third text is filled in the edge of the text in the expanded text image, so that the filled third text overlaps with the text in the expanded text image, and the overlapping part of the text is a small part. In this way, the text in the expanded text image serves as the foreground text, and the third text serves as the background text. The embodiments of the present application are not limited to this.
[0159] S703, cropping the third text image to obtain a cropped text image, wherein the size of the cropped text image is the same as the size of the text image.
[0160] In the embodiment of the present application, the third text image is cropped to obtain a cropped text image, wherein the size of the cropped text image is the same as the size of the text image, which means that the third text image is cropped to the same size as the text image.
[0161] It should be noted that, after the above-mentioned cropping process, the filled third text may form an incomplete text, which is not limited in the embodiment of the present application.
[0162] S704, performing enhancement processing on the cropped text image to obtain a sample image, and performing enhancement processing on the text image to obtain a sample standard image.
[0163] In the embodiment of the present application, the cropped text image is enhanced to obtain a sample image, and the text image is enhanced to obtain a sample standard image. The same enhancement processing method is used for the cropped text image and the text image, and the obtained sample image and sample standard image can be as follows: Figure 8 As shown, Figure 8 The upper screenshot shown is a sample image, and the lower screenshot is a sample standard image.
[0164] It should be noted that the enhancement processing may specifically include one or more of the following methods: adding noise (such as shadows, moiré, compression noise, blur), randomly bending images, perspective transformation, cropping, etc. The embodiments of the present application do not limit this, nor do they limit the processing order of adding noise, randomly bending images, perspective transformation, cropping, etc.
[0165] In one embodiment of the present application, a cropped text image is enhanced to obtain a sample image, including sequentially performing noise addition, random image bending, perspective transformation and cropping operations on the cropped text image to obtain a sample image. Similarly, noise addition, random image bending, perspective transformation and cropping operations are sequentially performed on the text image to obtain a sample standard image.
[0166] For the third generation method, such as Fig. 9 FIG. 1 is a schematic diagram of an implementation flow of another method for generating a sample training pair provided in an embodiment of the present application. The method is applied to an electronic device and may specifically include the following steps:
[0167] S901, obtaining a fourth text image and a fifth text image, wherein there is no text overlap between the fourth text image and the fifth text image.
[0168] In the embodiment of the present application, the fourth text image and the fifth text image are obtained, and there is no text overlap between the fourth text image and the fifth text image, which means obtaining two text images that do not have text overlap. The fourth text image and the fifth text image can be obtained randomly, and the embodiment of the present application does not limit this.
[0169] S902, intercepting a local text image block from the fourth text image, processing the local text image block and pasting it to the fifth text image to obtain a sixth text image with overlapping text.
[0170] In the embodiment of the present application, a local text image block is captured from the fourth text image, processed and then pasted to the fifth text image to obtain a sixth text image with overlapping text.
[0171] Partial text image blocks may be randomly captured from the fourth text image, and the partial text image blocks may be made transparent according to a certain transparency, and then pasted to the fifth text image to obtain a sixth text image with overlapping text.
[0172] It should be noted that after the partial text image block in the fourth text image is intercepted and made transparent, it is pasted to the fifth text image so as to overlap with the text in the fifth text image.
[0173] S903, performing enhancement processing on the sixth text image to obtain a sample image, and performing enhancement processing on the fifth text image to obtain a sample standard image.
[0174] In the embodiment of the present application, the sixth text image is enhanced to obtain a sample image, and the fifth text image is enhanced to obtain a sample standard image. The same enhancement processing method is used for the sixth text image and the fifth text image, and the obtained sample image and sample standard image can be as follows: Fig.10 As shown, Fig.10 The upper screenshot shown is a sample image, and the lower screenshot is a sample standard image.
[0175] It should be noted that the enhancement processing may specifically include one or more of the following methods: adding noise (such as shadows, moiré, compression noise, blur), randomly bending images, perspective transformation, cropping, etc. The embodiments of the present application do not limit this, nor do they limit the processing order of adding noise, randomly bending images, perspective transformation, cropping, etc.
[0176] In one embodiment of the present application, the sixth text image is enhanced to obtain a sample image, including sequentially performing noise addition, random image bending, perspective transformation and cropping operations on the sixth text image to obtain a sample image, and similarly, sequentially performing noise addition, random image bending, perspective transformation and cropping operations on the fifth text image to obtain a sample standard image.
[0177] In addition, the above is to intercept a partial text image block in the fourth text image and perform a transparency process, and then paste it into the fifth text image, so as to overlap with the text in the fifth text image. Conversely, it is also possible to intercept a partial text image block in the fifth text image and perform a transparency process, and then paste it into the fourth text image, so as to overlap with the text in the fourth text image.
[0178] Based on this, a fourth text image and a fifth text image are obtained, where there is no text overlap between the fourth text image and the fifth text image, a local text image block is captured from the fifth text image, the local text image block is made transparent, and the local text image block after the transparency processing is pasted to the fourth text image to obtain a seventh text image with text overlap, the seventh text image is enhanced to obtain a sample image, and the fourth text image is enhanced to obtain a sample standard image.
[0179] It should be noted that, although the above three generation methods are different, they all follow a principle, that is, they all fill overlapping texts on one text, and the overlapping texts can only occupy part of the length of the image, which is not limited in the embodiments of the present application.
[0180] S402, inputting the sample image into a generator, generating a predicted image with background text removed, and determining a first loss between the predicted image and the sample standard image.
[0181] In an embodiment of the present application, a sample image is input into a generator, a predicted image with background text removed is generated, and a first loss between the predicted image and the sample standard image is determined.
[0182] It should be noted that the generator can be a convolutional neural network similar to U-Net, which consists of an encoder and a decoder. The encoder gradually reduces the spatial resolution of the input sample image and extracts high-level features. The decoder gradually restores the spatial resolution and restores the extracted high-level features to the sample image size.
[0183] In addition, in one embodiment of the present application, for the generator, in order to improve the global perception ability of the generator, a self-attention mechanism can also be used. For example, Swin-transformer can be used as a generator. Of course, other models that use a self-attention mechanism can also be used as a generator. This embodiment of the present application is not limited to this.
[0184] For the above-mentioned first loss, it can be a weighted sum of L1 Loss and perceptual loss. To this end, the absolute value loss (i.e., L1 Loss) and perceptual loss between the predicted image and the sample standard image are determined, and the absolute value loss and the perceptual loss are weightedly summed to obtain the first loss between the predicted image and the sample standard image.
[0185] In addition, since text is high-frequency information and the gradient of high-frequency information is very obvious, during the network training process, the image gradient between the predicted image and the sample standard image can also be calculated, and L1 Loss can be used to supervise the image gradient of the two, so that the results generated by the generator are more in line with expectations.
[0186] That is, in one embodiment of the present application, when the generator is trained, the first absolute value loss L1 Loss between the predicted image and the sample standard image and the gradient loss function are used together to supervise the model training, so that the results generated by the generator are more in line with expectations.
[0187] Therefore, the first absolute value loss between the predicted image and the sample standard image, as well as the image gradient between the predicted image and the sample standard image can also be determined; L1 Loss is used to determine the second absolute value loss corresponding to the image gradient, and the first absolute value loss and the second absolute value loss are weightedly summed to obtain the first loss, as shown below.
[0188]
[0189] Among them, Loss is the first loss, output is the predicted image, Ground truth is the sample standard image, L1 is L1 Loss, the first absolute value loss, the distance from the predicted image to the sample standard image; is the image gradient function, β is the weight coefficient, The loss function that represents the gradient of the image.
[0190] S403, training the generator according to the first loss, and stopping the training when the first loss converges to obtain a pre-trained generator.
[0191] In an embodiment of the present application, the generator is trained according to the first loss, and when the first loss converges (for example, the first loss is less than a certain loss threshold), the training is stopped to obtain a pre-trained generator.
[0192] S404: input the predicted image to the first discriminator to obtain a first discrimination result, and determine a second loss between the first discrimination result and the first sample discrimination result.
[0193] S405, training the first discriminator according to the second loss, and stopping the training when the second loss converges, to obtain a pre-trained first discriminator.
[0194] In the embodiment of the present application, the predicted image is input to the first discriminator to obtain the first discrimination result, and the second loss between the first discrimination result and the first sample discrimination result (for example, the first sample discrimination result can be 1) is determined, and the first discriminator is trained according to the second loss, and when the second loss converges (for example, the second loss is less than a certain loss threshold), the training is stopped to obtain the pre-trained first discriminator. Among them, the first discriminator is responsible for judging whether there is text overlap in the predicted image generated by the generator from the whole.
[0195] S406: Input the predicted image to a second discriminator to obtain a second discrimination result for each local area in the predicted image.
[0196] S407, determining a third loss corresponding to each local area in the predicted image according to the second discrimination result and the second sample discrimination result, and obtaining a loss sum of the third losses.
[0197] S408, training the second discriminator according to the sum of losses, and stopping the training when the sum of losses converges, to obtain a pre-trained second discriminator.
[0198] In an embodiment of the present application, the predicted image is input into the second discriminator to obtain a second discrimination result for each local area in the predicted image, and the third loss corresponding to each local area in the predicted image is determined according to the second discrimination result and the second sample discrimination result (for example, 1), and the sum of the losses of the third losses is obtained. According to the sum of the losses, the second discriminator is trained, and when the sum of the losses converges (for example, the sum of the losses is less than a certain loss threshold), the training is stopped to obtain a pre-trained second discriminator.
[0199] Among them, in one embodiment of the present invention, the second discriminator can be a Markov discriminator. When the second discriminator is a Markov discriminator, the predicted image is regarded as a combination of a series of local regions, and the Markov hypothesis is used to focus on whether each local region of the predicted image is reasonable, rather than the overall predicted image. This method is called a Markov discriminator or PatchGAN. For the Markov discriminator, the predicted image is divided into small local regions p1, p2, ..., pn, and each local region is discriminated separately. The final loss function is the sum of all local region losses, and its formula is as follows.
[0200]
[0201] Wherein, pi represents the i-th local area, D(pi) is the second discrimination result of the i-th local area, and Loss is the sum of the losses of all local areas.
[0202] After the above training, a pre-trained generative adversarial network can be obtained. Subsequently, the text lines in the original image are detected and traversed. For the traversed text lines, when they overlap with any other text lines, the local image block corresponding to the text line is input into the pre-trained generator to generate a target local image block for removing the background text, and then the valid text corresponding to the text line is obtained based on this. In this way, the background text in the local image block corresponding to the text line is removed, and the valid text corresponding to the text line is obtained according to the target local image block. This can avoid the interference of background text, improve the recognition accuracy of the valid text, and can deal with the situation where two texts overlap.
[0203] Corresponding to the above method embodiment, the present application embodiment also provides an image processing device, such as Fig.11 As shown, the device may include: an image acquisition module 1110 , a text line traversal module 1120 , a text removal module 1130 , and a text recognition module 1140 .
[0204] The image acquisition module 1110 is used to acquire an original image, detect text lines in the original image, and store them in a text line set;
[0205] A text line traversal module 1120, configured to traverse the text lines in the text line set;
[0206] A text removal module 1130 is used to intercept a local image block corresponding to the text line from the original image when the text line overlaps with any other text line in the text line set, remove background text in the local image block, and obtain a target local image block;
[0207] The text recognition module 1140 is used to obtain the valid text corresponding to the text line according to the target local image block.
[0208] In an optional implementation, the device further includes: an overlap detection module, specifically configured to:
[0209] Obtaining a first minimum bounding rectangular frame of the text line, and obtaining a second minimum bounding rectangular frame of any other text line in the text line set;
[0210] Determining, according to the first minimum bounding rectangular frame and the second minimum bounding rectangular frame, whether the text line overlaps with any other text line in the text line set;
[0211] Wherein, determining, according to the first minimum bounding rectangular frame and the second minimum bounding rectangular frame, whether the text line overlaps with any other text line in the text line set comprises:
[0212] Determine an intersection-and-union ratio of the first minimum bounding rectangular frame and the second minimum bounding rectangular frame;
[0213] When the IoU is greater than a preset IoU threshold, it is determined that the text line overlaps with any other text line in the text line set.
[0214] In an optional implementation, the text removal module is specifically used to:
[0215] Inputting the local image block into a pre-trained generator to generate a target local image block with background text removed;
[0216] The background text is incomplete text in the local image block and / or text in the local image block whose area ratio is less than a preset ratio threshold.
[0217] In an optional embodiment, the device further comprises:
[0218] A sample acquisition module, used to acquire a sample image and a sample standard image, wherein the sample image has overlapping texts, and the sample standard image is an image in which background texts are removed;
[0219] An image generation module, used for inputting the sample image into a generator to generate a predicted image with background text removed;
[0220] A first loss determination module, configured to determine a first loss between the predicted image and the sample standard image;
[0221] The first training module is used to train the generator according to the first loss, and stop the training when the first loss converges to obtain a pre-trained generator.
[0222] In an optional embodiment, the device further comprises:
[0223] a second loss determination module, configured to input the predicted image into the first discriminator, obtain a first discrimination result, and determine a second loss between the first discrimination result and a first sample discrimination result;
[0224] The second training module is used to train the first discriminator according to the second loss, and stop the training when the second loss converges to obtain a pre-trained first discriminator.
[0225] In an optional embodiment, the device comprises:
[0226] An image discrimination module, used for inputting the predicted image into a second discriminator to obtain a second discrimination result for each local area in the predicted image;
[0227] a loss sum determination module, configured to determine a third loss corresponding to each local area in the predicted image according to the second discrimination result and the second sample discrimination result, and obtain a loss sum of the third losses;
[0228] The third training module is used to train the second discriminator according to the sum of losses, and stop the training when the sum of losses converges to obtain a pre-trained second discriminator.
[0229] In an optional implementation, the sample acquisition module is specifically used to:
[0230] Acquire a background image, fill a first text in the background image to obtain a first text image, wherein no text exists in the background image; fill a second text near the first text to obtain a second text image, wherein the first text and the second text overlap; perform enhancement processing on the second text image to obtain a sample image, and perform enhancement processing on the first text image to obtain a sample standard image;
[0231] Alternatively, a text image is obtained, pixels are filled at the edge of the text image to obtain an extended text image, wherein there is no text overlap in the text image; a third text is filled near the text in the extended text image to obtain a third text image, wherein there is text overlap between the third text and the text in the extended text image; the third text image is cropped to obtain a cropped text image, wherein the size of the cropped text image is the same as the size of the text image; the cropped text image is enhanced to obtain a sample image, and the text image is enhanced to obtain a sample standard image;
[0232] Alternatively, a fourth text image and a fifth text image are obtained, wherein there is no text overlap between the fourth text image and the fifth text image; a local text image block is captured from the fourth text image, and the local text image block is processed and pasted to the fifth text image to obtain a sixth text image with text overlap; the sixth text image is enhanced to obtain a sample image, and the fifth text image is enhanced to obtain a sample standard image.
[0233] In an optional implementation, the first loss determination module is specifically configured to:
[0234] Determining an absolute value loss and a perceptual loss between the predicted image and the sample standard image;
[0235] A weighted sum is performed on the absolute value loss and the perceptual loss to obtain a first loss between the predicted image and the sample standard image.
[0236] In an optional embodiment, the generator supervises the training using a first absolute value loss between the predicted image and the sample standard image and a loss function of a gradient between the predicted image and the sample standard image; the first loss determination module is specifically used for:
[0237] determining a first absolute value loss between the predicted image and the sample standard image, and an image gradient between the predicted image and the sample standard image;
[0238] Determine a second absolute value loss corresponding to the image gradient, and perform weighted summation on the first absolute value loss and the second absolute value loss to obtain a first loss.
[0239] The present application also provides an electronic device, such as Fig.12 As shown, it includes a processor 121, a communication interface 122, a memory 123 and a communication bus 124, wherein the processor 121, the communication interface 122, and the memory 123 communicate with each other through the communication bus 124.
[0240] Memory 123, used for storing computer programs;
[0241] The processor 121 is used to execute the program stored in the memory 123 to implement the following steps:
[0242] Acquire an original image, detect text lines in the original image, and store them in a text line set; traverse the text lines in the text line set; when the text line overlaps with any other text line in the text line set, intercept a local image block corresponding to the text line from the original image, remove background text in the local image block, and obtain a target local image block; obtain valid text corresponding to the text line based on the target local image block.
[0243] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0244] The communication interface is used for communication between the above electronic device and other devices.
[0245] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0246] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0247] In another embodiment provided in the present application, a storage medium is provided. The storage medium stores instructions. When the instructions are executed on a computer, the computer executes the image processing method described in any one of the above embodiments.
[0248] In another embodiment provided in the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute the image processing method described in any one of the above embodiments.
[0249] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a storage medium, or transmitted from one storage medium to another storage medium, for example, the computer instructions may be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.
[0250] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0251] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0252] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire an original image, detect text lines in the original image, and store them in a text line set; Traversing the text lines in the text line set; When the text line overlaps with any other text line in the text line set, intercepting a local image block corresponding to the text line from the original image, removing background text in the local image block, and obtaining a target local image block; The valid text corresponding to the text line is obtained according to the target local image block.
2. The method according to claim 1, characterized in that After traversing the text lines in the text line set, the method further includes: Obtaining a first minimum bounding rectangular frame of the text line, and obtaining a second minimum bounding rectangular frame of any other text line in the text line set; Determining, according to the first minimum bounding rectangular frame and the second minimum bounding rectangular frame, whether the text line overlaps with any other text line in the text line set; Wherein, determining, according to the first minimum bounding rectangular frame and the second minimum bounding rectangular frame, whether the text line overlaps with any other text line in the text line set comprises: Determine an intersection-and-union ratio of the first minimum bounding rectangular frame and the second minimum bounding rectangular frame; When the IoU is greater than a preset IoU threshold, it is determined that the text line overlaps with any other text line in the text line set.
3. The method according to claim 1, characterized in that The removing of the background text in the local image block to obtain the target local image block includes: Inputting the local image block into a pre-trained generator to generate a target local image block with background text removed; The background text is incomplete text in the local image block and / or text in the local image block whose area ratio is less than a preset ratio threshold.
4. The method according to claim 3, characterized in that The pre-trained generator is obtained in the following way: Acquire a sample image and a sample standard image, wherein the sample image has text overlap, and the sample standard image is an image in which background text is removed; Inputting the sample image into a generator, generating a predicted image with background text removed, and determining a first loss between the predicted image and the sample standard image; The generator is trained according to the first loss, and when the first loss converges, the training is stopped to obtain a pre-trained generator.
5. The method according to claim 4, characterized in that The method further comprises: Inputting the predicted image into a first discriminator to obtain a first discrimination result, and determining a second loss between the first discrimination result and a first sample discrimination result; The first discriminator is trained according to the second loss, and when the second loss converges, the training is stopped to obtain a pre-trained first discriminator.
6. The method according to claim 4, characterized in that The method further comprises: Inputting the predicted image into a second discriminator to obtain a second discriminant result for each local area in the predicted image; Determine, according to the second discrimination result and the second sample discrimination result, a third loss corresponding to each local area in the predicted image, and obtain a loss sum of the third losses; The second discriminator is trained according to the sum of losses, and when the sum of losses converges, the training is stopped to obtain a pre-trained second discriminator.
7. The method according to claim 4, characterized in that The obtaining of the sample image and the sample standard image comprises: Acquire a background image, fill a first text in the background image to obtain a first text image, wherein no text exists in the background image; fill a second text near the first text to obtain a second text image, wherein the first text and the second text overlap; perform enhancement processing on the second text image to obtain a sample image, and perform enhancement processing on the first text image to obtain a sample standard image; Alternatively, a text image is obtained, pixels are filled at the edge of the text image to obtain an extended text image, wherein there is no text overlap in the text image; a third text is filled near the text in the extended text image to obtain a third text image, wherein there is text overlap between the third text and the text in the extended text image; the third text image is cropped to obtain a cropped text image, wherein the size of the cropped text image is the same as the size of the text image; the cropped text image is enhanced to obtain a sample image, and the text image is enhanced to obtain a sample standard image; Alternatively, a fourth text image and a fifth text image are obtained, wherein there is no text overlap between the fourth text image and the fifth text image; a local text image block is captured from the fourth text image, and the local text image block is processed and pasted to the fifth text image to obtain a sixth text image with text overlap; the sixth text image is enhanced to obtain a sample image, and the fifth text image is enhanced to obtain a sample standard image.
8. The method according to claim 4, characterized in that The determining a first loss between the predicted image and the sample standard image comprises: Determining an absolute value loss and a perceptual loss between the predicted image and the sample standard image; A weighted sum is performed on the absolute value loss and the perceptual loss to obtain a first loss between the predicted image and the sample standard image.
9. The method according to claim 4, characterized in that: The generator is supervised for training using a first absolute value loss between a predicted image and a sample standard image and a loss function of a gradient between the predicted image and the sample standard image; Determining a first loss between the predicted image and the sample standard image includes: determining a first absolute value loss between the predicted image and the sample standard image, and an image gradient between the predicted image and the sample standard image; Determine a second absolute value loss corresponding to the image gradient, and perform weighted summation on the first absolute value loss and the second absolute value loss to obtain a first loss.
10. An image processing device, characterized in that: The device comprises: An image acquisition module, used to acquire an original image, detect text lines in the original image, and store them in a text line set; A text line traversal module, used for traversing the text lines in the text line set; A text removal module, configured to intercept a local image block corresponding to the text line from the original image when the text line overlaps with any other text line in the text line set, remove background text in the local image block, and obtain a target local image block; A text recognition module is used to obtain the valid text corresponding to the text line according to the target local image block.
11. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.