A method, apparatus, device, and medium for erasing text from an image.
By constructing a text erasure model and combining visual and textual features for cross-modal fusion, the problem of text erasure disrupting the continuity of geological information in existing technologies has been solved, achieving high-precision and semantically consistent text erasure results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG LAB
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-26
AI Technical Summary
Existing image restoration methods can easily disrupt the continuity of the original geological information in an image when erasing text annotations, resulting in inconsistencies between the restored area and the surrounding geological content, which affects the professional accuracy and usability of the drawings.
By constructing a text erasure model, text features and contextual semantic information in the image are extracted, and cross-modal feature fusion is performed by combining visual features to generate a target image that does not include the original text. The LaMa image inpainting framework and OCR model are used for iterative training, and the loss function is optimized to improve erasure accuracy.
It achieves high-quality text erasure while maintaining the professional semantic consistency of images, avoiding the generation of erroneous content that contradicts the surrounding geological elements, and improving the accuracy and quality of text erasure.
Smart Images

Figure CN121707878B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and medium for erasing text from an image. Background Technology
[0002] In fields such as geological surveys, mineral exploration, and engineering mapping, specialized drawings serve as crucial carriers for recording and transmitting geological information. They comprehensively express key geological elements such as stratigraphic structure, fault structures, and lithological distribution through symbols, colors, and text annotations. With the accelerating pace of geoscience informatization, digitizing these drawings and enabling the re-editing and reuse of information has become a key requirement for improving work efficiency and supporting comprehensive analysis. In practice, it is often necessary to erase text on the drawings for different purposes to generate clearer and more adaptable maps.
[0003] However, existing general image restoration methods mostly rely on local texture and color features of the image for content inference, lacking the ability to understand the semantics of professional elements in geological maps. When erasing text annotations, such methods easily disrupt the continuity of the original geological information in the map, leading to inconsistencies or semantic conflicts between the restored area and the surrounding lithology, structure, and other geological content, thus affecting the professional accuracy and usability of the map.
[0004] Therefore, how to achieve high-quality erasure and content restoration of text annotations while maintaining the consistency of professional semantics in images, and avoid generating erroneous restoration content with logical inconsistencies, is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, one aspect of this application provides a method for erasing text from an image, the method comprising:
[0006] Acquire the image to be processed and the pre-built text erasure model;
[0007] Text information is extracted from the image to be processed to obtain text features; the text information includes the original text in the image to be processed and the contextual semantic information between the original texts.
[0008] The region containing the original text in the image to be processed is marked as the first pixel value, and the non-text region is marked as the second pixel value to obtain a mask image;
[0009] The text erasure model is used to perform cross-modal feature fusion of the image to be processed and the text features to obtain fused features; and a target image excluding the original text is generated based on the mask image and the fused features.
[0010] Optionally, the text erasure model is a model based on the LaMa image restoration framework with an embedded feature fusion module; the LaMa image restoration framework includes an encoder and a decoder;
[0011] The encoder is used to extract visual features from the image to be processed; the feature fusion module is used to fuse the visual features and the text features to obtain the fused features; the decoder is used to reconstruct the geological content of the region where the original text is located based on the mask image and the fused features to obtain the target image.
[0012] Optionally, the feature fusion module includes a multi-layer attention mechanism; and at least a portion of the output of the cross-modal feature fusion module is fused with the feature processing path of the encoder.
[0013] Optionally, constructing the text erasure model includes the following steps:
[0014] Construct an initial text erasure model and a loss function for the initial text erasure model; the loss function includes at least one of reconstruction loss, perceptual loss, and adversarial loss;
[0015] Obtain training image pairs; the training image pairs include textless images and textless images obtained by adding specified text to the textless images;
[0016] With the goal of minimizing the loss function, the initial text erasure model is iteratively trained using the training image pairs to obtain the text erasure model.
[0017] Optionally, after generating the target image that does not include the original text, the process further includes:
[0018] The target image is used to extract text using an OCR model to obtain the first feature;
[0019] When it is determined that the target image includes recognizable text based on the first feature, a penalty signal is generated based on the first feature;
[0020] Based on the penalty signal, generate a consistency loss; add the consistency loss to the loss function; and return to the step of obtaining training image pairs and subsequent steps.
[0021] Optionally, generating a penalty signal based on the first feature includes:
[0022] Obtain a truth image of the user input, excluding the original text;
[0023] The OCR model is used to extract text from the ground truth image to obtain the second feature;
[0024] The penalty signal is generated based on the difference between the first feature and the second feature.
[0025] Optionally, the text erasure method for the image further includes:
[0026] A confidence level is determined for assessing the credibility of the first feature; the confidence level is positively correlated with the credibility level.
[0027] Determine whether the confidence level is greater than the confidence threshold;
[0028] If the value is greater than the target image, the target image is determined to contain recognizable text.
[0029] If the confidence threshold is not greater than the specified value, after outputting the prompt signal, if a verification signal input by the user is obtained, the confidence threshold is adjusted; the verification signal is used to characterize that the target image contains recognizable text.
[0030] Another aspect of this application provides a text erasing device for an image, the device comprising:
[0031] The target acquisition module is used to acquire the image to be processed and the pre-built text erasure model;
[0032] The text information extraction module is used to extract text information from the image to be processed to obtain text features; the text information includes the original text in the image to be processed and the contextual semantic information between the original text.
[0033] The mask image generation module is used to mark the area where the original text is located in the image to be processed as the first pixel value and the non-text area as the second pixel value to obtain a mask image;
[0034] The text erasure module is used to perform cross-modal feature fusion on the image to be processed and the text features through the text erasure model to obtain fused features; and to generate a target image that does not include the original text based on the mask image and the fused features.
[0035] Another aspect of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor executes the computer program to implement the steps of a text erasure method for the image.
[0036] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a text erasure method for the image.
[0037] The text erasure method, apparatus, device, and medium provided in this application have the following beneficial effects: when erasing text in an image using a text erasure model, it is possible not only to rely on the visual information of the context, but also to understand the professional meaning of the area where the erased text is located. That is, by fusing and analyzing visual and textual information, the accuracy and quality of text erasure are improved, and erroneous content that contradicts the surrounding image content elements is avoided. For example, for geological images, erroneous content that contradicts the surrounding geological elements (such as strata, faults, rock masses, etc.) is avoided. Attached Figure Description
[0038] Figure 1 A schematic flowchart illustrating a text erasure method for an image provided in an embodiment of this application;
[0039] Figure 2 A schematic diagram illustrating the principle of a text erasure method for an image provided in an embodiment of this application;
[0040] Figure 3 A schematic diagram illustrating the effect of a text erasure method for an image provided in an embodiment of this application;
[0041] Figure 4 This is a schematic diagram of the structure of a text erasure model provided in an embodiment of this application;
[0042] Figure 5 A schematic diagram illustrating the principle of a text erasure method for an image provided in another embodiment of this application;
[0043] Figure 6 A schematic diagram of the structure of an image text erasing device provided in an embodiment of this application;
[0044] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0045] The reference numerals in the attached figures are as follows: 30 is the encoder, 31 is the feature fusion module, 32 is the decoder, 33 is the OCR model, 60 is the target acquisition module, 61 is the text information extraction module, 62 is the mask image generation module, 63 is the text erasure module, 70 is the memory, 71 is the processor, 72 is the display screen, 73 is the input / output interface, 74 is the communication interface, 75 is the power supply, 76 is the communication bus, 701 is the computer program, 702 is the operating system, and 703 is the data. Detailed Implementation
[0046] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0047] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0048] Figure 1 This is a flowchart illustrating a text erasure method for an image provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes:
[0049] S10: Obtain the image to be processed and the pre-built text erasure model;
[0050] S11: Extract text information from the image to be processed to obtain text features; the text information includes the original text in the image to be processed and the contextual semantic information between the original texts.
[0051] In a specific embodiment of image text erasure, the image to be processed that needs to have its text erased is obtained. The image to be processed can be an image from any field, such as a geological image, and this application does not limit it.
[0052] Figure 2 This is a schematic diagram illustrating the principle of a text erasure method for an image provided in this application embodiment. Further, to clearly identify the text information included in the image to be processed, the text information is extracted from the image in step S11. In an optional embodiment, such as... Figure 2 As shown, a pre-trained Optical Character Recognition (OCR) model can be used to recognize text information in the image to be processed and obtain text features.
[0053] To ensure that the text erasure model can accurately erase the text in the image to be processed, in an optional embodiment, the text information includes not only the identified original text, i.e., the text content included in the image, but also the contextual semantic information between the original text corresponding to the context.
[0054] This ensures that the subsequent text erasure model, when erasing text, not only knows which text needs to be erased, but can also accurately reconstruct the image information of the original text's location by combining contextual semantic information. For example, when the image to be processed is a geological image, it can reconstruct the geological elements of the original text's location. Furthermore, it should be noted that, in an optional embodiment, the generated text features can be text-encoded features, including semantic embedding vectors or high-dimensional representations.
[0055] It is worth noting that the OCR model 33 is used to recognize the original text content and contextual semantic information in the image to be processed. In specific embodiments, the OCR model 33 can adopt the Transformer architecture or a mainstream OCR architecture model, which is not limited in this application. The text features obtained by the OCR model include not only the character recognition results (i.e., the original text), but also its contextual semantic information, which can be used to guide subsequent semantic consistency reconstruction.
[0056] S12: Mark the region containing the original text in the image to be processed as the first pixel value, and mark the non-text region as the second pixel value to obtain the mask image;
[0057] Furthermore, to further improve the text erasure accuracy of the text erasure model, while extracting text features, the image to be processed is converted into a binary mask image. Specifically, the region containing the original text to be erased in the image is marked as the first pixel value, and other non-text regions are marked as the second pixel value, thus obtaining the binary mask image. In specific embodiments, this application does not limit the specific values of the first and second pixel values, nor does it limit the method for generating the binary mask image.
[0058] S13: The text erasure model is used to perform cross-modal feature fusion of the image to be processed and the text features to obtain fused features; and the target image excluding the original text is generated based on the mask image and the fused features.
[0059] At this time, as Figure 2 As shown, the image to be processed is obtained, and after processing by the OCR model 33, text features and a mask image are obtained. All three images are then input into a pre-built text erasure model for processing. In a specific embodiment, the text erasure module performs cross-modal feature fusion of the visual features and text features provided by the image to be processed, obtaining fused features.
[0060] Furthermore, the mask image provides information such as the location of the original text in the image to be processed. Based on the fused features, the original text in the image to be processed is accurately erased to obtain the target image. In fact, the text erasure model realizes semantic feature fusion and semantic-aware image reconstruction, that is, after erasing the original text, the content elements (e.g., geological elements) of the area where the original text was located are reconstructed based on semantic awareness.
[0061] The text erasure method for images provided in this application includes, but is not limited to, the processing of professional drawings such as those for geology and surveying. It can solve the problem of traditional methods in repairing content with unreasonable semantics and visual abruptness on complex professional graphics, and achieves high-fidelity and high semantic consistency text erasure effect.
[0062] Figure 3 This is a schematic diagram illustrating the effect of a text erasure method for an image provided in an embodiment of this application. In an optional embodiment, for example... Figure 3 As shown, the image to be processed is a geological map. The original image includes text to be erased. After processing by the text erasure method provided in this application embodiment, the target image with the text erased on the right can be obtained.
[0063] In one optional embodiment, in order to further improve the text erasure effect, after obtaining the target image, the target image is optimized. The optimization process includes, but is not limited to, color smoothing, edge refinement and other optimization processes, so as to obtain the final image.
[0064] Therefore, the text erasure method for images provided in this application, when erasing text in an image using a text erasure model, can not only rely on the visual information of the context, but also understand the professional meaning of the area where the erased text is located. That is, it integrates and analyzes visual information and text information to improve the accuracy and quality of text erasure and avoids generating erroneous content that contradicts the surrounding image content elements. For example, for geological images, it avoids generating erroneous content that contradicts the surrounding geological elements (such as strata, faults, rock masses, etc.).
[0065] Figure 4 This is a schematic diagram of the structure of a text erasure model provided in an embodiment of this application. In an optional embodiment, such as... Figure 4 As shown, the text erasure model is based on the LaMa image restoration framework and incorporates a feature fusion module 31; the LaMa image restoration framework includes an encoder 30 and a decoder 32.
[0066] The encoder 30 is used to extract visual features from the image to be processed; the feature fusion module 31 is used to fuse visual features and text features to obtain fused features; the decoder 32 is used to reconstruct the geological content of the area where the original text is located based on the mask image and the fused features to obtain the target image.
[0067] In a specific embodiment, such as Figure 2 and Figure 4 The text erasure model is based on the LaMa (Large Mask Inpainting) image restoration framework, which includes encoder 30 and decoder 32. A cross-modal feature fusion module 31 is introduced in encoder 30 to fuse the text features of OCR model 33 with the visual features extracted by encoder 30.
[0068] For details, see Figure 2 and Figure 4 In a specific embodiment, the OCR model 33 extracts text features from the image to be processed and inputs the text features into the feature fusion module 31 in the text erasure model. Meanwhile, the encoder 30 extracts visual features from the image to be processed and transmits them to the feature fusion module 31 so that the feature fusion module 31 can generate fused features, which are visual features containing text semantic enhancement information.
[0069] Furthermore, the fused multimodal features are input into decoder 32, which generates content for the mask-covered area (i.e., decoder 32 reconstructs the original text-covered area). Since the visual features have been fused with textual semantic information, the model can "understand" the original geological meaning of the text during the repair process, thereby generating stratigraphic lines, fault symbols, or lithological fills consistent with the surrounding structure, achieving a high-quality erasing effect with semantic awareness.
[0070] Based on the above embodiments, as an optional embodiment, such as... Figure 4 As shown, the feature fusion module 31 includes a multi-layer attention mechanism; and at least a portion of the output of the cross-modal feature fusion module 31 is fused with the feature processing path of the encoder 30.
[0071] In a specific embodiment, to improve the text erasure accuracy of the text erasure model, the feature fusion module 31 is inserted into multiple scale levels of the encoder 30 to calculate the attention distribution of visual features to text features through a cross-attention mechanism. In an optional embodiment, the number of attention mechanism layers can be set to 3, but this application does not limit this.
[0072] In one optional embodiment, the input to the first layer of encoder 30 is the image to be processed, and the input data of the first-layer attention mechanism (cross-attention) of feature fusion module 31 includes the output of the first layer of encoder 30 and the text features output by OCR model 33. Except for the first layer of encoder 30, the input to other layers of encoder 30 is the fusion output of the attention mechanism of the previous layer with the same number of layers. Except for the first-layer cross-attention, the input to other layers of cross-attention is the output of encoder 30 with the same number of layers and the text features.
[0073] Therefore, based on Figure 4 As can be seen, the feature fusion module 31 is inserted into multiple different scale levels of the image encoder 30, so that the semantic information of the text is gradually fused in the image features from shallow to deep layers, and the semantic-visual full-scale association is realized.
[0074] In a specific embodiment, the visual feature map of the intermediate layer of encoder 30 is used as the Query, and the vector of text features is used as the Key and Value. The attention weight of the visual Query to the text Key-Value pair is calculated through a cross-attention mechanism, thereby generating visual features enhanced by text semantic information, i.e., generating fused features, and feeding the fused features back to the subsequent layers of encoder 30.
[0075] In one alternative embodiment, constructing a text erasure model includes the following steps:
[0076] Construct an initial text erasure model and its loss function; the loss function includes at least one of reconstruction loss, perceptual loss, and adversarial loss;
[0077] Obtain training image pairs; the training image pairs include images without text and images with text obtained by adding specified text to images without text.
[0078] With the goal of minimizing the loss function, the initial text erasure model is iteratively trained using training image pairs to obtain the text erasure model.
[0079] It is understandable that, in specific embodiments, in order to obtain a high-precision text erasure model, after introducing the feature fusion module 31 on the basis of the LaMa image restoration framework, it is also necessary to iteratively train the model with a large training dataset so that the model has high-precision text erasure capabilities.
[0080] Specifically, an initial text erasure model is pre-built. Simultaneously, training image pairs are constructed for training the model. These image pairs include images without text and images with text. The image with text is the image with specified text added to the original image without text. In other words, after adding specified text to the original image without text, the image without text and the processed image are used as an image pair.
[0081] Furthermore, the model's loss function is constructed using at least one of reconstruction loss, perceptual loss, and adversarial loss. During the model training phase, the initial text erasure model is iteratively supervised and trained with the goal of minimizing the loss function, thereby obtaining an initial text erasure model that can accurately erase text.
[0082] Figure 5 A schematic diagram illustrating the principle of a text erasure method for an image provided in another embodiment of this application, as shown below. Figure 5 As shown, based on the above embodiments, as an optional embodiment, after generating the target image that does not include the original text, the following steps are also included:
[0083] The text is extracted from the target image using OCR model 33 to obtain the first feature;
[0084] When it is determined that the target image contains recognizable text based on the first feature, a penalty signal is generated based on the first feature;
[0085] Based on the penalty signal, generate a consistency loss; add the consistency loss to the loss function; and return to execute the step of obtaining training image pairs and subsequent steps.
[0086] To further improve the accuracy of text erasure in images, after processing the target image using the text erasure model, the target image is input again into the OCR model 33 for text feature extraction to obtain the first feature. If the first feature indicates that the target image contains recognizable text, it means that the text erasure model has not completely erased the original text in the image to be processed, reflecting that the current text erasure model has a low erasure capability for such images.
[0087] At this time, as Figure 5 As shown, a penalty signal is generated based on the first feature, and a consistency loss is generated based on the penalty signal. It can be understood that the consistency loss is the consistency loss of the OCR model 33, which is used to constrain the model to optimize in the direction of not containing recognizable text, that is, to perform semantic consistency constraints on the erasure results of the model during the training phase.
[0088] Specifically, the consistency loss is added to the loss function during model training, and the steps of training the text erasure model using training images as described in the previous embodiment are executed again to adjust the parameters in the encoder 30, feature fusion module 31, and decoder 32 to optimize the model. At this point, fine-tuning training is performed with the goal of minimizing the loss function constructed by the model. This mechanism can prevent the model from leaving residual characters or generating text-like textures after erasure, thus ensuring complete semantic erasure.
[0089] In one alternative embodiment, generating a penalty signal based on a first feature includes:
[0090] Obtain a truth image of the user input, excluding the original text;
[0091] Text extraction is performed on the ground truth image using OCR model 33 to obtain the second feature;
[0092] A penalty signal is generated based on the difference between the first feature and the second feature.
[0093] In a specific embodiment, the consistency loss is used to characterize the degree of feature difference between the ground truth image and the target image after being processed again by the OCR model 33. Therefore, when generating the penalty signal, the ground truth image that does not include the original text input by the user is obtained. The ground truth image refers to the image that does not contain the original text for the image to be processed. It can be an image that has been manually erased by the user, or it can be the initial image that was originally intended to have the specified text added during the training phase. This application does not limit this.
[0094] In fact, the ground truth image is used to guide the remaining deviations in the image processed by OCR model 33. Therefore, OCR model 33 extracts a second feature from the ground truth image and generates a penalty signal based on the difference between the first and second features. In a specific embodiment, both the first and second features are text features, which can generate text feature vectors. Therefore, the difference between the first and second features can be represented by vector similarity. After generating the penalty signal, a consistency loss for OCR model 33 is constructed based on this penalty signal to minimize the semantic information perceived by OCR model 33 before and after erasure.
[0095] In one optional embodiment, the text erasure method for an image further includes:
[0096] Determine the confidence level used to assess the credibility of the first feature; the confidence level is positively correlated with the credibility level.
[0097] Determine whether the confidence level is greater than the confidence threshold;
[0098] If the value is greater than 1, the target image is determined to contain identifiable text.
[0099] If the value is not greater than the specified value, after outputting the prompt signal, if the user input verification signal is obtained, the confidence threshold is adjusted; the verification signal is used to characterize that the target image contains recognizable text.
[0100] In a specific embodiment, when the OCR model 33 processes the target image again, if it finds that there is still high-confidence text in the target image, it may mean that the text in the image to be processed has not been completely erased, or it may be a judgment error. This may lead to repeated training of the model or repeated erasing of the image, wasting computing resources, or it may also lead to the loss of images with good erasing effects.
[0101] To avoid the aforementioned technical problems and to improve the accuracy of text erasure, in an optional embodiment, after extracting the text from the target image using the OCR model 33 to obtain the first feature, the confidence level of the first feature is determined. The higher the confidence level, the more reliable the representation of the first feature.
[0102] If the confidence level is greater than the confidence threshold, the current judgment result is output as "high-confidence text still exists in the target image." If the confidence level is not greater than the confidence threshold, the current judgment result is output as "no high-confidence text exists in the target image," meaning the text in the image to be processed has been completely erased. At this point, the current judgment result needs to be verified, and a prompt signal should be output to the user for verification.
[0103] Furthermore, if the user verifies that the target text does not contain high-confidence text, the text erasure of the current image to be processed can be terminated. If, after verification, high-confidence text still exists in the target text, but the system concludes that high-confidence text does not exist, a conflict exists.
[0104] At this point, the confidence threshold may be set inappropriately. As a sustainable implementation, the confidence threshold can be adjusted based on the verification signal input by the user. The verification signal is used to characterize that there is still identifiable text in the current target image.
[0105] In one alternative embodiment, the confidence threshold can be lowered based on the verification signal, and reduced to a level less than the confidence level corresponding to the first feature. For ease of understanding, examples will be provided below.
[0106] For example, if the confidence level of the first feature is 0.89 and the confidence threshold is 0.9, it indicates that the target image does not contain identifiable text. However, after user verification, the input verification signal indicates that the target image contains identifiable text. This may be because the confidence threshold is not set appropriately. In this case, the confidence threshold can be appropriately reduced to less than the confidence level of the first feature, which is 0.89. For example, the confidence threshold can be set to 0.88.
[0107] In one optional embodiment, when the confidence level is greater than a confidence threshold, indicating that there is still high-confidence text in the target image, the text erasure model is triggered to perform a secondary erasure. Alternatively, a prompt signal can be sent to the terminal so that the user can verify and correct the mask image, thereby improving erasure accuracy.
[0108] Therefore, the text erasure method for images provided in this application improves the accuracy of text erasure by dynamically adjusting the confidence threshold, correcting the mask image, and performing secondary erasure on the target image.
[0109] In the above embodiments, the method for erasing text from images has been described in detail. This application also provides an embodiment of an image text erasing device.
[0110] Figure 6 This is a schematic diagram of the structure of an image text erasing device provided in an embodiment of this application, as shown below. Figure 6 As shown, the device includes:
[0111] The target acquisition module 60 is used to acquire the image to be processed and the pre-built text erasure model;
[0112] The text information extraction module 61 is used to extract text information from the image to be processed and obtain text features; the text information includes the original text in the image to be processed and the contextual semantic information between the original texts.
[0113] The mask image generation module 62 is used to mark the area where the original text is located in the image to be processed as the first pixel value and the non-text area as the second pixel value to obtain a mask image;
[0114] The text erasure module 63 is used to perform cross-modal feature fusion of the image to be processed and the text features through the text erasure model to obtain fused features; and to generate a target image that does not include the original text based on the mask image and the fused features.
[0115] Furthermore, the image text erasing device provided in this application embodiment also includes:
[0116] The building block is used to construct the initial text erasure model and its loss function; the loss function includes at least one of reconstruction loss, perceptual loss, and adversarial loss.
[0117] The training data acquisition module is used to acquire training image pairs; the training image pairs include images without text and images with text obtained by adding specified text to images without text.
[0118] The model training module is used to iteratively train the initial text erasure model by training image pairs with the goal of minimizing the loss function, so as to obtain the text erasure model.
[0119] The first feature extraction module is used to extract text from the target image using an OCR model to obtain the first feature;
[0120] The penalty signal generation module is used to generate a penalty signal based on the first feature when it is determined that the target image includes recognizable text based on the first feature.
[0121] The first processing module is used to generate a consistency loss based on the penalty signal; add the consistency loss to the loss function; and return to execute the step of obtaining training image pairs and subsequent steps.
[0122] The truth image acquisition module is used to acquire the truth image input by the user, excluding the original text;
[0123] The second feature extraction module is used to extract text from the ground truth image using an OCR model to obtain the second feature;
[0124] The second processing module is used to generate a penalty signal based on the difference between the first feature and the second feature.
[0125] The confidence level determination module is used to determine the confidence level used to evaluate the credibility of the first feature; the confidence level is positively correlated with the credibility level.
[0126] The third processing module is used to determine whether the confidence level is greater than the confidence level threshold. If it is greater, it determines that the target image contains identifiable text. If it is not greater, after outputting a prompt signal, if a verification signal input by the user is obtained, the confidence level threshold is adjusted. The verification signal is used to characterize that the target image contains identifiable text.
[0127] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 7 As shown, the electronic device includes: a memory 70 for storing computer programs;
[0128] The processor 71 is configured to execute a computer program to implement the steps of the image text erasure method as described in the above embodiments.
[0129] The electronic devices provided in this embodiment may include, but are not limited to, tablet computers, laptop computers, or desktop computers.
[0130] The processor 71 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 71 may be implemented using at least one of the following hardware forms: Digital Signal Processor (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 71 may also include a main processor and a coprocessor. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 71 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 71 may also include an Artificial Intelligence (AI) processor, which is used to handle computational operations related to machine learning.
[0131] The memory 70 may include one or more computer-readable storage media, which may be non-transitory. The memory 70 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 70 is used to store at least the following computer program 701, which, after being loaded and executed by the processor 71, is capable of implementing the relevant steps of the image text erasure method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 70 may also include an operating system 702 and data 703, etc., and the storage method may be temporary storage or permanent storage. The operating system 702 may include Windows, Unix, Linux, etc. The data 703 may include, but is not limited to, relevant data involved in the image text erasure method.
[0132] In some embodiments, the electronic device may further include a display screen 72, an input / output interface 73, a communication interface 74, a power supply 75, and a communication bus 76.
[0133] Those skilled in the art will understand that Figure 7 The structures shown do not constitute a limitation on electronic devices and may include more or fewer components than those shown.
[0134] The electronic device provided in this application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the text erasure method of the image in the above embodiments.
[0135] It should be noted that although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
Claims
1. A method for erasing text from an image, characterized in that, The method includes: Acquire the image to be processed and the pre-built text erasure model; Text information is extracted from the image to be processed to obtain text features; the text information includes the original text in the image to be processed and the contextual semantic information between the original text; the text features are semantic embedding vectors. The region containing the original text in the image to be processed is marked as the first pixel value, and the non-text region is marked as the second pixel value to obtain a mask image; The text erasure model is used to perform cross-modal feature fusion of the image to be processed and the text features to obtain fused features; and a target image excluding the original text is generated based on the mask image and the fused features. The text erasure model is based on the LaMa image restoration framework and incorporates a feature fusion module; the LaMa image restoration framework includes an encoder and a decoder; the feature fusion module is inserted into multiple layers of the encoder at different scales. The encoder is used to extract visual features from the image to be processed; the feature fusion module is used to fuse the visual features and the text features to obtain the fused features; the decoder is used to reconstruct the geological content of the region where the original text is located based on the mask image and the fused features to obtain the target image.
2. The text erasure method for an image as described in claim 1, characterized in that, The feature fusion module includes a multi-layer attention mechanism; and at least a portion of the output of the cross-modal feature fusion module is fused with the feature processing path of the encoder.
3. The text erasure method for an image as described in claim 1, characterized in that, Constructing the text erasure model includes the following steps: Construct an initial text erasure model and a loss function for the initial text erasure model; the loss function includes at least one of reconstruction loss, perceptual loss, and adversarial loss; Obtain training image pairs; the training image pairs include textless images and textless images obtained by adding specified text to the textless images; With the goal of minimizing the loss function, the initial text erasure model is iteratively trained using the training image pairs to obtain the text erasure model.
4. The method for erasing text from an image as described in claim 3, characterized in that, After generating the target image that does not include the original text, the process also includes: The target image is used to extract text using an OCR model to obtain the first feature; When it is determined that the target image includes recognizable text based on the first feature, a penalty signal is generated based on the first feature; Based on the penalty signal, generate a consistency loss; add the consistency loss to the loss function; and return to the step of obtaining training image pairs and subsequent steps.
5. The text erasure method for an image as described in claim 4, characterized in that, Generating a penalty signal based on the first feature includes: Obtain a truth image of the user input, excluding the original text; The OCR model is used to extract text from the ground truth image to obtain the second feature; The penalty signal is generated based on the difference between the first feature and the second feature.
6. The method for erasing text from an image as described in claim 4, characterized in that, The method further includes: A confidence level is determined for assessing the credibility of the first feature; the confidence level is positively correlated with the credibility level. Determine whether the confidence level is greater than the confidence threshold; If the value is greater than the target image, the target image is determined to contain recognizable text. If the confidence threshold is not greater than the specified value, after outputting the prompt signal, if a verification signal input by the user is obtained, the confidence threshold is adjusted; the verification signal is used to characterize that the target image contains recognizable text.
7. A text erasing device for an image, characterized in that, The device includes: The target acquisition module is used to acquire the image to be processed and the pre-built text erasure model; The text information extraction module is used to extract text information from the image to be processed to obtain text features; the text information includes the original text in the image to be processed and the contextual semantic information between the original text; the text features are semantic embedding vectors. The mask image generation module is used to mark the area where the original text is located in the image to be processed as the first pixel value and the non-text area as the second pixel value to obtain a mask image; The text erasure module is used to perform cross-modal feature fusion on the image to be processed and the text features through the text erasure model to obtain fused features; and to generate a target image that does not include the original text based on the mask image and the fused features. The text erasure model is based on the LaMa image restoration framework and incorporates a feature fusion module; the LaMa image restoration framework includes an encoder and a decoder; the feature fusion module is inserted into multiple layers of the encoder at different scales. The encoder is used to extract visual features from the image to be processed; the feature fusion module is used to fuse the visual features and the text features to obtain the fused features; the decoder is used to reconstruct the geological content of the region where the original text is located based on the mask image and the fused features to obtain the target image.
8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the text erasure method for the image according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the text erasure method for the image according to any one of claims 1 to 6.