Method and device for processing text in image, readable storage medium and program product

By extracting the target text structure and style features in the image, and using the diffusion model to generate target graphic text, the problem of low efficiency of graphic text editing in the prior art is solved, and automated and efficient graphic text editing is achieved.

CN120339462APending Publication Date: 2025-07-18XIAMEN MEITUZHIJIA TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510410251.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, graphic text editing is inefficient and depends on the user's image processing experience, so it is impossible to efficiently automate the processing of graphic text in the image.

Method used

By obtaining the target text area and content retention area in the initial image, extracting the structural features and text style features of the target text, using the diffusion model to generate the target graphic text, replacing the initial graphic text and retaining the image content of the content retention area.

Benefits of technology

Automatic editing of graphic text is realized, and the editing efficiency is improved. The generated target graphic text can accurately express the structure of the target text and retain the content style of the initial image, avoiding the inefficiency of manual editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339462A_ABST
    Figure CN120339462A_ABST
Patent Text Reader

Abstract

The invention relates to a method and device for processing a text in an image, a readable storage medium and a program product. The method comprises the following steps: acquiring an initial image, and determining a target text region and a content retention region in the initial image; the target text area comprises an initial graphic text of a target text style; obtaining a target text, and performing structural feature extraction based on an image formed by the target text to obtain structural features of the target text; performing text style feature extraction according to the target text area to obtain text style features of the initial graphic text; and performing image generation based on the initial image, the structural features and the text style features to replace an initial graphic text in the initial image with a target graphic text formed by the target text according to a target text style, and retaining the image content of the content retaining area to generate a target image. By adopting the method, the graphic text editing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and in particular, to a method, an apparatus, a readable storage medium, and a program product for processing text in an image. Background Art

[0002] With the development of computer technologies, the demand for image processing has become increasingly diverse. In some scenarios, there is a need to edit graphic text in a text area of an image. For example, in the field of graphic design, there is a need to edit some graphic text presented in graphic design drawings such as posters, leaflets, magazine covers, etc. Also, in the field of post-production of films and television, there is a need to edit graphic text in image areas such as sign image areas and license plate image areas displayed in video frame images of film and television works. In response to this demand for graphic text editing, in traditional methods, image processing applications are usually used to manually perform the editing. Specifically, the initial graphic text in the image can be first filled with blanks, and then new graphic text can be designed according to the style of the initial graphic text, and the new graphic text is overlaid on the image.

[0003] However, the above-mentioned method of manually performing editing using an image processing application depends on the user's image processing experience and has the problem of low efficiency in graphic text editing. Summary of the Invention

[0004] Based on this, the present application provides a method, an apparatus, a readable storage medium, and a program product for processing text in an image, which can improve the efficiency of graphic text editing.

[0005] On the one hand, the present application provides a method for processing text in an image, including:

[0006] Obtain an initial image, and determine a target text area and a content retention area in the initial image; the target text area contains initial graphic text of a target text style;

[0007] Obtain a target text, extract structural features of the target text based on an image formed by the target text, and obtain the structural features of the target text;

[0008] Extract text style features of the initial graphic text according to the target text area, and obtain the text style features of the initial graphic text;

[0009] Perform image generation based on the initial image, the structural features, and the text style features, so as to replace the initial graphic text in the initial image with target graphic text formed by the target text according to the target text style, and retain the image content of the content retention area to generate a target image.

[0010] On the one hand, the present application also provides a device for processing text in an image, including:

[0011] An acquisition module, configured to acquire an initial image and determine a target text region and a content retention region in the initial image; the target text region contains initial graphic text in a target text style;

[0012] A first feature extraction module, configured to acquire target text and perform structural feature extraction on an image formed based on the target text to obtain structural features of the target text;

[0013] A second feature extraction module, configured to perform text style feature extraction according to the target text region to obtain text style features of the initial graphic text;

[0014] An image generation module, configured to perform image generation based on the initial image, the structural features, and the text style features, so as to replace the initial graphic text in the initial image with target graphic text formed by the target text in the target text style, and retain the image content of the content retention region to generate a target image.

[0015] On the one hand, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0016] Acquire an initial image and determine a target text region and a content retention region in the initial image; the target text region contains initial graphic text in a target text style;

[0017] Acquire target text and perform structural feature extraction on an image formed based on the target text to obtain structural features of the target text;

[0018] Perform text style feature extraction according to the target text region to obtain text style features of the initial graphic text;

[0019] Perform image generation based on the initial image, the structural features, and the text style features, so as to replace the initial graphic text in the initial image with target graphic text formed by the target text in the target text style, and retain the image content of the content retention region to generate a target image.

[0020] On the one hand, the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0021] Acquire an initial image and determine a target text region and a content retention region in the initial image; the target text region contains initial graphic text in a target text style;

[0022] Obtain the target text, perform structural feature extraction on the image formed based on the target text, and obtain the structural features of the target text;

[0023] Extract text style features according to the target text area, and obtain the text style features of the initial graphic text;

[0024] Perform image generation based on the initial image, the structural features, and the text style features, so as to replace the initial graphic text in the initial image with the target graphic text formed by the target text according to the target text style, and retain the image content of the content retention area to generate a target image.

[0025] In the above method, device, readable storage medium, and program product for processing text in an image, the initial graphic text with the target text style included in the target text area of the initial image is the graphic text to be edited. Since the structural features are obtained by performing structural feature extraction on the image formed based on the target text, the structural features can represent the structural information on which the accurate recognition of the target text depends. The text style features are obtained by performing text style feature extraction based on the target text area, and the text style features can represent the target text style of the initial graphic text. Then, based on the initial image, the structural features, and the text style features, image generation is performed, and the target graphic text in the generated target image can not only accurately express the structure of the target text, so that the target graphic text can be accurately recognized visually as the target text, but also have the target text style of the initial graphic text. Moreover, the target image retains the content retention area of the initial image. In this way, without manual editing of the graphic text, it can be automatically realized that the initial graphic text in the initial image is edited into a target graphic text with accurate structure and retaining the target text style, thereby improving the efficiency of graphic text editing. Brief Description of the Drawings

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0027] Figure 1 It is a schematic flowchart of a method for processing text in an image in an embodiment;

[0028] Figure 2 It is a schematic flowchart of the brief steps of a method for processing text in an image in an embodiment;

[0029] Figure 3Schematic diagram of the architecture of the algorithm processing module in an embodiment;

[0030] Figure 4 Schematic diagram of the process flow of the structural feature extraction step in an embodiment;

[0031] Figure 5 Schematic diagram of the process flow of the text style feature extraction step in an embodiment;

[0032] Figure 6 Schematic diagram of the architecture of the diffusion model in an embodiment;

[0033] Figure 7 Structural block diagram of the device for processing text in an image in an embodiment;

[0034] Figure 8 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0035] In order to make the objectives, technical solutions and beneficial effects of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0036] In one embodiment, as Figure 1 shown, a method for processing text in an image is provided. In this embodiment, an example is given where this method is applied to a computer device. The computer device can be a terminal or a server. It can be understood that this method can also be applied to a system including a terminal and a server and be implemented through the interaction between the terminal and the server. The terminal can be a personal computer, a laptop, a smartphone, a tablet computer, or others. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the method includes the following steps 102 to step 108:

[0037] Step 102, obtain an initial image, and determine the target text area and the content retention area in the initial image; the target text area contains the initial graphic text of the target text style.

[0038] Among them, the initial image is the image to be edited for graphic text. The initial image can be the input image passed in by the user, or the image obtained after image preprocessing of the input image. The initial image can be an image of various image types with graphic text. For example, various image types can include scene images, graphic design drawings, interface screenshots, book cover images, or others. A scene image is an image that captures a real-world environment. A scene image can contain rich object regions, background regions, text regions, or others, and the information in a scene image is usually complex. The text region in a scene image, for example, can be an image region formed by a billboard, a license plate, a store signboard, or other objects presenting text.

[0039] The target text region is the image region where the initial graphic text to be edited in the initial image is located. The target text region can be the region formed by the irregular contour of the whole initial graphic text, and this region can be the smallest region surrounding the whole initial graphic text. The target text region can also be a region of a regular shape surrounding the initial graphic text. Regular shapes such as circles and rectangles. The content retention region is the image region outside the target text region in the initial image.

[0040] The initial graphic text is the graphic text to be edited. The initial graphic text can be all the graphic text contained in the initial image, or a part of all the graphic text contained in the initial image. Graphic text is text presented in the form of an image. Graphic text is data of the image type. The target text style is the text style of the initial graphic text. The text style can include the visual style of the graphic text and can also include the position of the graphic text in the image. The visual style can include at least one of font, color, font size, transparency, spacing between characters, whether there is an underline, or other text styles. The font can include the shape of the characters, the line size of the characters, and the inclination of the characters.

[0041] Exemplarily, the computer device can obtain the input image, perform image preprocessing on the input image to obtain the initial image, determine the selected target text region in response to the region selection operation on the initial image, and determine the image region outside the target text region in the initial image as the content retention region.

[0042] Among them, the input image is an image uploaded by the user, which can be obtained by uploading an existing image or taking a real-time photo. Image preprocessing may include at least one of text existence detection, image content review, and resolution adjustment. Image preprocessing may also be to perform text existence detection, image content review, and resolution adjustment in sequence. Text existence detection is a process of detecting whether graphic text exists in the input image. Text existence detection can be achieved through a text detection model. When it is detected that there is no text in the input image, the user can be prompted to re-enter. Image content review is a process of performing security review on the image content of the input image. The security review may be to review whether the image content includes sensitive information, inappropriate information, or other information not suitable for display. Resolution adjustment is a process of adjusting the resolution of the input image. For example, adjusting the resolution of the input image to 1024*1024.

[0043] The region selection operation can be a single operation, such as a smearing operation, a lasso operation, a bounding box operation, or others. The region selection operation can also be a combined operation, such as a combination formed by a smearing operation and an erasing operation. The computer device can use the region selected by the region selection operation as the selected target text region. The computer device can also use the region formed after fine-tuning the edge of the region selected by the region selection operation as the selected target text region. For example, the edge of the region selected by the lasso operation or the bounding box operation can be adjusted to follow the contour surrounding the entire graphic text in the selected region, and the region surrounded by this contour is used as the target text region.

[0044] Step 104, obtain the target text, extract the structural features from the image formed based on the target text, and obtain the structural features of the target text.

[0045] Among them, the target text is used to replace the text represented by the initial graphic text. The structural features of the target text can characterize the structural information on which the accurate recognition of the target text depends. The structural features of the target text are used to distinguish the target text from other texts different from the target text. The structural features of the target text may not be affected by the text style. That is to say, even if the target text is formed into different graphic texts according to different text styles, the structural features of the target text can make these different graphic texts still be recognized as being formed by the target text.

[0046] Exemplarily, the computer device can obtain the input text, determine the target text based on the input text, and use the trained text recognition model to extract the structural features from the image formed by the target text to obtain the structural features of the target text.

[0047] Among them, the computer device can determine the input text as the target text, or perform text preprocessing on the input text to obtain the target text. The text preprocessing may include natural language processing for improving the semantic quality of the text, character filtering processing for improving the visual quality, or others. Natural language processing can include, for example, at least one of traditional and simplified Chinese conversion, synonym replacement, spelling correction, etc. Character filtering processing can include, for example, at least one of special character filtering, Emoji filtering, etc. Character filtering processing can be used to ensure that the generated graphic text is visually compatible with the target text style. The text recognition model can include the text recognition part in the OCR (Optical Character Recognition) model, the text recognition part in the multi-modal large model (such as GPT-4 with Vision provided by OpenAI, Gemini Pro provided by Google), or others.

[0048] In one embodiment, the computer device can determine the length of the target text that the target text area can accommodate based on the size of the target text area and the initial graphic text, and obtain the input text when the maximum text length that can be input is controlled to be the length of the target text. Among them, the text length can be the width occupied by the text in the character writing direction, or the number of characters in a specified language. The character writing direction is, for example, from left to right, from top to bottom, or others. The specified language is, for example, Chinese, English, Japanese, Korean, or others. It can be understood that under the same-sized area, for different languages, the number of characters that can be accommodated is different. For example, in an area that can accommodate one Chinese character, two English characters can be accommodated.

[0049] In one embodiment, the computer device can determine the initial language of the initial graphic text, determine the average character size of the initial graphic text, determine the number of characters that the target text area can accommodate in the initial language according to the size of the target text area and the average character size, determine the number of characters that the target text area can accommodate in a preset language different from the initial language based on the number of characters that the target text area can accommodate in the initial language, and determine the number of characters corresponding to the initial language and the preset language respectively as the target text length. Among them, when the initial language is Chinese, the preset language can be, for example, English, Japanese, Korean, or others.

[0050] Step 106, extract text style features according to the target text area to obtain the text style features of the initial graphic text.

[0051] Among them, the text style features can be used to characterize the target text style of the initial graphic text. The text style features can include the texture features, color features, spatial layout features between characters, position features relative to the initial image, or others of the initial graphic text.

[0052] Exemplarily, the computer device can obtain a first mask image that matches the image size of the initial image, and perform feature extraction on the initial image and the first mask image through a trained Convolutional Neural Networks (CNN) to obtain the text style features of the initial graphic text.

[0053] Among them, in the first mask image, each pixel in the image area corresponding to the position of the target text area is a non-zero pixel, and each pixel in the image area corresponding to the position of the content retention area is a zero pixel. The non-zero pixel can specifically be a pixel with a preset value. The preset value is, for example, 1, 255, or others. The non-zero pixel can be a pixel with a white color. The zero pixel can represent a black pixel. For the image area corresponding to the position of the target text area in the first mask image, it can be understood that the relative position of the image area corresponding to the position of the target text area in the first mask image with respect to the first mask image is the same as the relative position of the target text area with respect to the initial image. The meaning of the term "position corresponding" mentioned in other places in this application is similar to the meaning of the term "position corresponding" in the "image area corresponding to the position of the target text area in the first mask image" above, and will not be elaborated further hereinafter.

[0054] In one embodiment, the target text area can be a rectangular area. The computer device can expand from the target text area to the surrounding in the initial image, determine an expanded image block that includes the target text area after expansion, crop the expanded image block from the initial image to obtain a local image, and perform text style feature extraction based on the local image to obtain the text style features of the initial graphic text.

[0055] Among them, the size of the expanded image block can be a preset multiple of the size of the target text area. The preset multiple can be greater than 1. For example, the preset multiple can be 1.5, 2, or others. By cropping out the local image and performing text style feature extraction on the local image, for the scenario where the size of the target text area is much smaller than the size of the initial image, the accuracy of the extracted text style features can be improved. The size of the target text area being much smaller than the size of the initial image can be that the ratio value of the size of the target text area to the size of the initial image is less than a preset ratio value. The preset ratio value is, for example, 10%, 5%, or others.

[0056] Step 108, perform image generation based on the initial image, glyph features, and text style features to replace the initial graphic text in the initial image with a target graphic text formed by the target text according to the target text style, and retain the image content of the content retention area to generate a target image.

[0057] Exemplarily, the computer device can obtain a first mask image that matches the image size of the initial image, and based on the initial image, the first mask image, the glyph features, and the text style features, perform image generation through a trained image generation model to replace the initial graphic text in the initial image with the target graphic text formed by the target text in the target text style, and retain the image content in the content retention area to generate the target image.

[0058] Among them, when performing image generation through the image generation model, the initial image and the first mask image can be used as the input of the image generation model, and the glyph features and the text style features can be used as the conditional guidance information of the image generation model. The image generation model can specifically be a diffusion model. The diffusion model can be, for example, the image-to-image architecture in DDIM (Denoising Diffusion Implicit Models), Stable Diffusion (Steady State Diffusion Model), or others.

[0059] In one embodiment, the target text area can be a rectangular area. The computer device can expand from the target text area to the surrounding areas in the initial image to determine an extended image block that includes the target text area after expansion, and crop the extended image block from the initial image to obtain a local image; input the local image into the trained image generation model, and use the glyph features and the text style features as the conditional guidance information of the image generation model, and perform image generation through the image generation model to replace the initial graphic text in the target text area of the local image with the target graphic text formed by the target text in the target text style to generate a local intermediate image; fuse the target text area in the local intermediate image to the position of the target text area in the initial image to obtain the target image.

[0060] In the above method for processing text in an image, the initial graphic text of the target text style included in the target text area of the initial image is the graphic text to be edited. Since the structural features are obtained by extracting structural features from the image formed based on the target text, the structural features can represent the structural information on which the accurate recognition of the target text depends. The text style features are obtained by extracting text style features based on the target text area, and the text style features can represent the target text style of the initial graphic text. Then, based on the initial image, the structural features, and the text style features, an image is generated. The target graphic text in the generated target image can not only accurately express the structure of the target text, so that the target graphic text can be visually recognized as the target text accurately, but also have the target text style of the initial graphic text. Moreover, the target image retains the content retention area of the initial image. In this way, without manual editing of the graphic text, it can automatically realize the editing of the initial graphic text in the initial image into a target graphic text with accurate structure and retaining the target text style, thereby improving the efficiency of graphic text editing.

[0061] In an exemplary embodiment, the target text includes at least two characters, and step 104 may include: respectively determining character images formed by each of the at least two characters; respectively extracting features from the character images formed by the at least two characters to obtain character structure features of each of the at least two characters; and determining the structural features of the target text based on the local structural features obtained by splicing the character structure features of each of the at least two characters.

[0062] Among them, the characters may include symbols and characters. The characters may include numbers, Chinese characters, kana, letters, or others. The character image may be an image formed by the character according to a preset text style. The preset text style may be, for example, the font is equal-width font and the character color is white. Specifically, the character image may be an image formed by drawing the character into a black blank image according to the preset text style. For example, if the target text is "AGI will come tomorrow,", then the target text includes 8 characters, and the character image sequence of the target text may include the character images formed by each character in the target text (8 character images), and these character images are arranged according to the sorting positions of the characters in the target text.

[0063] Character structural features may include structural features of different levels extracted based on character images, such as low-level contour features and texture features, mid-level stroke topological features and spatial relationship features, and high-level overall shape structural features and semantic features. Stroke topological features may characterize the connection method of strokes in a character, the order of strokes, the geometric relationship between strokes, or other features. Spatial relationship features may characterize the relative position, proportion, arrangement, or other features between the components within a character. Components may be divisible components within a character, for example, the components of “八” may include “丿” and “乀”. Overall shape structural features may characterize the abstract structure of a character in terms of shape. For example, the shape structure of “A” is a triangular structure.

[0064] Exemplarily, the computer device may respectively determine the character image formed by each of the at least two characters contained in the target text, use a pre-trained OCR model to perform feature extraction on the character images formed by the at least two characters respectively, obtain the character structural features of the at least two characters respectively, and splice the character structural features of the at least two characters respectively into local structural features according to the character sorting positions of the at least two characters respectively in the target text, and determine the local structural features as the structural features of the target text.

[0065] The OCR model can be PaddleOCR-v4 (the fourth generation of the open source OCR toolkit developed by Baidu based on the Paddle deep learning framework), RapidOCR (an OCR model developed by RapidAI), or others. The character structure feature can be an intermediate feature vector extracted by the feature extractor of the OCR model based on the character image. The feature extractor of the OCR model can be a text recognition model of the OCR model, for example, SVTR (SwinTransformer for Text Recognition, a text recognition model based on deep learning) in PaddleOCR-v4.

[0066] In this embodiment, by determining the character image formed by each character in the target text and performing feature extraction on each character image, the character structural features of each character in the target text can be extracted. The character structural features can represent the information on which the character depends for accurate recognition, thereby determining the structural features of the target text based on the local structural features obtained by splicing the character structural features of each character in the target text. In this way, the structural features of the target text can accurately represent the information on which each character in the target text depends for accurate recognition, creating conditions for the subsequent generation of accurate target graphic text.

[0067] In an exemplary embodiment, the steps of determining the structural feature of the target text based on the local structural features spliced from the respective character structural features of at least two characters may include: determining a text image formed by the target text in the form of a single-line text; extracting features from the text image to obtain global structural features; and performing feature interaction based on the local structural features and the global structural features to obtain the structural feature of the target text.

[0068] Among them, in the text image formed by the target text, there may be a graphic text formed by the target text in the form of a single-line text according to a preset text style. The global structural features may include the respective character structural features of each character extracted from the text image, and may also include text-level features. The text-level features may characterize the spacing between characters, the aspect ratio of the text line width and height, the order of each character in the text, the semantics of the text, the segmentation method of the characters in the text, or others.

[0069] It can be understood that since the character image only presents a single character, while the text image can present multiple characters, there may be differences between the character structural features extracted from the character image and the character structural features extracted from the text image. The character structural features extracted from the text image can also focus on the structural relationship between characters, the semantics formed by adjacent characters, etc.

[0070] Performing feature interaction based on the local structural features and the global structural features can be achieved by using the attention mechanism to interact the respective information of the local structural features and the global structural features. For example, the local structural features and the global structural features can be input into a Transformer model (a neural network model based on the attention mechanism proposed by Google in 2017) for feature interaction. Specifically, the Transformer model may include an attention module and a Multilayer Perceptron (MLP). The attention module can process the local structural features and the global structural features to obtain attention features, and then the MLP can map the attention features to the structural feature of the target text. The attention module can adopt a cross-attention mechanism, a multi-head cross-attention mechanism, or a combination of these two attention mechanisms. The structural feature of the target text may be in the form of a feature vector.

[0071] Exemplarily, the computer device can determine a text image formed by the target text in the form of a single-line text, use a pre-trained OCR model to extract features from the text image to obtain global structural features, and perform feature interaction on the local structural features and the global structural features through the Transformer model to obtain the structural feature of the target text. Among them, the global structural features may be intermediate feature vectors extracted by the feature extractor of the OCR model based on the text image.

[0072] In this embodiment, a text image formed by the target text in the form of a single-line text is determined, and then feature extraction is performed on the text image to obtain global structural features. The global structural features can represent the structural features of characters from a global perspective and can also supplement the structural features of the overall text. Furthermore, the local structural features and the global structural features are subjected to feature interaction to obtain the structural features of the target text, which can improve the accuracy of the structural features of the target text.

[0073] In an exemplary embodiment, step 106 may include: obtaining a first mask image that matches the image size of the initial image. In the first mask image, each pixel in the image region corresponding to the target text region is a non-zero pixel, and each pixel in the image region corresponding to the content retention region is a zero pixel; performing feature extraction based on the initial image and the first mask image to obtain the text visual features of the initial graphic text; and determining the text style features of the initial graphic text based on the text visual features.

[0074] Among them, the non-zero pixel may specifically be a pixel with a preset value. The preset value is, for example, 1, 255, or others. The non-zero pixel may be a pixel with a white color. The zero pixel may represent a black pixel. The text visual features can be used to characterize the visual style of the initial graphic text. The visual style may include at least one of font, color, font size, transparency, spacing between characters, whether there is an underline, or other text styles.

[0075] Exemplarily, the computer device may obtain a first mask image that matches the image size of the initial image, and perform feature extraction on the initial image and the first mask image through a trained visual encoder to obtain the text visual features of the initial graphic text, and determine the text visual features as the text style features of the initial graphic text. Among them, the visual encoder may adopt a convolutional neural network (Convolutional Neural Networks, abbreviated as CNN), a ViT model (Vision Transformer, a model that applies the Transformer architecture to image processing tasks), or others.

[0076] In this embodiment, since each pixel in the image region corresponding to the target text region in the first mask image is a non-zero pixel, and each pixel in the image region corresponding to the content retention region is a zero pixel, feature extraction based on the initial image and the first mask image can focus on extracting the text visual features of the initial graphic text in the target text region of the initial image while ignoring the information in the content retention region, thereby improving the accuracy of extracting the text visual features of the initial graphic text.

[0077] In an exemplary embodiment, the steps of determining the text style features of the initial graphic text based on the text visual features may include: extracting features from the first masked image to obtain text position features, where the text position features are used to characterize the relative position of the target text region with respect to the initial image; and performing feature interaction based on the text visual features and the text position features to obtain the text style features of the initial graphic text.

[0078] Among them, performing feature interaction based on the text visual features and the text position features can be achieved by using a self-attention mechanism to interact the information of the text visual features and the text position features respectively. For example, the text visual features and the text position features can be input into a Transformer model for feature interaction. Specifically, the Transformer model may include a self-attention module and a multi-layer perceptron. The self-attention module can process the text visual features and the text position features to obtain self-attention features, and then the multi-layer perceptron can map the self-attention features to the text style features of the initial graphic text. The text style features may be in the form of feature vectors.

[0079] Exemplarily, the computer device can extract features from the first masked image through a position encoder to obtain text position features, and perform feature interaction on the text visual features and the text position features through a Transformer model to obtain the text style features of the initial graphic text.

[0080] Among them, the position encoder can be implemented based on a spatial attention mechanism. For example, the position encoder can be a Convolutional Block Attention Module (CBAM for short). The position encoder can also be implemented based on position encoding technology. For example, the position encoder is a component in a ViT model that implements the position encoding function, or a Position Encoding Generator (PEG for short).

[0081] In this embodiment, since the text position features are obtained by extracting features from the first masked image, the text position features can characterize the relative position of the target text region with respect to the initial image. Furthermore, by performing feature interaction based on the text visual features and the text position features, the text style features of the initial graphic text obtained contain the information after the interaction of the text visual features and the text position features, improving the richness of the text style features.

[0082] In an exemplary embodiment, step 108 may include: downsampling the initial image to obtain a downsampled image; the downsampled image includes a downsampled graphic text formed by downsampling the initial graphic text in the initial image; based on the downsampled image, the structural features, and the text style features, performing image generation through a trained image generation model, replacing the downsampled graphic text in the downsampled image with an intermediate graphic text formed according to the target text style of the target text to generate an intermediate image; upsampling the intermediate image to the image size of the initial image to obtain an upsampled image; the image area in the upsampled image corresponding to the target text area includes a target graphic text formed by upsampling the intermediate graphic text; fusing the image area in the upsampled image corresponding to the target text area into the target text area position in the initial image to obtain a target image.

[0083] Among them, downsampling is a process of reducing the image size of the initial image to reduce the resolution of the initial image. The image size of the downsampled image may be the same as the image size supported by the image generation model. For example, the image size of the downsampled image may be 64*64. The trained image generation model may be a pre-trained image generation model or a model obtained by fine-tuning training on the basis of a pre-trained image generation model. The image generation model may adopt, for example, the image-to-image architecture in DDIM (Denoising Diffusion Implicit Models), StableDiffusion (Steady-State Diffusion Model), or others.

[0084] It can be understood that when the initial image is downsampled to obtain a downsampled image, the target text area in the initial image is also downsampled to a downsampled text area, so that the initial graphic text in the initial image is also downsampled to a downsampled graphic text in the downsampled text area. The relative size of the downsampled text area with respect to the target text area may be the same as the relative size of the downsampled image with respect to the initial image.

[0085] The intermediate graphic text may be generated according to the target text style of the target text in the downsampled text area of the downsampled graphic text. The intermediate image includes the intermediate graphic text. The image size of the intermediate image may be the same as the image size of the downsampled image. Upsampling is a process of enlarging the image size of the intermediate image. It can be understood that when the intermediate image is upsampled to obtain an upsampled image, the intermediate graphic text in the intermediate image is also upsampled to a target graphic text in the upsampled image.

[0086] Exemplarily, the computer device can downsample the initial image to obtain a downsampled image; acquire a second mask image that matches the image size of the downsampled image; generate an intermediate image through a trained image generation model based on the second mask image, the downsampled image, the structural features, and the text style features; upsample the intermediate image to the image size of the initial image to obtain an upsampled image; and use a pre-configured image fusion algorithm to fuse the image region corresponding to the target text region position in the upsampled image into the target text region position in the initial image to obtain a target image.

[0087] Among them, in the second mask image, the pixels in the image region corresponding to the target text region position are non-zero pixels, and the pixels in the image region corresponding to the content retention region position are zero pixels. When generating an image through a trained image generation model, the downsampled image and the second mask image can be used as the inputs of the image generation model, and the glyph features and the text style features can be used as the conditional guidance information of the image generation model.

[0088] The image fusion algorithm can be Poisson Image Editing, Alpha fusion, or others. The core idea of Poisson Image Editing is to maintain the gradient information of the source image while making the boundary of the fusion region smoothly transition with the target image, so as to achieve seamless image fusion. In the obtained target image, the image region corresponding to the target text region position can contain the image information of the image region corresponding to the target text region position in the upsampled image, and in the target image, the image region corresponding to the content retention region position is consistent with the content of the content retention region.

[0089] The core formula of Poisson Image Editing can be expressed as: , satisfying . Among them, can represent the target image to be solved, can represent the image region corresponding to the target text region position in the upsampled image (which can be denoted as the source image), can represent the initial image, can represent the fusion region (that is, the image region corresponding to the target text region in the target image), can represent the boundary of the fusion region. The above core formula can be equivalent to the following Poisson equation: . Among them, can represent the Laplace operator, which represents the sum of second-order partial derivatives. In a two-dimensional image, the Laplace operator describes the local curvature of each pixel point in the image. can represent the divergence operator, which describes the divergence degree of the gradient field. can represent 's gradient field.

[0090] In the discrete domain, the Laplace operator can be approximated using the finite difference method. Therefore, the above Poisson equation can be expressed as a discrete-domain Poisson equation: The discrete Poisson equation applicable to each pixel in the fusion region can be: For each pixel on the boundary of the fusion region, the pixel value of the pixel at the same position in the initial image can be used, expressed as: .

[0091] By solving the discrete-domain Poisson equation, the pixel values in the fusion region can be determined, thereby obtaining the target image. The solution process of the discrete-domain Poisson equation can include: First, construct a linear equation system: For each pixel in the fusion region, there is a discrete Poisson equation applicable to each pixel in the above fusion region, and these discrete Poisson equations form a large sparse linear equation system; for the boundary pixels ( the pixels on), directly use the pixel values of the pixels at the same position in the initial image as known quantities; Second, solve the equation system: An iterative method (such as the Jacobi iterative method (Jacobi iteration method), Gauss-Seidel iterative method (Gauss-Seidel iteration method)) or a direct method (such as the conjugate gradient method) can be used to solve this large sparse linear equation system; Finally, result processing: The obtained values can be clipped and quantized to ensure that they are within the effective pixel value range (which can be 0 to 255).

[0092] When performing image Poisson fusion, some extended processing can also be carried out, including: 1. Blended gradient: A more natural transition effect can be created by interpolating between the gradients of the source image and the initial image. 2. Local color adjustment: Local color adjustment can be added during the fusion process to better match the colors of the source image and the initial image. 3. Multi-resolution analysis: By performing fusion at multiple resolution levels, large-scale and small-scale details can be better processed.

[0093] In this embodiment, by downsampling the initial image to obtain a downsampled image, the data volume of the image can be reduced, thereby improving the efficiency of the image generation model for image generation; generating an intermediate image through the image generation model, the intermediate image can include the replaced intermediate graphic text, and the intermediate graphic text is formed according to the target text style of the target text. Thus, in the upsampled image obtained by upsampling the intermediate image to the image size of the initial image, the image region corresponding to the target text region can include the target graphic text, and fusing the image region containing the target graphic text to the position of the target text region in the initial image can not only include the generated target graphic text but also retain the original resolution and detail texture of the content retention region.

[0094] In an exemplary embodiment, based on the downsampled image, the structural features, and the text style features, image generation is performed through a trained image generation model to replace the downsampled graphic text in the downsampled image with an intermediate graphic text formed by the target text according to the target text style. The steps of generating the intermediate image may include: obtaining a second mask image matching the image size of the downsampled image, in which the pixels in the image region corresponding to the target text region in the second mask image are non-zero pixels, and the pixels in the image region corresponding to the content retention region are zero pixels; performing noise addition processing on the combined image formed by the downsampled image and the second mask image to obtain a noisy feature; performing iterative denoising processing based on the noisy feature according to the structural features and the text style features to replace the downsampled graphic text in the downsampled image with an intermediate graphic text formed by the target text according to the target text style, thereby generating an intermediate image.

[0095] Among them, the image generation model may include an image encoder, an image information creator, and an image decoder. The image information creator may include a noise prediction network and a scheduler. The image generation model may generate image data based on a forward diffusion process and a reverse diffusion process. The forward diffusion process starts from a clear image without added noise and gradually adds noise, causing the clear image to gradually transform into a noise distribution. The reverse diffusion process starts from the noise distribution and gradually restores the original clear image or generates a new image. During the reverse diffusion process, the image generation model learns how to restore the original clear image or generate a new image from the noise data, which can be achieved by training the image generation model.

[0096] The image encoder may be an encoder in TAESD (Tiny AutoEncoder for Stable Diffusion), VAE (Variational Autoencoder), or others. The image decoder may be a Consistency Decoder, a decoder in VAE (Variational Autoencoder), or others. The image information creator may perform iterative denoising processing through the noise prediction network and the scheduler to generate an image. The noise prediction network may be a UNet network (a network including an encoder and a decoder with skip connections between the encoder and the decoder). The scheduler may include a computational process for adding noise or removing noise. For example, the scheduler may adopt the noise addition formula and the denoising formula in the DDIM model, or adopt the Scheduler component in the Stable Diffusion model.

[0097] Adding noise to the downsampled image can be achieved based on the forward diffusion process; specifically, the downsampled image can be input into the image encoder of the image generation model to obtain the implicit features of the downsampled image output by the image encoder in the latent space. Then, the scheduler is used to iteratively add noise to these implicit features to obtain the noisy features. The combined image formed by the downsampled image and the second mask image can be obtained by performing element-wise multiplication on the downsampled image and the second mask image.

[0098] Iterative denoising processing based on the noisy features according to the structural features and text style features can be achieved based on the reverse diffusion process; specifically, the iterative denoising processing can include processing of multiple denoising time steps. In the first denoising time step, the noisy features are used as the input to the noise prediction network, and the structural features and text style features are used as the conditional vectors (conditional guiding information) of the noise prediction network. Through the noise prediction network, the predicted noise for the current denoising time step is obtained. The scheduler is used to remove the predicted noise of the current denoising time step from the noisy features to obtain the implicit features generated in the first denoising time step. From the second denoising time step to the Nth denoising time step, in each denoising time step, the implicit features generated in the previous denoising time step are used as the input to the noise prediction network, and the structural features and text style features are used as the conditional vectors of the noise prediction network. Through the noise prediction network, the predicted noise for the current denoising time step is obtained. The scheduler is used to remove the predicted noise of the current denoising time step from the implicit features generated in the previous denoising time step to obtain the implicit features generated in the current denoising time step.

[0099] Among them, N can represent the number of iterations, and N can be a positive integer, such as 100, 200, or others. The fused features obtained by fusing the structural features and text style features can be used as the conditional vectors of the noise prediction network. The fused features can be obtained by concatenating the structural features and text style features, or by performing feature interaction on the structural features and text style features based on the attention mechanism. For example, the structural features and text style features can be input into a Transformer model for feature interaction. In the noise prediction network, information fusion between the conditional vectors and the input of the noise prediction network can be performed based on the cross-attention mechanism. For example, in the UNet network in Stable Diffusion, the conditional vectors can be input into each Transformer2DModel module of the UNet network, and the Transformer2DModel module can include a CrossAttention layer.

[0100] In this embodiment, the downsampled image is denoised to obtain a noisy feature. Since in the second mask image, each pixel in the image region corresponding to the target text region is a non-zero pixel, and each pixel in the image region corresponding to the content retention region is a zero pixel, the combined image formed by the downsampled image and the second mask image is denoised to obtain a noisy feature that can focus on the image region corresponding to the target text region for image generation. Based on the structural features and text style features, iterative denoising processing is performed on the noisy feature, and the image details can be gradually refined by gradually removing the noise, so that the generated intermediate image contains clear and sharp graphic text, improving the image quality.

[0101] In a specific application scenario, a schematic diagram of the simplified step flow for processing text in an image can be as follows Figure 2 shown, and a schematic diagram of the architecture of the algorithm processing module can be as follows Figure 3 shown. The algorithm processing module may include a structural feature extraction module, a text style feature extraction module, and an image generation model. The image generation model may use a diffusion model. The structural feature extraction module may include a feature extractor in the OCR model and a Transformer model; the text style feature extraction module may include a visual encoder, a position encoder, and a Transformer model; the diffusion model may include an image encoder, an image information creator, and an image decoder. The image information creator may include a UNet network and a scheduler; among them, each component in the above two modules and the diffusion model (such as the feature extractor in the OCR model) may be implemented using the corresponding specific pre-trained model of each component (such as SVTR in the pre-trained PaddleOCR-v4), or may be implemented using a model obtained by fine-tuning the corresponding pre-trained model with a training dataset. The training dataset may include images with graphic text in a preset language, and the preset language may be at least one of languages such as Chinese, Japanese, Korean, English, etc. In this way, the model obtained after fine-tuning training can have the ability to accurately process graphic text in the preset language. Based on Figure 2 and Figure 3 , the method for processing text in an image specifically may include the following steps.

[0102] The computer device can obtain the input image passed in by the user, perform image preprocessing on the input image to obtain an initial image; in response to the region selection operation (user interaction) on the initial image, determine the selected target text region, and determine the image region outside the target text region in the initial image as the content retention region; obtain the input text passed in by the user, and perform preprocessing on the input text to obtain the target text. Among them, the initial image may specifically be a scene image with graphic text.

[0103] The process of extracting the structural features of the target text by the structural feature extraction module can be referred to as shown in Figure 4 the schematic diagram of the structural feature extraction step process shown. The target text contains at least two characters. The computer device can respectively determine the character images formed by each character in the at least two characters contained in the target text, obtain the character image sequence of the target text composed of the character images formed by each character, and determine the text image formed by the target text in the form of a single-line text; use the feature extractor in the OCR model to respectively extract features from the text image and each character image in the character image sequence, obtain the global structural features extracted from the text image, and the local structural features extracted from each character image; input the local structural features and the global structural features into the Transformer model for feature interaction to obtain the structural features of the target text.

[0104] The process of extracting the text style features of the initial graphic text by the text style feature extraction module can be referred to as shown in Figure 5 the schematic diagram of the text style feature extraction step process shown. The computer device can obtain a first mask image matching the image size of the initial image; through the vision encoder, extract features according to the initial image and the first mask image to obtain the text visual features of the initial graphic text; through the position encoder, extract features from the first mask image to obtain the text position features; input the text visual features and the text position features into the Transformer model for feature interaction to obtain the text style features of the initial graphic text.

[0105] Refer to as shown in Figure 6 the schematic diagram of the diffusion model architecture shown. The computer device can downsample the initial image to obtain a downsampled image; obtain a second mask image matching the image size of the downsampled image; perform an operation on the downsampled image and the second mask image through element-wise multiplication to obtain a combined image formed by the downsampled image and the second mask image, input the combined image into the image encoder of the diffusion model to obtain an implicit feature, and perform iterative noise addition on the implicit feature through the scheduler of the diffusion model to obtain a noisy feature; use the structural features and the text style features as the conditional vectors of the UNet network in the diffusion model, and perform iterative denoising processing on the noisy feature through the UNet network and the scheduler to obtain a denoised feature; input the denoised feature into the image decoder in the diffusion model to output the generated intermediate image.

[0106] The computer device can perform image post-processing. Specifically, the image Poisson fusion algorithm is used to fuse the image area corresponding to the target text area in the upsampled image to the target text area in the initial image to obtain the target image. The target image is presented to the terminal used by the user, and can be specifically displayed in the image processing application running on the terminal. The image processing application can also provide export and sharing functions. The image processing application can record the user's editing behavior and feedback on the target image for subsequent iterative optimization of the algorithm and model.

[0107] In the traditional method of processing text in images, there is also the case of using GAN network (generative adversarial network) for processing. However, there are some problems in using GAN network for processing. The present application can solve these problems accordingly through the above steps. Specifically, 1. GAN network faces problems such as training instability and mode collapse, and is prone to produce blurred and distorted text effects; while the present application adopts a diffusion model as an image generation model, and the image details can be gradually refined through iterative denoising steps to generate clearer and sharper graphic text. 2. GAN network is limited by training data and is difficult to adapt to new fonts, languages or styles; the images used in the training data set of the present application can have graphic texts with a variety of different text styles. 3. GAN network is difficult to handle complex background textures and occlusion problems, and may produce abrupt splicing traces; the present application adopts OCR model, position encoder, mask image, etc., which can accurately extract features such as structure and position of text. 4. GAN network is limited to Latin language system and is difficult to expand to other languages with huge differences in glyph structure; the images in the training data set of the present application can include graphic texts with complex glyphs such as Chinese, Japanese, Korean, etc., which can improve the editing effect of graphic text. 5. The GAN network relies on end-to-end black box optimization and lacks control and explanation of the generation details; the architecture of this application can be divided into a structural feature extraction module, a text style feature extraction module and a diffusion model, which facilitates a more intuitive explanation and control of the generation process. 6. The GAN network is an integrated design with highly coupled modules, poor flexibility and scalability. The architecture of this application is divided and can be flexibly replaced and expanded.

[0108] Based on the above method for processing text in an image, graphic text editing is performed on a specific initial image example. For example, the initial graphic text to be edited may be the graphic text corresponding to “I believe in ten thousand hours” in the initial image, and the target text may be “Eight hundred standard soldiers rushing to the north slope”. Based on this, a target image is obtained by performing graphic text editing. The target graphic text formed by “Eight hundred standard soldiers rushing to the north slope” replaces the graphic text corresponding to “I believe in ten thousand hours”, and the target graphic text formed by “Eight hundred standard soldiers rushing to the north slope” has the same text style as the graphic text corresponding to “I believe in ten thousand hours”.

[0109] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this document, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0110] Based on the same inventive concept, an embodiment of the present application further provides an apparatus for processing text in an image for implementing the method for processing text in an image described above. The solution for solving the problem provided by this apparatus is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the apparatus for processing text in an image provided below can refer to the limitations on the method for processing text in an image in the above text, and will not be repeated here.

[0111] In an exemplary embodiment, as Figure 7 shown, an apparatus 700 for processing text in an image is provided, including: an acquisition module 710, a first feature extraction module 720, a second feature extraction module 730, and an image generation module 740, where:

[0112] The acquisition module 710 is configured to acquire an initial image, and determine a target text area and a content retention area in the initial image; the target text area includes the initial graphic text of the target text style.

[0113] The first feature extraction module 720 is configured to acquire the target text, and perform structural feature extraction on the image formed by the target text to obtain the structural features of the target text.

[0114] The second feature extraction module 730 is configured to perform text style feature extraction according to the target text area to obtain the text style features of the initial graphic text.

[0115] The image generation module 740 is configured to perform image generation based on the initial image, the structural features, and the text style features, so as to replace the initial graphic text in the initial image with the target graphic text formed by the target text according to the target text style, and retain the image content of the content retention area to generate a target image.

[0116] In an exemplary embodiment, the target text contains at least two characters, and the first feature extraction module 720 is further configured to respectively determine character images formed by each of the at least two characters; respectively perform feature extraction on the character images formed by the at least two characters to obtain character structure features of each of the at least two characters; and determine the structure feature of the target text based on the local structure feature obtained by splicing the character structure features of the at least two characters respectively.

[0117] In an exemplary embodiment, the first feature extraction module 720 is further configured to determine a text image formed by the target text in the form of a single-line text; perform feature extraction on the text image to obtain a global structure feature; and perform feature interaction based on the local structure feature and the global structure feature to obtain the structure feature of the target text.

[0118] In an exemplary embodiment, the second feature extraction module 730 is further configured to obtain a first mask image that matches the image size of the initial image. In the first mask image, each pixel in the image region corresponding to the target text region is a non-zero pixel, and each pixel in the image region corresponding to the content retention region is a zero pixel; perform feature extraction based on the initial image and the first mask image to obtain the text visual feature of the initial graphic text; and determine the text style feature of the initial graphic text based on the text visual feature.

[0119] In an exemplary embodiment, the second feature extraction module 730 is further configured to perform feature extraction on the first mask image to obtain a text position feature, where the text position feature is used to characterize the relative position of the target text region with respect to the initial image; and perform feature interaction based on the text visual feature and the text position feature to obtain the text style feature of the initial graphic text.

[0120] In an exemplary embodiment, the image generation module 740 is further configured to downsample the initial image to obtain a downsampled image; the downsampled image contains a downsampled graphic text formed by downsampling the initial graphic text in the initial image; perform image generation through a trained image generation model based on the downsampled image, the structure feature, and the text style feature, replace the downsampled graphic text in the downsampled image with an intermediate graphic text formed by the target text in the target text style to generate an intermediate image; upsample the intermediate image to the image size of the initial image to obtain an upsampled image; the image region corresponding to the target text region in the upsampled image contains a target graphic text formed by upsampling the intermediate graphic text; and fuse the image region corresponding to the target text region in the upsampled image into the target text region position in the initial image to obtain a target image.

[0121] In an exemplary embodiment, the image generation module 740 is further configured to obtain a second mask image that matches the image size of the downsampled image. In the second mask image, each pixel in the image region corresponding to the position of the target text region is a non-zero pixel, and each pixel in the image region corresponding to the position of the content retention region is a zero pixel; perform noise addition processing on the combined image formed by the downsampled image and the second mask image to obtain a noisy feature; perform iterative denoising processing based on the noisy feature according to the structural feature and the text style feature, and replace the downsampled graphic text in the downsampled image with an intermediate graphic text formed by the target text according to the target text style to generate an intermediate image.

[0122] Each module in the above apparatus for processing text in an image can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of the processor, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0123] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data required to be stored in the method for processing text in an image described above. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for processing text in an image.

[0124] Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0125] In one embodiment, a computer device is further provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented.

[0126] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0127] In one embodiment, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0128] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant regulations.

[0129] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.

[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.

[0131] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for processing text in an image, characterized in that The method includes: Obtain an initial image, and determine a target text region and a content retention region in the initial image; the target text region contains initial graphic text of a target text style. Obtain target text, perform structural feature extraction on an image formed based on the target text, and obtain structural features of the target text. Perform text style feature extraction according to the target text region, and obtain text style features of the initial graphic text. Perform image generation based on the initial image, the structural features, and the text style features, so as to replace the initial graphic text in the initial image with target graphic text formed by the target text according to the target text style, and retain the image content of the content retention region to generate a target image.

2. The method according to claim 1, wherein The target text contains at least two characters. Performing structural feature extraction on an image formed based on the target text to obtain structural features of the target text includes: Respectively determine character images formed by each of the at least two characters. Respectively perform feature extraction on the character images formed by the at least two characters, and obtain character structural features of the at least two characters respectively. Based on local structural features obtained by splicing the character structural features of the at least two characters respectively, determine structural features of the target text.

3. The method according to claim 2, wherein Determining structural features of the target text based on local structural features obtained by splicing the character structural features of the at least two characters respectively includes: Determine a text image formed by the target text in the form of a single-line text. Perform feature extraction on the text image to obtain global structural features. Perform feature interaction based on the local structural features and the global structural features to obtain structural features of the target text.

4. The method according to claim 1, characterized in that Performing text style feature extraction according to the target text region to obtain text style features of the initial graphic text includes: Obtain a first mask image matching the image size of the initial image. In the first mask image, pixels in the image region corresponding to the position of the target text region are non-zero pixels, and pixels in the image region corresponding to the position of the content retention region are zero pixels. Perform feature extraction according to the initial image and the first mask image to obtain text visual features of the initial graphic text. Based on the text visual features, determine text style features of the initial graphic text.

5. The method according to claim 4, characterized in that, Determining text style features of the initial graphic text based on the text visual features includes: Perform feature extraction on the first mask image to obtain text position features, where the text position features are used to characterize the relative position of the target text region relative to the initial image. Perform feature interaction based on the text visual features and the text position features to obtain text style features of the initial graphic text.

6. The method according to any one of claims 1-5, characterized in that, Performing image generation based on the initial image, the structural features, and the text style features to replace the initial graphic text in the initial image with the target graphic text formed by the target text in the target text style and retain the image content in the content retention area, generating a target image, including: Downsampling the initial image to obtain a downsampled image; the downsampled image includes the downsampled graphic text formed by downsampling the initial graphic text in the initial image; Based on the downsampled image, the structural features, and the text style features, performing image generation through a trained image generation model to replace the downsampled graphic text in the downsampled image with the intermediate graphic text formed by the target text in the target text style, generating an intermediate image; Upsampling the intermediate image to the image size of the initial image to obtain an upsampled image; the image area corresponding to the target text area position in the upsampled image includes the target graphic text formed by upsampling the intermediate graphic text; Fusing the image area corresponding to the target text area position in the upsampled image into the target text area position in the initial image to obtain a target image.

7. The method according to claim 6, wherein The performing image generation based on the downsampled image, the structural features, and the text style features through a trained image generation model to replace the downsampled graphic text in the downsampled image with the intermediate graphic text formed by the target text in the target text style, generating an intermediate image, includes: Obtaining a second mask image matching the image size of the downsampled image, in the second mask image, each pixel in the image area corresponding to the target text area position is a non-zero pixel, and each pixel in the image area corresponding to the content retention area position is a zero pixel; Performing noise addition processing on the combined image formed by the downsampled image and the second mask image to obtain a noisy feature; Performing iterative denoising processing based on the noisy feature according to the structural features and the text style features to replace the downsampled graphic text in the downsampled image with the intermediate graphic text formed by the target text in the target text style, generating an intermediate image.

8. An apparatus for processing text in an image, characterized in that, The apparatus includes: An acquisition module, configured to acquire an initial image and determine a target text area and a content retention area in the initial image; the target text area includes the initial graphic text in the target text style; A first feature extraction module, configured to acquire a target text and perform structural feature extraction on the image formed based on the target text to obtain the structural features of the target text; A second feature extraction module, configured to perform text style feature extraction according to the target text area to obtain the text style features of the initial graphic text; An image generation module, configured to perform image generation based on the initial image, the structural features, and the text style features, so as to replace the initial graphic text in the initial image with a target graphic text formed by the target text according to the target text style, and retain the image content in the content retention area to generate a target image.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.