Image processing method and device

By acquiring style attributes from images and using diffusion processing to generate target text, the problem of obvious modification traces in image modification is solved, achieving uniform text style and improved viewing experience.

CN121582380APending Publication Date: 2026-02-27ZHUHAI KINGSOFT OFFICE SOFTWARE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511981002.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

When modifying text in an image using existing technology, there are obvious traces of modification between the modified text and the unmodified text, resulting in a poor viewing experience for the reader.

Method used

By acquiring the initial image and processing information, the style attributes of the text are determined, and diffusion processing is used to generate target text with style attributes, which overwrites the original text, ensuring that the modified text and the unmodified part have the same style.

Benefits of technology

This achieves stylistic consistency between the modified text and the unmodified parts, eliminates traces of modification, and enhances the reader's viewing experience of the new image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582380A_ABST
    Figure CN121582380A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image processing method and device. According to the image processing method, an initial image needing to be processed, a processing area of the initial image and target characters needing to be presented in the processing area are obtained, visual attributes of characters in the area except the processing area are determined, and the initial image is subjected to diffusion processing in the processing area to generate an area image; the content of the area image is the target character with the visual attribute of the character of the area outside the processing area, and the target character generated through diffusion processing covers the original character of the processing area and has the same style attribute as the character which is not modified in the initial image, so that the modification of the existing character in the initial image is realized; and the overall styles of the modified characters and the characters in the initial image are unified, so that the modification trace between the modified characters and the non-modified parts is eliminated, and the viewing feeling of a reader on the modified new image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of image processing, and in particular, to an image processing method and device. BACKGROUND

[0002] In the field of image editing, it is often necessary to modify the text in the original image (such as correcting errors or updating versions), and it is required that the modified text in the modified image is highly consistent with the text before modification in the original image in terms of font, size, style, color, background texture, and overall visual style, achieving the effect of no trace modification. This modification requirement may face the situation that the original image only has a bitmap file, and the text that needs to be modified cannot be directly selected to maintain the original visual style and only modify the text content, but can only be modified by taking the original image corresponding bitmap file as the modification object.

[0003] The applicant found, after observing the new image obtained by using the related technology to modify the text in the original image, that there were obvious modification traces between the modified text and the unmodified part in the new image, and the reading experience of the new image was poor. SUMMARY

[0004] Embodiments of the present application provide an image processing method and device, which solve the technical problem that after modifying the text in the image, the modified text and the unmodified part have obvious modification traces, and the reading experience of the new image is poor.

[0005] In a first aspect, embodiments of the present application provide an image processing method, which comprises: obtaining an initial image and processing information, the processing information comprising region information and text information; determining the style attribute of the text in the first region in the initial image, the first region being an associated region of the second region, and the second region being determined according to the region information; generating target text with the style attribute by diffusion processing of the initial image in the second region to obtain a target image, the target text being determined according to the text information.

[0006] In a second aspect, embodiments of the present application provide an image processing device, which comprises: a data acquisition unit configured to obtain an initial image and processing information, the processing information comprising region information and text information; an information extraction unit configured to determine the style attribute of the text in the first region in the initial image, the first region being an associated region of the second region, and the second region being determined according to the region information; a text redrawing unit configured to generate target text with the style attribute by diffusion processing of the initial image in the second region to obtain a target image, the target text being determined according to the text information.

[0007] In a third aspect, the embodiments of the present application provide an electronic device, the electronic device comprising: one or more processors; a memory for storing one or more computer programs; when the one or more computer programs are executed by the one or more processors, the electronic device implements the image processing method according to the embodiments of the present application.

[0008] In a fourth aspect, the embodiments of the present application provide a non-volatile storage medium storing computer executable instructions, which, when executed by a computer processor, are configured to perform the image processing method according to the embodiments of the present application.

[0009] In the embodiments of the present application, the initial image to be processed is obtained, and the style attribute of the text in the region outside the processing region is determined according to the processing region of the initial image and the target text to be presented in the processing region. The region image is generated in the processing region of the initial image by diffusion processing, and the content of the region image is the target text with the style attribute of the text in the region outside the processing region. The target text generated by diffusion processing covers the original text in the processing region, and is the same as the style attribute of the text in the initial image that is not modified. The modification of the text in the initial image is realized, and the modified text is uniform with the overall style of the text in the initial image. The modification trace between the modified text and the unmodified part is eliminated, and the viewing experience of the reader for the new image after modification is improved. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0011] Figure 1 A method flowchart of an image processing method provided by the embodiments of the present application.

[0012] Figure 2 A process diagram for determining a style attribute provided by the embodiments of the present application.

[0013] Figure 3 A process diagram for redrawing target text provided by the embodiments of the present application.

[0014] Figure 4 A model architecture diagram of a character feature extraction model provided by the embodiments of the present application.

[0015] Figure 5 A schematic diagram of reconstructing a character image is provided for an embodiment of the present application.

[0016] Figure 6 A schematic diagram of a denoising process is provided for an embodiment of the present application.

[0017] Figure 7 A schematic diagram of a model architecture of a diffusion model is provided for an embodiment of the present application.

[0018] Figure 8 A comparison diagram of images before and after image processing is provided for an embodiment of the present application.

[0019] Figure 9 A data change schematic diagram of an image processing method is provided for an embodiment of the present application.

[0020] Figure 10 A structural schematic diagram of an image processing device is provided for an embodiment of the present application.

[0021] Figure 11 A structural schematic diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0022] The following description and drawings are illustrative of the specific embodiments of the present application and are not intended to limit the generality of the application. The embodiments are merely representative of possible variations. Individual components and functions are optional unless specifically required, and the order of operations can be varied. Portions and features of some embodiments can be included in, or substituted for, those of other embodiments. The scope of the embodiments of the present application encompasses the full range of equivalents of the claims, and all available extensions of the claims. In this document, the term "application" can be used to refer to one or more embodiments of the application, and the terms "application" and "application(s)" can be used interchangeably with the term "invention." In this document, relational terms such as first and second, and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. The terms "comprises," "comprising," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. The embodiments are described in a progressive manner, with each embodiment highlighting the differences from other embodiments, and the same or similar parts between embodiments are cross-referenced. For the structure, product, and the like disclosed in the embodiments, since they correspond to the parts disclosed in the embodiments, they are described simply, and the relevant parts are cross-referenced with the method part.

[0023] In the field of image editing, it is often necessary to modify the text in the original image (such as correcting errors, version updates), requiring the modified text in the modified image to be highly consistent with the text before modification in the original image in terms of font, size, style, color, background texture, and overall visual style to achieve the effect of no trace modification. This modification requirement may face the situation that the original image only has a bitmap file, and the text that needs to be modified cannot be directly selected to maintain the original visual style and only modify the text content. Only the original image corresponding bitmap file can be used as the modification object for modification.

[0024] The applicant observed the new image obtained by using related technologies to modify the text in the original image and found that there were obvious modification traces between the modified text and the unmodified part in the new image, and the reader's viewing experience of the new image was poor.

[0025] To solve the above technical problems, the applicant summarized the underlying implementation principles of related technologies for text modification in images based on bitmap files and found that there are mainly the following technical routes for implementing text modification in images: 1. Pixel-based image processing, the underlying implementation principle is to erase the original text and fill the background, and then superimpose new text; 2. Deep learning method, for example, combined with OCR (Optical Character Recognition, optical character recognition) positioning, using generation models such as GANs (Generative Adversarial Networks, generative adversarial networks) to generate new text images and fusion; 3. Diffusion model method, using Stable Diffusion and other general models to generate new text images according to prompt words.

[0026] The applicant further analyzed the problems caused by the underlying implementation principles of each technical route and found that pixel-based image processing easily produces artifacts when erasing the original text and filling the background, and it is difficult to match the background when superimposing new text, resulting in obvious modification traces and a clear sense of pasting new text on the original image in the modified image. Deep learning methods and diffusion model methods belong to artificial intelligence biased methods, and such methods rely too much on text encoders and lack sufficient control over specific font, font size, character spacing, and stroke thickness. Moreover, various methods in related technologies lack a display font feature guiding mechanism, resulting in inconsistent font style of the added new text with the overall image. In summary, the weaknesses of each related technology in different dimensions result in a sense of fragmentation between the modified area and the surrounding image in terms of texture, lighting, color, and style in the new image modified by related technologies, visual continuity is damaged, modification traces are obvious, and the reader's viewing experience of the new image is poor.

[0027] In view of the above technical problems, and based on analysis of the underlying causes of the technical problems, the applicant proposes an image processing method, acquires an initial image that needs to be processed, and determines the style attribute of the text in the region outside the processing region of the initial image and the target text that needs to be presented in the processing region, generates a region image in the processing region of the initial image through diffusion processing, the content of the region image is the target text with the style attribute of the text in the region outside the processing region, the target text generated through diffusion processing covers the original text in the processing region, and is the same as the style attribute of the text that is not modified in the initial image, which realizes modification of the text in the initial image, and the modified text is uniform with the overall style of the text in the initial image, eliminates modification marks between the modified text and the part that is not modified, and improves the viewing experience of readers for the new image after modification.

[0028] The image processing method in the present scheme is applied to an electronic device, which can be an integrated device for controlling and realizing human-computer interaction operation on the content displayed on the display screen through an external device, such as a smart phone, a personal computer, an interactive smart tablet, etc.

[0029] In the electronic device, at least one operating system is installed locally and / or in the cloud, wherein the operating system includes but is not limited to an Android system, a Linux system, a Windows system, and a Huawei Hongmeng system, which is used to control and coordinate the electronic device and external devices, so that various independent hardware in the electronic device can work coordinately as a stable whole, and the architecture level where the operating system is located is defined as a system layer. On the basis of the system layer, application programs developed to meet different fields and different problems of users are installed in the electronic device, and the corresponding architecture level thereof is an application layer.

[0030] The image processing method in the present application embodiment can be realized by a functional module integrated in the system layer in the electronic device, such as an artificial intelligence or image editing tool provided by the operating system; can also be realized by a separately installed application, such as an application developed specifically for text modification in images; or can be realized by a certain functional module of a certain application in the application layer, such as a functional interface provided by an application with file editing as the main function, and the user starts the functional module to realize image processing through the functional interface. The actual implementation mode is not limited here.

[0031] Please refer to Figure 1 which is a method flowchart of an image processing method provided by the present application embodiment. As Figure 1 shown, the image processing method includes but is not limited to steps S110-S130, and the image processing method provided by the present application embodiment is used in the electronic device described above.

[0032] Step S110: obtaining an initial image and processing information, the processing information including region information and text information.

[0033] In the embodiments of the present application, the initial image is an image uploaded by a user and needs to be modified, such as an image in a bitmap file of various formats. Since the user uploads the initial image for implementing the image processing method, the target is to modify the text in the initial image, and the initial image can have text by default. The user can upload a locally saved image, a cloud disk saved image, or a link to provide an image. Correspondingly, obtaining the initial image directly receives the bitmap file corresponding to the image or reads the corresponding bitmap file according to the link.

[0034] The processing information is provided by the user according to the modification needs of the initial image, that is, the user provides the initial image and information that can explicitly indicate the modification target of the initial image, including region information and text information. The region information is used to determine which part of the text in the initial image needs to be modified, and the text information is used to determine what content the text needs to be modified to.

[0035] Which part of the text in the initial image needs to be modified is determined by the region information. In an optional implementation, the region information is operation information determined by a region selection operation on the initial image. After obtaining the initial image, the initial image is presented to the user in a preview form, and the user performs a region selection operation in the initial image in the preview form, such as box selection by mouse operation, or box selection by touch screen operation interactive control, or scribbling by touch screen. The position information corresponding to the box selection or scribbling is the operation information determined by the region selection operation. The electronic device can determine which part of the text in the initial image needs to be modified according to the operation information, based on the known display position of the initial image. In another optional implementation, the region information is first text information determined by describing the text displayed in the initial image. For example, the user can indicate what text needs to be modified by inputting text at the same time or after uploading the initial image. Exemplary first text information: “modify the fifth line of text”, “modify ‘pattern’ ”. Based on the analysis of the first text information, it can be determined that the fifth line of text (the specific text content is determined according to the recognition of the initial image) or the word “pattern” in the initial image needs to be modified. In the specific implementation process, if the specific text is specified as the text to be modified, and multiple places with the same text are recognized in the initial image, the user can be prompted by a prompt box to further determine whether the text at one or more places needs to be modified, and the region information is finally determined according to the user's determination result.

[0036] The text to be modified is modified into what, which is determined by the text information. In an optional implementation, the text information is second text information based on the target text. For example, the user can indicate what the text to be modified is modified into by inputting the text at the same time of uploading the initial image or after uploading the initial image. An example of the second text information is: "modify into 'eyesight'". Based on the analysis of the second text information, it can be determined that the specified text in the initial image needs to be modified into "eyesight", that is, "eyesight" is the target text. In another optional implementation, the text information is image information including the target text. For example, a screenshot with text, and the target text is determined by recognizing the text of the screenshot.

[0037] In an optional implementation, the first text information and the second text information can be part of a whole paragraph of text. For example, a whole paragraph of text is received, which includes the initial image and "modify the text in the third row into'sky' " or "modify'sky' into'sky' ". Based on the analysis of the whole paragraph of text, it can be determined that the text in the third row of the initial image or "sky" in the initial image needs to be modified into "sky", that is, "sky" is the target text. The analysis of the first text information, the second text information and the whole paragraph of text containing the first text information and the second text information is a conventional capability of artificial intelligence, and the specific implementation process of the analysis itself is not described in the present application.

[0038] In addition, it should be noted that for a single initial image, there can be multiple modifications, and the multiple modifications can be completed at one time. In the embodiments of the present application, the region determined by the same region selection operation is defined as an independent complete region, and the corresponding text information is determined. The continuous region determined according to the first text information is also defined as an independent complete region. For example, "modify the text in the first row and the second row into 'the sun sets on the mountain, and the Yellow River flows into the sea'", because the text in the first row and the second row is continuous, it is regarded as an independent complete region here. For example, "modify the first word in the first row and the last word in the third row into 'hundred' and 'pasture' respectively", because the first word in the first row and the last word in the third row are not continuous, they are regarded as two independent complete regions, and the corresponding text information is determined. Each independent complete region corresponds to the independent completion of the image processing method in the embodiments of the present application.

[0039] Step S120: determining the style attribute of the text in the first region in the initial image, the first region being the associated region of the second region, and the second region being determined according to the region information.

[0040] In the embodiment of the present application, the area where the text to be modified is located, i.e., the area directly determined by the area information, is defined as the second area. Considering that the modified text needs to be consistent with the overall style of the text in the initial image that has not been modified, the text in the area that has not been modified (i.e., the first area determined with reference to the second area) is subjected to information extraction, and the corresponding style attribute is determined, so as to ensure that the subsequently added text is modified according to the style attribute and is consistent with the overall style of the text in the initial image that has not been modified. The style attribute is used to determine the color, shape, and other aesthetic effects of the text seen by the user from the image, regardless of the semantics of the text. The style attribute can include one or more of font, color, thickness, and font size. The font is, for example, Songti, Kaishu, etc.; the color is, for example, black, red, color with multiple color combinations, etc.; the thickness is, for example, regular thickness, bold, other specific thickness, etc.; and the font size is, for example, represented by the number of pixels. In addition to the exemplary style attributes described above, other style attributes such as light and shadow, whether underlined, whether framed, etc. can also be included, without limitation. Determining as many various style attributes as possible can ensure that the text in the initial image that has been modified is consistent with the text that has not been modified in various dimensions during the image processing process, thereby ensuring a high degree of visual style consistency of the text and reducing modification traces.

[0041] In the case where the text in the second area needs to be modified in a diffusion manner and the visual effect of the modified text is consistent with the text that has not been modified (i.e., the text outside the second area), the style attribute of the target text is determined based on the style attribute of the text in the first area, which can avoid irrelevant information (semantics, background) of the text features in the second area from being included in the style attribute, thereby excluding irrelevant noise and ensuring that the finally added text is the target text that focuses on pure visual style. The first area and the second area are associated and do not overlap each other. The specific association relationship can be that the first area is adjacent to the second area, the distance between the first area and the second area is less than a distance threshold, the text in the first area and the text in the second area have a semantic connection (i.e., the similarity of the text in the first area and the text in the second area is greater than a semantic similarity threshold), etc., without limitation.

[0042] In an optional implementation, the process of determining the style attribute in step S120 can include steps S121-S123: Step S121: determining the second area from the initial image according to the area information.

[0043] According to the region selection operation of the user or the indication range of the initial image in the input first text information, the second region is determined. For example, the region selection operation is the framing of a certain part of the initial image, and the second region is within the framing range. The region selection operation is the scribbling of a certain part of the initial image, and the second region is within the scribbling range or the region determined by the scribbling range (for example, the minimum circumscribed rectangle of the scribbling range). If the region information is the first text information, the second region can be within the minimum circumscribed rectangle corresponding to the to-be-modified text determined according to the first text information.

[0044] Step S122: Recognizing the text outside the second region to determine the first region corresponding to the preset number of texts closest to the second region.

[0045] In the determination of the first region based on the second region, the first region can be determined based on the text adjacent to the second region. The style attribute of the text in the first region determines the style attribute of the text in the second region after the modification, so as to ensure the style uniformity between the text in the second region after the modification and the text outside the second region. In the determination of the first region based on the second region, the text outside the second region is recognized, and the first region is determined according to the relative position relationship between the recognized text and the second region and the number of texts required for accurate analysis of the visual effect. The style attribute is analyzed for the text closest to the second region as possible, and after the target text with the style attribute is added to the second region, the text in the second region after the modification can be as uniform as possible with the visual effect of the adjacent text. Through the constraint of the number of texts by the preset number, the comprehensive and universal style attribute can be ensured, and the text in the second region after the modification can be as uniform as possible with the overall visual effect of the adjacent text, rather than the visual effect of one or a small number of texts. In the determination of the preset number of texts closest to the second region, the region determined by the texts is the first region, that is, the first region corresponding to the preset number of texts closest to the second region is determined.

[0046] In another optional implementation manner of determining the first region based on text recognition, the text outside the second region is subjected to semantic recognition, and the region in which the text with the same, opposite, similar, or other semantic association relationship with the target text is located is taken as the first region. In the target image obtained by the image processing based on the embodiments of the present application, the texts with the same, opposite, similar, or other semantic association relationship are displayed with the same visual effect, which improves the contrast effect of the texts in the target image. In addition, the first region can be determined based on the semantic recognition and the distance from the second region. That is, if it is determined that there is text with the same, opposite, similar, or other semantic association relationship with the target text, the first region is directly determined according to the region in which the text is located, and if it is determined that there is no such text, the first region is determined according to the text closest to the second region as possible.

[0047] In another optional implementation, the manner of determining the first region in step S122 can also be replaced by: determining the first region according to a preset distance parameter with reference to the second region. That is, after the second region is determined, the outer contour of the second region is expanded outward by a preset distance parameter to obtain a new outer contour, and the region between the new outer contour and the outer contour of the second region is the first region.

[0048] The distance parameter can be an absolute distance in pixels, that is, according to the statistics of the distance between the text that needs to be modified in the image that needs to be modified and the text in the adjacent row (or column), a distance in pixels is determined, and most images include at least one row (or column) of text within the distance range with the region where any text is located as the center. When the first region is determined according to the absolute distance in pixels, the second region is expanded outward by the absolute distance to obtain a new outer contour, and the region between the new outer contour and the outer contour of the second region is the first region.

[0049] The distance parameter can also be a relative distance in text, that is, the first region is determined according to the personalized distribution of specific text in the initial image. For this manner, in the process of specifically determining the first region, according to the recognition result of the text in the initial image, the position of one or more rows (or columns) of text is determined in one or more directions respectively, and then the distribution region of the determined text is taken as the first region. The number of specific directions and the number of rows (or columns) are taken as the parameter value of the distance parameter.

[0050] Step S123: performing visual effect analysis on the text in the first region to obtain style attributes.

[0051] On the basis of having determined the first region, the visual effect analysis is performed on the text in the first region to obtain one or more style attributes of font, color, thickness and font size, for example, the specific style attribute values such as “Songti, black, bold, 12-point font” are recognized for the text in the first region in a certain initial image. For the font, it can be recognized based on stroke features (for example, with or without serifs, corner shapes); the color can be recognized by extracting the RGB average, brightness, contrast and the like of the text display position; the thickness can be recognized by the text stroke edge gradient; the font size can be recognized by the pixel width and height of a single text; in addition, the horizontal / vertical spacing between texts can also be used to recognize the layout parameters. The specific recognition manner can be based on the recognition of graphics or based on various pre-trained networks, and is not limited specifically. The recognized style attribute values can be represented by specific names or by corresponding class labels.

[0052] Step S130: The initial image is processed in the second region to generate target text with style attributes, and the target text is determined based on the text information.

[0053] The initial image is structurally complete. To eliminate interference from the original content in the second region during diffusion processing, the second region of the initial image can be masked. Then, diffusion processing is performed on this second region using a diffusion model. The content processed by this model is constrained by the background, the second region, and the image features of the target text presented as style attributes in the initial image. Ultimately, after diffusion processing, the local image outside the second region in the target image is identical to the corresponding local image in the initial image. The background of the second region is generated based on the diffusion of the initial image, ensuring no modification traces in the background. Modified text in the second region is generated through diffusion under the constraints of the image features of the target text presented as style attributes, ensuring no modification traces in the foreground. This eliminates the modification traces between the modified text and the unmodified parts, improving the viewer's experience of the modified image.

[0054] In one optional implementation, such as Figure 3 As shown, the process of performing diffusion processing in the second region of the initial image to obtain the target image in step 130 may include, but is not limited to, steps S131-S132. Up to step S120, the target effect after modifying the initial image is still described at the data level through attribute values ​​or text. Steps S131-S132 demonstrate the detailed process of obtaining the target image based on the data level description.

[0055] Step S131: Extract features from the target text drawn according to style attributes to obtain the target font features.

[0056] The image processing method in this embodiment aims to add target text with a specific visual effect to an initial image. This process can be understood as redrawing the initial image based on another image. The redrawing method is exemplified by diffusion processing based on a diffusion model, where the other image is the text image obtained by displaying the target text according to style attributes. The subsequent redrawing process does not involve directly pasting the text image onto the initial image to cover the original text and background in the second area, as this pasting method might result in the background of the second area being missing, and the text style might not be consistent with other unmodified text.

[0057] Therefore, in the embodiments of the present application, the initial image is redrawn in a diffused manner based on the image features of the character image, and a dedicated character feature extraction model is required to extract features of the characters in the character image. When generating a character image according to a target character, the characters in the character image are drawn according to the style attribute, and in addition to ensuring the display of the target character, the image size also meets the requirements of the character feature extraction model for the image size. Each character in the target character corresponds to a character image, that is, the character image is a single character image, and the image size of the character image is, for example, 64x64, 128x128 or more (in pixels).

[0058] The feature extraction of the character image is completed by a pre-trained character feature extraction model. An example of the model architecture of the character feature extraction model is shown in Figure 4 The components of the character feature extraction model can include a character encoding module (i.e., an encoder), a position encoding module, a character classification module (i.e., a character classifier), and a position classification module (i.e., a position classifier). The character encoding module, which is equivalent to the encoder, is responsible for extracting the target font features of the character image; the position encoding module uses a learnable embedding vector (pos_embedding) to enhance the character feature extraction model's ability to perceive the position information of the character in the character sequence, to ensure the accurate position of a single character in the target character and avoid the situation that "hello" is displayed as "hello", and the position information can be set to a maximum character sequence length, for example, 16, 18, etc.; the character classification module is used to identify which of the predefined character categories the character in the character image belongs to, for example, the predefined character categories include 6000 common Chinese characters, if "you" is the 300th Chinese character in the character category, the corresponding character category is 300, and if "good" is the 499th Chinese character in the character category, the corresponding character category is 499; the position classification module is used to predict the specific position index of the character in the character sequence of the target character, and if the maximum character sequence length is 18, the position index is 0 to 17.

[0059] In a specific implementation, the character encoding module is composed of a 5-layer convolutional structure, and the number of channels increases layer by layer (32→64→128→256→512). Each layer includes a convolution operation (Conv2d), a batch normalization (BatchNorm), and a LeakyReLU activation function. Through 5 times of downsampling (with a step of 2), the size of the input character image is compressed from 64x64 to 2x2 feature maps. Finally, a convolutional layer is used instead of a fully connected layer to generate a mean vector μ representing the latent distribution and a log variance vector log_var.

[0060] In the pre-training phase, the character decoding module (also known as the decoder) is also involved in the training process of the character feature extraction model. Figure 4The character decoding module mainly participates in the training corresponding to the character encoding module in the character feature extraction model. The character decoding module is used to reconstruct the verification image according to the feature vector output by the character encoding module in the training process, so as to judge the feature extraction capability of the character encoding module through the verification image.

[0061] In an optional implementation, the character decoding module is designed to have two parallel outputs, i.e., a main output and an auxiliary output, which are generated based on the same set of features. The auxiliary output is an auxiliary supervision branch, which can enhance the font feature representation capability of the model and is finally used for font feature injection of the subsequent diffusion model. The two outputs share the latent features and position encoding information extracted by the encoder, realize decoding of the same set of features by using two sets of decoding paths, and thus enable the character feature extraction model to mine more rich font feature details through double reconstruction tasks, thereby avoiding the problem of insufficient feature learning caused by a single output. After the training is completed, the auxiliary output is discarded, and only the latent features corresponding to the main output are used for the subsequent character diffusion generation module, i.e., the font features carried by the main output are extracted as z_font, which is then injected into the diffusion model as a control signal to realize traceless text modification. The value of the auxiliary output lies in feature enhancement in the training phase, which does not affect the process in the inference phase. The training basis of the character feature extraction model is the first sample set, which is a training set constructed according to the feature extraction object of the character feature extraction model, i.e., the first sample set includes a large number of sample text images. The sample text images are input into the initial feature extraction model multiple times, and the similarity between each input sample text image and the corresponding reconstructed training process image is determined. If the similarity does not reach a preset similarity threshold value or the similarity can still be improved, the model parameters of the initial feature extraction model are adjusted, the feature extraction capability of the initial feature extraction model is optimized, and the sample text image is continuously input for training. When the similarity reaches the preset similarity threshold value or the similarity cannot be continuously improved, it is considered that the training of the initial feature extraction model is completed, the latest model parameters are retained, and the final available character feature extraction model is obtained. After the training of the character feature extraction model is completed, Figure 4 The decoder (i.e., the character decoding module described above) can be retained or discarded, which does not affect the feature extraction function in the inference phase.

[0062] In the specific model structure level, for example, Figure 4As shown, the character decoding module is a decoder corresponding to the encoder, and the decoder performs a symmetric upsampling operation, including a 5-layer ConvTranspose2d structure, and the number of channels decreases (512→256→128→64→32). The last layer includes a ConvTranspose2d, a batch normalization, a LeakyReLU activation, followed by a Conv2d layer and a Tanh activation function to constrain the output value range to [-1, 1], so as to reconstruct the original character image, for realizing the judgment of the feature extraction capability of the encoder.

[0063] The position embedding is a learnable parameter matrix (with dimensions max_len=18 x mid_dim=2048), which is used by the position encoding module to perceive the position information of the character in the character sequence.

[0064] The character feature extraction model further includes a classification part, and the classifier part is composed of two independent fully connected layers. The character classifier corresponding to the character classification module maps the 2048-dimensional features to 3848 character categories; and the position classifier corresponding to the position classification module maps the same features to 18 position categories.

[0065] The character feature extraction model extracts features from one or more character images corresponding to the target character (determined by the number of characters of the target character). For a modification of an independent complete region, the input of the character feature extraction module includes: a batch of character images x (with dimensions BxCxHxW), corresponding character IDs char_ids, and position information positions. The output includes a reconstructed character image x_recon, an auxiliary reconstructed image x_recon_aux (used to enhance feature representation), predicted logic values (logits) of character classification, and predicted logic values of position classification.

[0066] Based on the above architecture, the character feature extraction model obtained by training is used to extract features from character images to obtain a target feature vector. In this process, the character encoding module is used to extract features from the character images to obtain corresponding target font features, with each character image corresponding to a target character; the position encoding module is used to extract position features of the characters in the character sequence of the target character; the character classification module is used to identify target category features of the characters in the character image in the predetermined character categories; the position classification module is used to predict position index features of the characters in the character image in the character sequence of the target character; and the target feature vector is determined according to the target font features, the position features, the target category features, and the position index features.

[0067] The character feature extraction model is trained based on a multi-task joint loss function, and the multi-task joint loss function includes a reconstruction loss, a KL divergence loss, a gradient loss, etc., without limitation. The reconstruction loss is used to represent the pixel-level mean square error between the input encoder text image (i.e., the sample text image) and the output image reconstructed by the decoder corresponding to the sample text image, which can ensure visual similarity. The KL divergence loss can constrain the latent distribution z output by the encoder to be close to the standard normal distribution, promoting the regularity of the latent space. The gradient loss is used to represent the mean square error of the Sobel gradient map in the horizontal and vertical directions of the input text image and the reconstructed output image, which can maintain the font edge sharpness.

[0068] Suppose a target character is "A", and the style attributes set include the font (e.g., Songti). To achieve input of a 64x64 pixel "Songti" character "A" text image, the character ID is 100, and the position index in the text sequence is 5. The text image, character ID, and position index are input into the character feature extraction model. During the processing of the text image by the character feature extraction model, the character encoding module processes the text image and outputs the latent feature vector z and its distribution parameters μ and log_var; the position encoding module queries the learnable embedding vector pos_embedding[5] corresponding to position 5 and combines it with the character feature (e.g., concatenation or addition). The character classification module predicts the probability distribution of the character belonging to all supported character categories based on the fused features (expecting the highest probability corresponding to ID 100). The position classification module predicts the position probability distribution of the character in the sequence based on the same features (expecting the highest probability corresponding to index 5).

[0069] The decoder reconstructs the "Songti" character "A" text image and the auxiliary reconstructed image using the sampled latent vector z (or combined with the position embedding).

[0070] During the training of the character feature extraction model using the "Songti" character "A" text image as a training sample, all loss terms are calculated for each loss term constituting the total loss function: the reconstruction loss ensures that the reconstructed image is visually close to the original "Songti" "A"; the KL divergence loss ensures that the distribution of the latent vector z is close to the standard normal distribution; and the gradient loss ensures that the edges of the reconstructed character (such as the sharp corners and horizontal edges of "A") are as sharp and clear as the original character. During the training of the character feature extraction model, the loss function that integrates multiple loss terms is used as a constraint, and the parameters are updated through backpropagation and an optimizer, learning the ability to effectively encode "Songti" style features and accurately identify characters and their positions. For example, Figure 5 ​As shown, the character feature extraction model after training can reconstruct the corresponding character image according to the features extracted by the encoder under the condition that the corresponding character ID and position index in the text sequence are input at the same time. Figure 5 The order in which the character images are presented in the above table is not necessarily the presentation order determined according to the position index, but is only used to present the character images obtained after the decoder reconstructs the multiple texts in sequence.

[0071] In the embodiment of the present application, the total loss function is constructed by taking the loss term that may occur in each dimension between the character image reconstructed by the decoder and the input text image as a constraint on the training process of the character feature extraction model, so as to ensure that the character encoding module in the character feature extraction model can comprehensively and accurately extract the visual features of the text in the text image generated according to the text to be displayed, and in the case where the style attribute of the text in the text image is determined by the style attribute of the text that does not need to be modified in the initial image, the subsequent redrawing according to the features extracted by the character encoding module can present the same text content as the modification target in the image, and the visual effect of the modified text content is highly unified with the visual effect of the unmodified text in the initial image. The features extracted by the character encoding module from the input text image and related classification and position information are the target font features.

[0072] Step S132: Perform diffusion processing on the second region of the initial image under the constraint condition of the target font features to obtain a target image.

[0073] In the case where it has been determined that modification needs to be performed on the second region of the initial image, and the basis for the modification is not a direct image bitmap file or bitmap data, but target font features in the form of a feature vector, the second region of the initial image is first set as a mask, and the diffusion processing is performed based on a diffusion model when the second region is redrawn. The process of diffusion processing is constrained by the image features of the background, the second region, and the target text presented in the style attribute in the initial image, so as to eliminate the interference of the original image in the second region on the diffusion process and the diffusion result. Finally, after the diffusion processing is completed, the local image outside the second region in the target image obtained is completely the same as the local image in the corresponding region in the initial image; the background of the second region is generated based on the initial image diffusion, and no modification trace is guaranteed in the background part; the modified text in the second region is generated based on the image features of the target text presented in the style attribute, and no modification trace is guaranteed in the foreground part, so that the modification trace between the modified text and the unmodified part is finally eliminated, and the viewing experience of the reader for the modified new image is improved.

[0074] In an optional implementation, in step S132, during the process of obtaining the target image by performing diffusion processing in the second region of the initial image with the target font features as constraints, the second region of the initial image is set as a mask region, and then diffusion processing is performed in the mask region according to the target font features described above, thereby obtaining the target image. That is, in the initial image, the second region is used as the mask region, and diffusion processing is performed in the mask region with the target font features as constraints to obtain the target image. Setting the second region of the initial image as a mask region means setting the original image data in the second region of the initial image as a binary mask. Then, based on the image data outside the mask region and the target font features used as the basis for image restoration within the mask region, image restoration is performed within the mask region. Under the constraints of the image data outside the mask region and the target font features, the background style within the mask region can be restored to be consistent with the background style in the image data outside the mask region. Furthermore, the foreground added within the mask region is text corresponding to the target font features. The target font features are extracted from the text image of the target text displayed according to the style attributes of the text outside the mask region. The content of the text added during image restoration in the mask region is the same as the target text, and the visual effect of the added text is the same as the visual effect of the text outside the mask region. Based on the above data relationships, in the regional image obtained by image restoration in the masked area, the background and foreground correspond to the visual effects of the image outside the masked area. The text content in the foreground is the target text for modification of the initial image. This achieves the modification of the existing text in the initial image, and the modified text is consistent with the overall style of the text in the initial image. It eliminates the modification traces between the modified text and the unmodified parts, and improves the reader's viewing experience of the modified new image.

[0075] In another alternative implementation, in the initial image, a second region is used as the mask region, and the target font features are used as constraints to perform diffusion processing within the mask region to obtain the target image. In this process, noise can be added to the mask region of the initial image, and then the image data outside the mask region and the target font features are used as constraints. While keeping the image data outside the mask region unchanged, the image data within the mask region is denoised to complete the repair and obtain the target image.

[0076] In the initial image, the second region is used as the mask region, and diffusion processing is performed within the mask region based on the target font features as constraints. This process yields the target image, as follows: Figure 6 As shown, the steps may include, but are not limited to, steps S1321-S1323.

[0077] Step S1321: Add noise to the initial image to obtain a noisy image.

[0078] All image data in the initial image determines the actual meaning of the picture presented to the reader when the initial image is displayed, such as the text, patterns, background, etc. in the picture. Adding noise to the initial image, such as adding Gaussian noise, in the noise image obtained by adding noise to the initial image, the noise covers the image data in the initial image, and when the noise image is displayed, the reader may be presented with an image without meaning.

[0079] For the process of adding noise to the initial image to obtain the noise image in step S1321, when adding noise to the initial image, the noise can be added at one time to obtain a noise image composed of pure noise. Adding Gaussian noise to the image is often implemented in image processing, which will not be repeated here.

[0080] Step S1322: de-noising the noise image with the initial image, the mask area and the target font feature as the de-noising control condition to obtain a de-noised image.

[0081] De-noising the noise image with the initial image, the mask area and the target font feature as the de-noising control condition is to de-noise the noise image. The mask area controls the noise image to be processed in different regions during the de-noising process, and the initial image and the target font feature control the processing target of different regions during the de-noising process. In the case that the initial image and the noise image are of the same size, the initial image can be completely overlapped with the noise image, the mask area is determined from the initial image, and the distribution position of the mask area in the noise image can be used as a reference. The image data outside the mask area in the image generated by de-noising the noise image is completely the same as the image data at the corresponding position in the initial image, the background in the mask area is the same as the background of the mask area in the initial image, and the foreground in the mask area is the text generated based on the target font feature, that is, the background of the text outside the mask area in the initial image is used as the background in the mask area to generate text with the same visual effect as the text outside the mask area in the initial image, thereby realizing the modification of the text in the initial image, and the overall style of the modified text is unified with the text in the initial image, eliminating the modification traces between the modified text and the unmodified part, and improving the viewing experience of the reader for the new image after modification.

[0082] After adding noise to the initial image to obtain the noise image, de-noising the noise image with the initial image, the mask area and the target font feature as the de-noising control condition, the target font feature can be injected into the de-noising network of the pre-trained diffusion model as a condition signal in the de-noising process through the cross-attention mechanism; and the noise image is de-noised by the diffusion model, and the de-noising target is to process the noise image so that the image outside the mask area is the same as the image in the mask area of the initial image.

[0083] The diffusion model is a generative model that generates data by gradually adding noise and then learning a denoising process. In the embodiments of the present application, noise is gradually added in the image and then a denoising process is learned to generate a new image. The diffusion model that has completed training and learned the denoising process can generate a new image according to the input data (i.e., the initial image, the mask area, the target font feature, and the noise image). The data input into the diffusion model is either the data basis for generating a new image by denoising, such as the noise image, or the conditions for constraining the denoising process, such as the initial image, the mask area, and the target font feature. The diffusion model has the ability to generate an image with a given picture effect based on the noise image and the constraint conditions, that is, the ability to generate an image in which the overall style of the text in the generated image is consistent with the text in the initial image, and there is no modification trace between the modified text and the unmodified part. This ability is obtained through a corresponding training strategy.

[0084] In an optional implementation, the diffusion model is trained in the following manner: a training set is obtained, the training set including a plurality of training samples, each training sample including an initial sample image, and a mask sample, a text feature sample, and a plurality of noise sample images corresponding to the initial sample image, the text feature sample being a text feature of a region corresponding to the mask sample in the initial sample image; a sample training image is obtained by denoising a pure noise image multiple times under the constraints of the mask sample and the text feature sample according to the noise sample images through an initial diffusion model; and the initial diffusion model is trained to obtain the diffusion model according to a loss of the sample training image and a corresponding sample image, the sample image including the initial sample image and the corresponding noise sample image.

[0085] In this implementation, the diffusion idea and the diffusion model for implementing the diffusion idea are designed according to the analysis of the underlying causes of the technical problems, and the training process of the diffusion model is proposed for obtaining the diffusion model. The training process includes the design of the training samples in the training set, how the specific sample data in each training sample participates in the design of the diffusion model, and how to determine whether the training of the diffusion model is completed, thereby supporting the image processing method in the embodiments of the present application to solve the related technical problems in the image processing process.

[0086] In another optional implementation, the training sample is constructed in the following manner: noise is added to the initial sample image multiple times to obtain a plurality of noise sample images; a mask area is set in the initial sample image to obtain a corresponding mask sample; a text feature sample is obtained by performing feature extraction on the text in the mask area of the initial sample image; and the initial sample image, the corresponding noise sample image, the mask sample, and the text feature sample are taken as a training sample.

[0087] In the process of constructing the training set, for a training sample in the training set, an image with text can be generated as an initial sample image manually or automatically (for example, through a program or a generation process that records and automatically runs a manual operation once), and the position of the text is recorded. Noise is added to the initial sample image multiple times, and the initial sample image has multiple noise sample images. The position of the text in the initial sample image is determined as a corresponding mask region to determine a corresponding mask sample. Feature extraction is performed on the text in the mask region to obtain a corresponding text feature sample. The initial sample image and the corresponding noise sample image, mask sample, and text feature sample are taken as a training sample.

[0088] In an optional embodiment, according to the noise sample image, under the constraints of the mask sample and the text feature sample, the process of obtaining the sample training image by denoising the pure noise image multiple times through the initial diffusion model includes the following steps: determining a mask sample image according to the mask sample, the initial sample image, and the pure noise image; and under the constraint of the diffusion content of the mask region of the mask sample image by the text feature sample, denoising the mask sample image multiple times through the initial diffusion model to obtain the sample training image.

[0089] According to the diffusion idea proposed for the problems of the prior art, the content outside the mask region can be kept unchanged in the diffusion process. Thus, the image directly processed by the diffusion model can be regarded as an image obtained by adding local (i.e., mask region) noise to the initial sample image. The diffusion model not only diffuses the image outside the mask region, but also generates the complete image in the mask region. In the diffusion process, the text is generated in the mask region according to the text feature sample. The finally trained diffusion model can directly eliminate the influence of the original image content in the specified region of the input initial image when generating the image, generate the text with the specified text feature, and highly integrate the image outside the specified region without modification traces, thereby improving the viewing experience of the reader for the modified new image.

[0090] In another optional embodiment, determining the mask sample image according to the mask sample, the initial sample image, and the pure noise image includes the following steps: determining a first sub-sample image outside the mask region corresponding to the mask sample from the initial sample image; determining a second sub-sample image inside the mask region corresponding to the mask sample from the pure noise image; and splicing the first sub-sample image and the second sub-sample image to obtain the mask sample image. On the basis of having determined the mask sample, the initial sample image, and the pure noise image, the respective effective region images of the initial sample image and the pure noise image can be determined according to the mask sample to be spliced to obtain the mask sample image.

[0091] On the basis that various pre-processing of the training samples has been completed, the training set as a whole is divided into multiple times to train the initial diffusion model; according to the training loss of each time, the latest model parameters of the initial diffusion model are determined as the model parameters of the initial diffusion model of the next training or the model parameters of the final diffusion model. In the training process, one training sample can be used for each training, or multiple training samples can be used. After each training, the latest model parameters are determined according to the training loss, which are used as the model parameters of the initial diffusion model of the next training (if the training continues) or the model parameters of the final diffusion model (if it is determined that the training is completed).

[0092] In the process of determining the training loss each time, the loss of one or more sample training images generated in each training process and the corresponding sample image can be determined.

[0093] For one training sample, multiple denoising is required in the training process, and the sample training image obtained by each denoising can be used to determine the training loss with the corresponding sample image. For single training, the loss of one or more sample training images and the corresponding sample image can be used as the training loss of the training. For example, in the case of single training using only one training sample, the loss of one or more sample training images corresponding to the training sample and the corresponding sample image can be used to determine the training loss of the training. In the case of single training using multiple training samples, one or more of the losses of all sample training images corresponding to all training samples and the corresponding sample images can be used as the training loss of the training. It should be understood that if the training loss is determined according to multiple losses, the average of the multiple losses can be used as the training loss.

[0094] An optional implementation of determining the loss in the training process, that is, the process of training the initial diffusion model to obtain the diffusion model according to the loss of the sample training image and the corresponding sample image, can include: determining the similarity of the overall image features to obtain a first loss value of the spatial feature loss according to the sample training image and the corresponding sample image, determining the similarity of the local feature map of the mask region to obtain a second loss value of the mask feature loss, determining the similarity of the overall pixel data to obtain a third loss value of the overall pixel loss, determining the similarity of the local pixel data of the mask region to obtain a fourth loss value of the mask pixel loss, and determining the similarity of the recognized text of the mask region to obtain a fifth loss value of the text loss; determining a comprehensive loss according to the first loss value, the second loss value, the third loss value, the fourth loss value and the fifth loss value; determining whether the initial diffusion model is completed pre-training to obtain the diffusion model according to the comprehensive loss.

[0095] In this implementation, the diffusion model is trained based on a comprehensive loss function composed of multiple losses. The comprehensive loss function composed of multiple losses can ensure that the image generated by the diffusion model is consistent with the user's expectation represented by the input information in each dimension as much as possible, thereby achieving high uniformity of the overall style in the generated image, no modification traces between the modified part and the unmodified part, and improving the viewing experience of the reader on the modified new image.

[0096] Based on the overall description of the training process of the diffusion model above, in this embodiment, the training process of the diffusion model based on the comprehensive loss function composed of all losses described above is described. The diffusion model is trained based on sample images, each sample image including a base image and a plurality of corresponding noisy images. The base image is added with noise multiple times, and the corresponding plurality of noisy images are obtained. In the process of adding noise multiple times, each time of adding noise sequentially defines a different time step, and there is a noisy image with noise based on the initial image. For example, a total of 30 times of adding noise, 30 noisy images are obtained, and each noisy image corresponds to a time step from 1-30 in turn. Each sample image has a corresponding expected modified image, and the expected modified image corresponding to the sample image is obtained by modifying the text in the specified area in the editing mode, and the specified area is used as a mask area in the training stage. The expected modified image is different from the corresponding sample image only in the text content of the specified area, and the background and font style of the image are the same.

[0097] In the process of training the diffusion model, the input is mainly the noisy image x_t at different time steps, the corresponding noise ε, the time step t and the conditional information (z_font, mask). The output is the predicted noise ε_θ or the denoised image x_0_θ. The core of the comprehensive loss function is to supervise the prediction ability of the diffusion model on the noise or the original image. For the loss components of the comprehensive loss function mentioned in the foregoing, they are described as follows: Spatial feature loss (L_latent_mse): directly constrain the similarity between the latent feature z_latent (for example, the effective feature representation of the intermediate layer of the denoising network) predicted by the diffusion model and a target latent feature z_latent_target (for example, the feature extracted from the initial image, used to guide the style): L_latent_mse = ||z_latent - z_latent_target||², where ||||² represents the L2 norm, that is, the mean square error, used to measure the difference between the generated feature and the target feature. The spatial feature loss guides the diffusion model to learn the style of the font in the initial image at the feature level, and ensures that the generated image is consistent with the target style at the high-level feature level.

[0098] Masked latent feature loss (L_masked_latent): Focus on the alignment of latent features in the masked region: L_masked_latent = ||(z_latent - z_latent_target) x mask||2, where z_latent represents the effective feature representation of the intermediate layer output of the denoising network of the diffusion model, z_latent_target represents the target latent feature; mask represents a binary mask, the masked region (the region of the text that needs to be modified) takes 1, and the non-masked region (background) takes 0, x represents element-wise multiplication, only the feature difference of the masked region is calculated, ||||2 represents the L2 norm, i.e. mean square error, which is used to measure the alignment error of the latent features in the masked region. Ensure that the diffusion model pays more attention to the font feature learning of the region that needs to be modified (i.e. the masked region).

[0099] Overall pixel loss (L_pixel_mse): The "clean" image x_0_θ (or x_recon) predicted by the diffusion model, i.e. the image obtained after denoising for a given time step t, is input into a pre-trained decoder (which can be shared with the character decoding model in the character feature extraction model or independent), to obtain the reconstructed image. Calculate the pixel-level mean square error loss between this reconstructed image and the expected modified image x_target: L_pixel_mse = ||x_recon - x_target||2, the expected modified image is the standard image obtained after expected modification, which is determined by the technician who constructs the training sample during the training phase. ||||2 represents the L2 norm, i.e. mean square error, which is used to calculate the pixel-level mean square error between the reconstructed image and the target image. The overall pixel loss provides pixel-level supervision for the prediction results of the diffusion model.

[0100] Masked Pixel Loss (L_masked_pixel): Optimizes pixel-level reconstruction specifically for the masked region: L_masked_pixel = ||(x_recon - x_target) x mask||², which calculates the pixel-level mean squared error loss between the reconstructed image in the masked region and the expected modified image in the masked region, where x_recon represents the image reconstructed by the VAE decoder from the diffusion model's predicted clean image x_0_θ, with the same dimensions as the input image to the diffusion model, and x_target represents the expected modified image. Mask is a binary mask, with 1 for the masked region (text region that needs to be modified) and 0 for the non-masked region (background); x represents element-wise multiplication, only retaining the pixel difference in the masked region and ignoring the error in the non-modified region; || ||² represents the L2 norm (mean squared error), which calculates the pixel difference between the reconstructed image and the target image in the masked region, and the smaller the difference, the more pixel matches in the modified region. Again, pixel-level supervision is emphasized for the quality of the modified region.

[0101] OCR Loss (L_ocr): The key to ensuring the readability of generated text. Input the generated image x_recon or x_0_θ into a pre-trained optical character recognition model to obtain the recognized character class prediction ocr_logits. Calculate the cross-entropy loss between ocr_logits and the true target character label ocr_target (i.e. the character label in the expected modified image that needs to be modified): L_ocr = CrossEntropy(ocr_logits, ocr_target), where CrossEntropy represents the standard cross-entropy loss function, which measures the difference between the predicted probability distribution and the true label; ocr_logits represents the output of the pre-trained OCR model, which is the original score of the generated text belonging to all character categories; ocr_target represents the true character label of the target text (e.g. if the modification target is "deep research", it is the character category label corresponding to this text). This loss directly penalizes the case of recognizing errors in generated text, ensuring the semantic correctness of generated text.

[0102] A comprehensive loss function (L_total) is used to train the diffusion model by combining the above loss terms: L_total = L_latent_mse + L_masked_latent + L_pixel_mse + L_masked_pixel + L_ocr. The standard noise prediction loss of the diffusion model itself (e.g., ||ε - ε_θ||²) is usually a basic component of L_latent_mse or L_pixel_mse or jointly optimized with them. The combination of loss terms can also be a weighted combination to balance the influence of different loss terms.

[0103] In the inference process of the trained diffusion model, the data basis for inference is the noise image x_T, the initial image x_orig, the mask mask, and the target font feature z_font are constraints in the inference process. Among them, the initial image x_orig is mixed according to the mask mask and the non-mask region (1-mask): x_T_masked = x_orig × (1-mask) + x_T × mask. This ensures that the content of the non-modified region in the initial image x_orig remains unchanged from the beginning.

[0104] On the basis of adding noise to the initial image multiple times to obtain a noise image composed of pure noise, the process of denoising the noise image through the diffusion model, according to the order of adding noise, is performed multiple times in the reverse direction. In each denoising process through the diffusion model, the noise after adding noise is predicted according to the corresponding image and the target font feature, and the image features of the initial image and the target font feature are fused through the multi-head attention mechanism. The multiple denoising in the reverse direction is equivalent to iterative denoising from t=T to t=0. For a certain denoising in the iterative denoising process, the denoising network predicts the noise ε_θ of the current step or directly predicts the mean μ_θ (and possibly the variance) of x_{t-1}. x_{t-1} is calculated according to the sampling rule of the diffusion model, such as DDPM (Denoising Diffusion Probabilistic Models) or DDIM (Denoising Diffusion Implicit Models). After calculating x_{t-1}, the mask operation is applied again: x_{t-1} = x_orig× (1 - mask) + x_{t-1} × mask, to ensure that in each denoising step, the non-masked area is forced to restore the content of the original image x_orig, and only the content of the masked area is updated by the diffusion model according to the condition (i.e., the target font feature z_font). In each layer of the denoising network, such as in the cross-attention layer, the target font feature z_font is injected to guide the diffusion model to generate textures with the target font style in the masked area. When the iteration reaches t=0, the output image x_0 is obtained, which is the final result x_modified of the traceless modification. The non-masked area is completely consistent with the input x_orig, and the masked area is replaced by the target text in the z_font style.

[0105] In the process of denoising the noise image by the diffusion model based on the feature injection mechanism to generate an image, the overall input and output are the input initial image x_orig, the binary mask mask that accurately identifies the text area to be modified in the initial image x_orig, and the target font feature z_font provided by the character feature extraction module. The output is the modified image x_modified.

[0106] The process of feature injection can include feature fusion, temporal fusion, multi-head attention, and spatial position enhancement. Feature fusion takes the target font feature z_font as a key condition signal and injects it into the denoising network of the diffusion model through a cross-attention mechanism. An exemplary injection method is to use the features of the intermediate layers of the denoising network as Query, the target font feature z_font (usually after a projection layer) as Key and Value, and perform attention calculation. Temporal fusion means that at each time step t of the reverse denoising process, the diffusion model fuses the state of the current noisy image x_t, the time step embedding t, and the target font feature z_font when predicting noise or calculating the mean. Using a multi-head attention mechanism to fuse image features and font features allows the diffusion model to focus on different subspace information of the target font feature, improving the fusion effect. Combined with learnable position encoding (such as sinusoidal encoding or learnable spatial embedding), the diffusion model is provided with spatial coordinate information for spatial position enhancement to ensure that the generated text is accurately positioned within the mask region.

[0107] The above data generation process between the input and output under the constraint of feature injection can be modeled as: p(x_{t-1}| x_t, z_font, mask) = N(μ_θ(x_t, t, z_font, mask), Σ_θ(x_t, t)). Where x_{t-1} is the image at time step t-1; x_t is the noisy image at time step t; t is the current time step index; z_font is the target font feature; mask is the mask identifying the modified region; μ_θ is the conditional mean predicted by the denoising network; Σ_θ is the variance predicted by the denoising network or preset (usually fixed or related to the time step); N(·) represents a Gaussian distribution; the goal of the diffusion model θ is to learn to predict μ_θ (and sometimes Σ_θ) so that the denoising process can generate the target image under the conditional constraint. The trained diffusion model is the implementation result of this modeling. Based on this modeling idea and based on the diffusion model obtained by the training process described above, the model architecture of the diffusion model is as shown in Figure 7 .

[0108] Step S1323: Determine the image in the initial region of the denoised image as the target image, which is the same as the image in the corresponding region of the initial image, and the image in the generated region has the target font feature, the generated region is the region corresponding to the second region, and the initial region is the region outside the second region.

[0109] Corresponding to the reverse multiple denoising of multiple added noise, in the denoising process, if it is determined that the display requirements of the text in the generated image which is modified are met in the generated image, the denoising can be stopped, and the latest generated image is directly taken as the target image.

[0110] That is, from step S1321 to step S1323, Gaussian noise is added to the initial image x orig to obtain a pure noise image x T. The key reverse process is to gradually denoise from x T under the given condition information (original image, mask and target font feature z font obtained from the feature extraction module), and finally generate the target image x modified. The core task of the model is to generate new text content in the mask identified area according to the condition signal z font under the premise of keeping the area outside the mask (i.e. the part that does not need to be modified) unchanged, so as to conform to the target font style.

[0111] The image processing method in the embodiment of the application can be exemplarily described based on the processing of a specific image. Assuming that a user needs to modify a sentence of text "deep learning" in an initial image to "deep research", on the basis of completing the training of each model described in the foregoing and completing the application development for implementing the image processing method, to achieve this requirement based on the image processing method in the embodiment of the application, the user needs to provide the initial image through the application for implementing the image processing method in the embodiment of the application, and provide the modified target text through the operation (for example, frame selection of the display area of "deep learning") on the initial image or the input text (for example, "modify 'deep learning'"). The application for implementing the embodiment of the application determines the mask area from the initial image according to the frame selection range, or determines the display area of "deep learning" as the mask area through text recognition (for example, OCR) on the initial image. After determining the mask area, the style attribute of the text in the image content outside the mask area is recognized, the target text (i.e. "deep research") is drawn according to the style attribute, and a text image with the display content of "deep research" and the same visual effect as the text outside the mask area in the initial image is obtained. The target font feature is obtained by performing feature extraction on the text image, and the target font feature is represented by a feature vector. Thus, the data preparation required for generating a new image is completed.

[0112] On the basis that the data preparation has been completed, the mask region is filled with the mask outer region of the initial image + random noise to obtain a noise image. Start the iterative denoising (t = T -> 0): at each step t, the model receives the current noise image x t, the time step t, the mask mask and the target font feature. The denoising network in the diffusion model (internally through cross-attention fusion target font features) predicts how to update the mask region. After calculating x t-1, the mask outer region is immediately reset to the corresponding part of the initial image. Under the guidance of the target font feature, the diffusion model learns to gradually generate pixels with “deep research” stroke features in the mask region. At the same time, the L ocr loss (in the training stage) or the integrated OCR verification (optional in the inference stage) ensures that the generated text is recognized as “deep research” rather than other characters or random codes. After T steps of iteration, the modified image is obtained. The position of “deep learning” in the original image has been seamlessly replaced by “deep research”, while the image background, other text or non-textual regions remain completely unchanged, achieving “traceless” modification.

[0113] Figure 8 The comparison presents the comparison effect schematic diagram before and after the modification of two initial images based on the image processing method in the embodiments of the present application. The “CYBER” in the first initial image has been modified to “HELLO” in the corresponding modified target image, and the “non-heritage” in the second initial image has been modified to “flea” in the corresponding modified target image. Obviously, in the modified image, the modified text and the unmodified text are visually uniform, naturally integrated with the background, and the modification trace between the modified text and the unmodified part is eliminated. If it is not Figure 8 The initial image and the target image are indicated in the middle, it is difficult to detect with the naked eye which is the initial image and which is the target image, which obviously improves the viewing experience of the reader for the modified new image.

[0114] The image processing method in the present application starts from inputting an initial image to be modified. The initial image is sent to an image preprocessing module for character recognition and region segmentation, and a mask identifying the character region to be modified is accurately generated. Subsequently, the character feature extraction module starts working. It uses a single-branch variational autoencoder (VAE) encoder specially designed to extract key glyph features from the rendered image of the input character. These extracted glyph features, together with the original image and the mask, are input into the character diffusion generation module. This module is a diffusion model based on the principle of image inpainting. It gradually denoises and generates new character content conforming to the target font style in the region identified by the mask, using the extracted glyph features as a conditional control signal. The entire process ensures that the non-masked region remains unchanged and only the target region is seamlessly replaced with new text, thereby ensuring that the final output image with new character content after modification is a traceless modified image. The above processing process can be referred to as Figure 9 After the target character is determined as '350 pieces of painting felt', the target character can be modified to the specified position in the initial image, and the visual effect of the target character is consistent with that of the unmodified character.

[0115] Overall, in the image processing method, the initial image to be processed is obtained, and the processing region of the initial image and the target character to be presented in the processing region are determined. The style attribute of the character in the region outside the processing region is determined. The region image is generated in the processing region of the initial image through diffusion processing. The content of the region image is the target character with the style attribute of the character in the region outside the processing region. The target character generated by diffusion processing covers the original character in the processing region and has the same style attribute as the character in the initial image that is not modified. The modification of the character in the initial image is realized. The modified character is consistent with the overall style of the character in the initial image. The modification trace between the modified character and the unmodified part is eliminated. The viewing experience of the reader for the new image after modification is improved.

[0116] Please refer to Figure 10 , which is a structural schematic diagram of an image processing device provided by an embodiment of the present application, as Figure 10 shown, the image processing device comprises a data acquisition unit 810, an information extraction unit 820 and a character redrawing unit 830.

[0117] The data acquisition unit 810 is configured to acquire an initial image and processing information, and the processing information includes region information and text information; the information extraction unit 820 is configured to determine a style attribute of text in a first region in the initial image, and the first region is a neighboring region of a second region, and the second region is determined according to the region information; and the text redrawing unit 830 is configured to redraw target text with the style attribute in the second region of the initial image to obtain a target image, and the target text is determined according to the text information.

[0118] On the basis of the above embodiment, the text redrawing unit 830 includes: The feature extraction subunit is configured to perform feature extraction on the target text drawn according to the style attribute to obtain target font features. The feature redrawing subunit is configured to perform diffusion processing on the second region of the initial image with the target font features as a constraint condition to obtain the target image.

[0119] On the basis of the above embodiment, the feature redrawing subunit is specifically configured to perform diffusion processing in a mask region with the target font features as a constraint condition, to obtain the target image.

[0120] On the basis of the above embodiment, the feature redrawing subunit includes: The noise adding module is configured to add noise to the initial image to obtain a noise image. The denoising processing module is configured to perform denoising on the noise image with the initial image, the mask region, and the target font features as denoising control conditions to obtain a denoised image. The target determination module is configured to determine, as the target image, an image in an initial region of the denoised image that is the same as an image in a corresponding region of the initial image and has the target font features in a generated region, and the generated region is a region corresponding to the second region, and the initial region is a region outside the second region.

[0121] On the basis of the above embodiment, the denoising processing module includes: The feature input sub-module is configured to inject the target font features into a denoising network of a pre-trained diffusion model as a conditional signal in a denoising process through a cross-attention mechanism. The diffusion processing sub-module is configured to perform denoising on the noise image under the constraint of the conditional signal through the diffusion model to obtain the denoised image, and an image outside the mask region of the denoised image is the same as an image outside the mask region of the initial image.

[0122] On the basis of the above-mentioned embodiment, the diffusion processing submodule is specifically configured to: obtain an initial denoising image by fusing image features of an initial image in an initial region through a multi-head attention mechanism and fusing target font features in a generated region for initial denoising through a diffusion model; and obtain a corresponding intermediate denoising image by fusing image features of the initial image in the initial region through the multi-head attention mechanism and fusing the target font features in the generated region for at least one time of intermediate denoising through the diffusion model, the input image being a denoising image obtained by denoising the previous time.

[0123] On the basis of the above-mentioned embodiment, the target determination module is specifically configured to: in a case where the total number of initial denoising and intermediate denoising reaches a preset total number of times, take the intermediate denoising image obtained the last time as the target image.

[0124] On the basis of the above-mentioned embodiment, the diffusion model is obtained by training in the following manner: obtain a training set, the training set including a plurality of training samples, each training sample including an initial sample image, a mask sample corresponding to the initial sample image, a text feature sample, and a plurality of noise sample images, the text feature sample being a text feature of a region corresponding to the mask sample in the initial sample image; obtain a sample training image by denoising a pure noise image through an initial diffusion model under the constraints of the mask sample and the text feature sample according to the noise sample image; train the initial diffusion model to obtain the diffusion model according to a loss of the sample training image and a corresponding sample image, the sample image including the initial sample image and the corresponding noise sample image.

[0125] On the basis of the above-mentioned embodiment, the sample training image is obtained by denoising the pure noise image through the initial diffusion model under the constraints of the mask sample and the text feature sample according to the noise sample image, including: determine a mask sample image according to the mask sample, the initial sample image, and the pure noise image; obtain the sample training image by denoising the mask sample image through the initial diffusion model under the constraint of diffusion content of a mask region of the mask sample image by the text feature sample.

[0126] On the basis of the above-mentioned embodiment, the mask sample image is determined according to the mask sample, the initial sample image, and the pure noise image, including: determine a first sub-sample image outside a mask region corresponding to the mask sample from the initial sample image; determine a second sub-sample image within the mask region corresponding to the mask sample from the pure noise image; splice the first sub-sample image and the second sub-sample image to obtain the mask sample image.

[0127] On the basis of the above-mentioned embodiments, the initial diffusion model is trained according to the loss of the sample training image and the corresponding sample image to obtain the diffusion model, including: According to the sample training image and the corresponding sample image, the similarity of the overall image features is determined to obtain a first loss value about the spatial feature loss, the similarity of the local feature map of the mask region is determined to obtain a second loss value about the mask feature loss, the similarity of the overall pixel data is determined to obtain a third loss value about the overall pixel loss, the similarity of the local pixel data of the mask region is determined to obtain a fourth loss value about the mask pixel loss, and the similarity of the recognized text of the mask region is determined to obtain a fifth loss value about the text loss. The comprehensive loss is determined according to the first loss value, the second loss value, the third loss value, the fourth loss value and the fifth loss value. Whether the initial diffusion model is pre-trained to obtain the diffusion model is determined according to the comprehensive loss.

[0128] On the basis of the above-mentioned embodiments, the training sample is constructed by the following method: Noise is added to the initial sample image respectively multiple times to obtain multiple noise sample images; Mask regions are set in the initial sample image respectively to obtain corresponding mask samples; The text in the mask region of the initial sample image is feature extracted to obtain a corresponding text feature sample; The initial sample image and its corresponding noise sample image, mask sample and text feature sample are taken as a training sample. The image processing device provided in the embodiments of the present application is included in an electronic device of a device, and can be used to execute any image processing method provided in the above-mentioned embodiments, has the corresponding functions and beneficial effects.

[0129] It is worth noting that in the above-mentioned embodiments of the image processing device, each unit and module included is only divided according to the function logic, but is not limited to the above-mentioned division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for easy mutual differentiation, and does not limit the protection scope of the present application.

[0130] Figure 11 A structural schematic diagram of an electronic device provided in an embodiment of the present application is shown in FIG. 10. Figure 11 As shown in FIG. 10, the electronic device includes a processor 910 and a memory 920, and in a possible product form of the electronic device, can further include an input device 930, an output device 940 and a communication device 950; the number of processors 910 in the electronic device can be one or more, Figure 11The processor 910 in the electronic device is taken as an example; the processor 910, the memory 920, the input device 930, the output device 940, and the communication device 950 in the electronic device can be connected through a bus or other means, Figure 11 The processor 910 in the electronic device is taken as an example; the processor 910, the memory 920, the input device 930, the output device 940, and the communication device 950 in the electronic device can be connected through a bus or other means,

[0131] The memory 920 can be used to store software programs, computer executable programs, and modules, such as program instructions / modules corresponding to the image processing method in the embodiments of the present application. The processor 910 executes the software programs, instructions, and modules stored in the memory 920, thereby performing various functions and data processing of the electronic device, that is, implementing the image processing method described above.

[0132] The memory 920 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required by at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 920 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state memory device. In some examples, the memory 920 can further include a memory remotely arranged with respect to the processor 910, and these remote memories can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0133] The input device 930 can be used to receive network configuration information. The output device 940 can include a display screen and other electronic devices.

[0134] The above-described electronic device can be used to execute any image processing method, and has corresponding functions and advantages.

[0135] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device and equipment can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0136] In addition, the embodiments of the present application also provide a storage medium containing computer executable instructions, which are used to perform the related operations in the image processing method provided in any embodiment of the present application when executed by a computer processor, and have corresponding functions and advantages.

[0137] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product.

[0138] Accordingly, embodiments of the present application can be embodied in the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, and the like) embodying computer-readable instructions. Embodiments of the present application are described in terms of flowcharts and / or block diagrams in which each block and / or combination of blocks can represent a module, segment, or portion of computer program instructions. Blocks can also represent one or more processes, procedures, or steps performed by one or more computing devices. The computer program instructions can be provided to a processor of the computing device to produce a machine, such that the instructions, which execute via the processor of the computing device, create means for implementing the functions specified in the flowcharts and / or block diagrams. The computer program instructions can also be loaded onto a computing device to cause one or more processors in the computing device to perform the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more processes, procedures, or steps performed by one or more computing devices. The computer program instructions can be provided to a processor of the computing device to produce a machine, such that the instructions, which execute via the processor of the computing device, create means for implementing the functions specified in the flowcharts and / or block diagrams. The computer program instructions can also be loaded onto a computing device to cause one or more processors in the computing device to perform the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more processes, procedures, or steps performed by one or more computing devices. The computer program instructions can be provided to a processor of the computing device to produce a machine, such that the instructions, which execute via the processor of the computing device, create means for implementing the functions specified in the flowcharts and / or block diagrams. The computer program instructions can also be loaded onto a computing device to cause one or more processors in the computing device to perform the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more processes, procedures, or steps performed by one or more computing devices. The computer program instructions can be provided to a processor of the computing device to produce a machine, such that the instructions, which execute via the processor of the computing device, create means for implementing the functions specified in the flowcharts and / or block diagrams. The computer program instructions can also be loaded onto a computing device to cause one or more processors in the computing device to perform the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more processes, procedures, or steps performed by one or more computing devices. The computer program instructions can be provided to a processor of the computing device to produce a machine, such that the instructions, which execute via the processor of the computing device, create means for implementing the functions specified in the flowcharts and / or block diagrams. The computer program instructions can also be loaded onto a computing device to cause one or more processors in the computing device to perform the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more processes, procedures, or steps performed by one or more computing devices. The computer program instructions can be provided to a processor of the computing device to produce a machine, such that the instructions, which execute via the processor of the computing device, create means for implementing the functions specified in the flowcharts and / or block diagrams. The computer program instructions can also be loaded onto a computing device to cause one or more processors in the computing device to perform the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more processes, procedures, or steps performed by one or more computing devices. The computer program instructions can be provided to a processor of the computing device to produce a machine, such that the instructions, which execute via the processor of the computing device, create means for implementing the functions specified in the flowcharts and / or block diagrams. The computer program instructions can also be loaded onto a computing device to cause one or more processors in the computing device to perform the functions specified in the flowcharts and / or block diagrams.

[0139] In one typical configuration, the computing device includes one or more processors, input / output interfaces, network interfaces, and memory. The memory can include non-persistent memory, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory, among others. The memory is an example of computer-readable media.

[0140] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0141] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that processes, methods, articles or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed, or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.

[0142] The above specific embodiments further illustrate the purposes, technical solutions and beneficial effects of the present application. It should be understood that the above is only a specific embodiment of the present application and is not intended to limit the protection scope of the present application. It is particularly pointed out that any modification, equivalent replacement, improvement, etc. made by those skilled in the art within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An image processing method, characterized in that, include: Acquire an initial image and processing information, the processing information including region information and text information; The style attributes of the text in the first region of the initial image are determined, the first region is the associated region of the second region, and the second region is determined based on the region information; The initial image is processed in the second region to generate target text with the style attributes, thereby obtaining a target image. The target text is determined based on the text information.

2. The image processing method according to claim 1, characterized in that, The step of generating target text with the style attributes from the initial image in the second region through diffusion processing to obtain the target image includes: The target font features are obtained by extracting features from the target text drawn according to the style attributes. In the second region of the initial image, diffusion processing is performed with the target font features as constraints to obtain the target image.

3. The image processing method according to claim 2, characterized in that, The process of performing diffusion processing on the second region of the initial image with the target font features as constraints to obtain the target image includes: In the initial image, the second region is used as the mask region, and the target font features are used as constraints to perform diffusion processing within the mask region to obtain the target image.

4. The image processing method according to claim 3, characterized in that, The process involves using the second region as a mask region in the initial image and performing diffusion processing within the mask region using the target font features as constraints to obtain the target image, including: Noise is added to the initial image to obtain a noisy image; The noisy image is denoised using the initial image, mask region, and target font features as denoising control conditions to obtain a denoised image. In the denoised image, the image in the initial region is the same as the image in the corresponding region of the initial image, and the image in the generated region has the target font features is determined as the target image. The generated region is the region corresponding to the second region, and the initial region is the region outside the second region.

5. The image processing method according to claim 4, characterized in that, The step of denoising the noisy image using the initial image, mask region, and target font features as denoising control conditions to obtain a denoised image includes: The target font features are injected into the denoising network of the pre-trained diffusion model through a cross-attention mechanism as a conditional signal in the denoising process; The noisy image is denoised using the diffusion model under the constraint of the conditional signal to obtain a denoised image, wherein the image outside the mask region of the denoised image is the same as the image outside the mask region in the initial image.

6. The image processing method according to claim 5, characterized in that, The step of denoising the noisy image using the diffusion model under the constraint of the conditional signal to obtain a denoised image includes: Using the diffusion model, the image features of the initial image are fused in the initial region through a multi-head attention mechanism, and the target font features are fused in the generated region for initial denoising, resulting in an initial denoised image. The diffusion model fuses the image features of the initial image in the initial region through a multi-head attention mechanism, and fuses the target font features in the generated region to perform at least one intermediate denoising operation to obtain the corresponding intermediate denoised image. The input image is the denoised image obtained from the previous denoising operation.

7. The image processing method according to claim 6, characterized in that, The step of determining the image in the generated region that is identical to the image in the corresponding region of the initial image and possesses the target font features in the denoised image as the target image includes: If the total number of initial denoising and intermediate denoising reaches a preset total number of times, the most recently obtained intermediate denoised image is used as the target image.

8. The image processing method according to any one of claims 5-7, characterized in that, The diffusion model was trained in the following manner: Obtain a training set, which includes multiple training samples. Each training sample includes an initial sample image, a mask sample, a text feature sample, and multiple noise sample images corresponding to the initial sample image. The text feature sample is the text feature of the region corresponding to the mask sample in the initial sample image. Based on the noise sample image, and under the constraints of the mask sample and the text feature sample, the pure noise image is denoised multiple times using an initial diffusion model to obtain the sample training image. The initial diffusion model is trained based on the loss between the training images and the corresponding sample images to obtain the diffusion model. The sample images include the initial sample images and the corresponding noise sample images.

9. The image processing method according to claim 8, characterized in that, The step of obtaining sample training images by performing multiple denoising operations on the pure noise image using an initial diffusion model, under the constraints of the mask sample and text feature sample, based on the noise sample image, includes: The mask sample image is determined based on the mask sample, the initial sample image, and the pure noise image; Under the constraint of the diffusion content of the mask region of the mask sample image by the text feature samples, the mask sample image is denoised multiple times by the initial diffusion model to obtain the sample training image.

10. The image processing method according to claim 9, characterized in that, Determining the mask sample image based on the mask sample, the initial sample image, and the pure noise image includes: A first sub-sample image outside the mask region corresponding to the mask sample is determined from the initial sample image; Determine a second sub-sample image within the mask region corresponding to the mask sample from the pure noise image; The masked sample image is obtained by stitching together the first sub-sample image and the second sub-sample image.

11. The image processing method according to claim 8, characterized in that, The step of training the initial diffusion model based on the loss between the sample training image and the corresponding sample image to obtain the diffusion model includes: Based on the sample training image and the corresponding sample image, the similarity of the overall image features is determined to obtain a first loss value for spatial feature loss, the similarity of the local feature maps of the mask region is determined to obtain a second loss value for mask feature loss, the similarity of the overall pixel data is determined to obtain a third loss value for overall pixel loss, the similarity of the local pixel data of the mask region is determined to obtain a fourth loss value for mask pixel loss, and the similarity of the text recognized in the mask region is determined to obtain a fifth loss value for text loss. The comprehensive loss is determined based on the first loss value, the second loss value, the third loss value, the fourth loss value, and the fifth loss value; The initial diffusion model is determined based on the comprehensive loss to determine whether it has completed pre-training and obtained a diffusion model.

12. The image processing method according to claim 8, characterized in that, The training samples are constructed in the following manner: Noise is added to the initial sample image multiple times to obtain multiple noisy sample images; Mask regions are set in the initial sample image to obtain corresponding mask samples; Feature extraction is performed on the text in the masked region of the initial sample image to obtain the corresponding text feature samples; The initial sample image and its corresponding noise sample image, mask sample, and text feature sample are used as a training sample.

13. An image processing apparatus, characterized in that, include: A data acquisition unit is used to acquire an initial image and processing information, wherein the processing information includes region information and text information; An information extraction unit is used to determine the style attributes of text in a first region of the initial image, wherein the first region is an associated region of the second region, and the second region is determined based on the region information; The text redrawing unit is used to generate a target image by performing diffusion processing on the initial image in the second region to obtain target text with the style attributes, wherein the target text is determined according to the text information.