Image processing method and device, storage medium and electronic equipment

Through the diffusion model generation image processing method, the problem of poor visual fit between new text and background is solved, and image editing with consistent font and light and shadow effects is achieved.

CN120495468APending Publication Date: 2025-08-15ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510495802.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, in image text editing, the visual fit between the new text and the background is poor, resulting in poor visual effects.

Method used

Using the diffusion model, the text area of ​​the to-process image is removed by generating a mask image, and the random noise image and the mask image are input into the diffusion model, and the features of the text to be added are combined as the generation conditions to generate the final image.

Benefits of technology

The visual effects of the text and background to be added to the generated final image are similar to the original image, and the font and light and shadow effects are consistent, which improves the visual fit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495468A_ABST
    Figure CN120495468A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image processing method, and the method comprises the steps: inputting a to-be-processed image in which a to-be-edited character region is removed, a mask image of a to-be-added character, and a random noise image into a pre-trained diffusion model, and enabling the features of the to-be-added character to serve as a generation condition to be injected into the diffusion model, the diffusion model generates a final image under the guide control of the features of the to-be-added characters, the visual effect of the final image is that the to-be-added characters are added in the character area of the to-be-processed image, and the added to-be-added characters are basically the same as the original to-be-processed image in terms of visual effects such as fonts and shadows. And the visual effect of the obtained final image can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to an image processing method, device, storage medium, and electronic device. Background Art

[0002] In today's era of information explosion, the combination of images and text has become an important carrier of information dissemination, widely used in various fields such as advertising, social media, and document management. Therefore, users often need to edit the text content in images, such as modifying the address, date, or event content in a poster image.

[0003] In the prior art, when editing text in an image, it is usually necessary to erase the original text in the image by means of inpainting, and then fit the new text into the image by drawing the text.

[0004] However, when using the above method to process images, each time you edit the text in the image, you need to find a font that is similar to the text in the original image. In addition, since the text is pasted on the background by drawing the text, it is affected by light and shadow, and the visual fit between the text and the background is poor. Figure 1 shown.

[0005] Figure 1 This is a schematic diagram of editing text in an image using existing technology. Figure 1 In the example above, the original image contains the text "ABC" against a background of three simple geometric shapes: a square, a circle, and a triangle. Furthermore, due to the influence of light and shadow, both the text and the background have their own shadows in the lower right corners. To replace the text "ABC" with "DEF" using existing methods, the text "ABC" in the original image must first be erased using inpainting. A font similar to the original "ABC" is then found and used to create the new text "DEF" onto the original image background using this font.

[0006] from Figure 1 The final image shows that even though the font of the new text "DEF" is similar to that of the original text "ABC", the existing technology can only fit the new text "DEF" directly onto the background of the original image. Therefore, the lower right corner of the text "DEF" in the final image does not have a shadow, causing the text "DEF" to be out of tune with the entire image and having a poor visual effect. Summary of the Invention

[0007] The embodiments of this specification provide an image processing method, apparatus, storage medium, and electronic device to partially solve the problems existing in the above-mentioned prior art.

[0008] The embodiments of this specification adopt the following technical solutions:

[0009] This specification provides an image processing method, the method comprising:

[0010] Get the image to be processed and determine the text to be added;

[0011] Determining a text area that needs to be edited in the image to be processed;

[0012] generating a mask image corresponding to the text to be added according to the position of the text area in the image to be processed, and removing the text area from the image to be processed;

[0013] A random noise image is generated, and the random noise image, the mask image, and the image to be processed with the text area removed are input into a pre-trained diffusion model. The features of the text to be added are injected into the diffusion model as generation conditions, so that the diffusion model denoises the random noise image based on the generation conditions to obtain a final image.

[0014] This specification provides an image processing device, the device comprising:

[0015] An acquisition module is used to acquire the image to be processed and determine the text to be added;

[0016] A determination module, configured to determine a text area to be edited in the image to be processed;

[0017] a pre-processing module, configured to generate a mask image corresponding to the text to be added according to the position of the text area in the image to be processed, and remove the text area from the image to be processed;

[0018] A generation module is configured to generate a random noise image, input the random noise image, the mask image, and the image to be processed with the text area removed into a pre-trained diffusion model, and inject the features of the text to be added as generation conditions into the diffusion model, so that the diffusion model denoises the random noise image based on the generation conditions to obtain a final image.

[0019] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned image processing method is implemented.

[0020] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned image processing method is implemented.

[0021] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:

[0022] An embodiment of the present specification discloses an image processing method, which inputs an image to be processed with the text area to be edited removed, a mask image of the text to be added, and a random noise image into a pre-trained diffusion model, and injects the features of the text to be added as generation conditions into the diffusion model, so that the diffusion model generates a final image under the guidance and control of the features of the text to be added. The visual effect of the final image is that the text to be added is added to the text area of the image to be processed, and the added text to be added is basically the same as the original image to be processed in terms of visual effects such as font, light and shadow, which can effectively improve the visual effect of the final image obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:

[0024] Figure 1 A schematic diagram of editing text in an image using existing technology;

[0025] Figure 2 A flowchart of the image processing method provided in the embodiments of this specification;

[0026] Figure 3 An exemplary image to be processed provided in the embodiments of this specification;

[0027] Figure 4A The embodiments of this specification provide Figure 3 Schematic diagram of a mask image generated by the image to be processed;

[0028] Figure 4B The removal provided in the embodiments of this specification Figure 3 The schematic diagram behind the text area in the image to be processed is shown;

[0029] Figure 5 A schematic diagram of the process of generating a final image using a pre-trained diffusion model provided in an embodiment of this specification;

[0030] Figure 6 This is a rendering of the final image obtained in the embodiment of this specification;

[0031] Figure 7 A schematic diagram of a method for training a diffusion model provided in an embodiment of this specification;

[0032] Figure 8A schematic diagram of an image processing device provided in an embodiment of this specification;

[0033] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0034] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0035] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0036] Figure 2 The image processing method flow chart provided in the embodiment of this specification includes the following steps:

[0037] S200: Acquire an image to be processed and determine text to be added.

[0038] In the embodiments of this specification, the Figure 2 The device that processes the image to be processed in the method shown can be any electronic device that can run a diffusion model, such as a personal computer, a mobile phone, a server, etc. Hereinafter, the device that processes the image to be processed is referred to as a processing device.

[0039] The image processing described in the embodiments of this specification refers to the process of replacing the text originally present in a certain area of the image to be processed with the text to be added. Therefore, the processing device first obtains the image to be processed and the text to be added, such as Figure 3 shown.

[0040] Figure 3 The exemplary image to be processed provided in the embodiment of this specification is Figure 1 The same is an image with the text "ABC" and a background of three simple geometric shapes: square, circle and triangle. The subsequent steps are explained using the example of the text to be added "DEF". That is, now we need to add Figure 3 The text "ABC" in the image to be processed is replaced by "DEF".

[0041] S202: Determine a text area that needs to be edited in the image to be processed.

[0042] After the processing device obtains the image to be processed and determines the text to be added, it can first determine the text area that needs to be edited in the image to be processed. Specifically, the processing device can first identify the text in the image to be processed through the optical character recognition (OCR) method, and determine the text line image in the image to be processed based on the identified text, and display the area where each determined text line image is located as a candidate text area to the user, so that the user can select the text area to be edited from the candidate text areas. Then, the processing device can determine the selected candidate text area as the text area to be edited in response to the user's selection operation.

[0043] Continuing with the above example, Figure 3 The area in the dotted box is the text area that needs to be edited in the image to be processed.

[0044] S204: generating a mask image corresponding to the text to be added according to the position of the text area in the image to be processed, and removing the text area from the image to be processed.

[0045] In the embodiment of the present specification, the mask image generated by the processing device is an image of the same size as the image to be processed. The mask image is used to represent the position and size of the text area determined in step S202 in the image to be processed. If the mask image is represented by a matrix, that is, a mask matrix, then the elements of the mask matrix corresponding to the above text area are 1, and the elements corresponding to other areas are 0. If the mask image is represented by an image, then the pixels of the mask image corresponding to the above text area are 255, and the pixels corresponding to other areas are 0, such as Figure 4A shown.

[0046] Figure 4A The embodiments of this specification provide Figure 3 Schematic diagram of the mask image generated by the image to be processed, Figure 4A In the example, the mask image and Figure 3 The images to be processed shown have the same size, and the pixels corresponding to the text area are 255, that is, pure white, and the pixels corresponding to other areas are 0, that is, pure black.

[0047] In addition, the processing device also needs to remove the text area obtained in step S202 from the image to be processed. The removal method can also be performed by using a mask to set all pixels of the text area in the image to be processed to 0, that is, pure black, such as Figure 4B shown.

[0048] S206: Generate a random noise image, input the random noise image, the mask image, and the image to be processed with the text area removed into a pre-trained diffusion model, and inject the features of the text to be added as generation conditions into the diffusion model, so that the diffusion model denoises the random noise image based on the generation conditions to obtain a final image.

[0049] The random noise image described in the embodiments of this specification may be a noise image generated based on any random distribution, and is visually similar to a “snowflake image”.

[0050] like Figure 5 As shown, after generating a random noise image, the processing device can first extract the features of the text to be added, and input the random noise image, the mask image and the image to be processed with the text area removed into a pre-trained diffusion model, and inject the extracted features of the text to be added into the diffusion model as the generation conditions of the diffusion model, so that the diffusion model denoises the random noise image under the guidance and control of the above generation conditions to obtain the final image.

[0051] The visual effect of the final image obtained by the above method is that the text to be added is added to the text area of the image to be processed, and the added text to be added is basically the same as the original image to be processed in terms of font, light and shadow and other visual effects, which can effectively improve the visual effect of the final image. Figure 6 shown. Figure 6 The final image provided in this specification is compared with the Figure 3 The image to be processed and Figure 6 The final image is visible. Figure 6 The final image shown will be Figure 3 The text "ABC" in the image to be processed is replaced by the text "DEF", and the font is basically the same, and, Figure 6 The final image shown also has a shadow on the lower right side of the text "DEF", which is consistent with the Figure 3 The text in the image to be processed and the combined image in other backgrounds have the same visual effect of shadows in the lower right corner, which is more natural and does not feel out of place.

[0052] Furthermore, in the embodiment of the present specification, when extracting the features of the text to be added as the generation conditions from the text to be added through step S206, the extraction can be performed through a pre-trained feature alignment model. Specifically, the processing device can generate a text line image corresponding to the text to be added based on the text to be added, and then extract image features from the text line image through a pre-trained feature alignment model as the features of the text to be added, and then the features of the text to be added can be injected into the diffusion model as the generation conditions. The reason why the embodiment of the present specification injects the image features extracted from the text line image as the generation conditions into the diffusion model, rather than directly extracting character features or text features from the text to be added as the generation conditions into the diffusion model, is because the embodiment of the present specification needs to use a diffusion model to generate images, rather than processing characters of text or processing text. Therefore, extracting image features of the text line image as the generation conditions and injecting them into the diffusion model can improve the quality of the final image generated by the diffusion model.

[0053] When pre-training the above-mentioned feature alignment model, a contrastive learning method can be used to align the character features of the text with the image features of the text line image, so that the image features can well reflect the text in the text line image. Specifically, when training the feature alignment model, a sample text line image can be first obtained, and the sample text contained in the sample text line image can be determined. Then, the feature alignment model to be trained can be used to extract image features from the sample text line image, and the feature alignment model to be trained can be used to extract character features from the sample text. Finally, the model parameters of the feature alignment model to be trained can be adjusted with the training goal of minimizing the difference between the image features extracted from the sample text line image and the character features extracted from the sample text.

[0054] That is, the feature alignment model described in the embodiments of this specification has at least two encoders: a character encoder for extracting character features from text, and an image encoder for extracting image features from text line images. The character encoder can be a pre-trained character encoder. When training the feature alignment model, some or all of the model parameters of the character encoder can be fixed, and the model parameters of the image encoder can be adjusted with the training objective of minimizing the difference between the image features extracted from the sample text line images by the image encoder and the character features extracted from the sample text by the character encoder.

[0055] The feature alignment model trained by the above-mentioned contrastive learning method can extract image features that can accurately express text from text line images.

[0056] Since the embodiment of this specification aims to generate a final image with the same visual effect as the original image to be processed and with the text to be added through the diffusion model, the diffusion model also needs to be trained in advance. The training method is as follows: Figure 7 shown.

[0057] Figure 7 The method for training a diffusion model provided in the embodiment of this specification includes the following steps:

[0058] S700: Acquire a sample image, and determine text contained in the sample image as annotated text.

[0059] In this embodiment, the diffusion model is trained using a self-supervised learning approach. Specifically, the text originally contained in the sample image is used as the text to be subsequently added to the sample image, causing the diffusion model to generate an image. This allows the original sample image to serve as a supervisory signal for training the diffusion model. Therefore, in this embodiment, when training the diffusion model, sample images are obtained, and the text contained in the sample images is used as the annotated text.

[0060] S702: Determine the text area where the annotated text is located in the sample image as an editing area.

[0061] The method for determining the editing area in step S702 is the same as Figure 2 The method of determining various text areas that need to be edited in the image to be processed in step S202 is the same as shown, and will not be repeated here.

[0062] S704: Generate a mask image corresponding to the annotation text as a sample mask image according to the position of the editing area in the sample image, and remove the editing area from the sample image.

[0063] The method of generating a mask image corresponding to the annotation text and removing the editing area in the sample image in step S704 is the same as that of the method of Figure 2 The method of generating the mask image corresponding to the text to be added in step S204 is the same as the method of removing the text area in the image to be processed, and will not be repeated here.

[0064] S706: Generate a random noise image, input the random noise image, the sample mask image, and the sample image with the editing area removed into the diffusion model to be trained, and inject the features of the annotated text as sample generation conditions into the diffusion model to be trained, so that the diffusion model to be trained denoises the random noise image based on the sample generation conditions to obtain a final sample image.

[0065] In the embodiment of this specification, if the feature alignment model is not trained in advance, the character features or text features of the annotated text can be directly extracted and injected into the diffusion model to be trained as sample generation conditions.

[0066] If the above-mentioned feature alignment model has been trained, it is also possible to generate a text line image corresponding to the annotated text based on the annotated text. Through the pre-trained feature alignment model, image features are extracted from the text line image corresponding to the annotated text as the features of the annotated text, and the features of the annotated text are injected into the diffusion model to be trained as sample generation conditions.

[0067] Since the diffusion model is trained using a self-supervised learning method, the diffusion model to be trained needs to denoise the random noise image under the guidance and control of the features of the annotated text to obtain a sample final image with the annotated text in the editing area and the same visual effect as the sample image, and the other areas are the same as the sample image.

[0068] S708: Determine a sample designated area in the sample final image, where a position of the sample designated area in the sample final image is the same as a position of the editing area in the sample image.

[0069] S710: Extracting features of text from the designated area of the sample, and determining a text-image alignment loss value based on the features of the text extracted from the designated area of the sample and the features of the annotated text.

[0070] In the embodiment of this specification, the process of determining the image-text alignment loss value through steps S708 to S710 and the process of determining the diffusion loss value through step S712 are executed in any order and can be executed simultaneously.

[0071] Since, ideally, the final image of the sample output by the diffusion model to be trained should have the annotated text in the editing area, have the same visual effect as the sample image, and have the same other areas as the sample image, therefore, the features of the text extracted from the sample specified area of the sample final image should ideally be the features of the annotated text.

[0072] Therefore, if the above-mentioned feature alignment model has not been trained in advance, the character features or text features of the text can be directly extracted from the sample specified area of the sample final image as the features of the text extracted from the sample specified area of the sample final image. According to the difference between the features of the text extracted from the sample specified area of the sample final image and the features of the annotated text, the image-text alignment loss value can be determined.

[0073] If the above-mentioned feature alignment model has been trained in advance, the pre-trained feature alignment model can be used to extract image features from the specified area of the sample as the features of the text extracted from the specified area of the sample, and the image-text alignment loss value can be determined based on the difference between the features of the text extracted from the specified area of the sample and the features of the annotated text.

[0074] Among them, the greater the difference between the features of the text extracted from the specified area of the sample and the features of the annotated text, the greater the image-text alignment loss value; the smaller the difference between the features of the text extracted from the specified area of the sample and the features of the annotated text, the smaller the image-text alignment loss value.

[0075] S712: Determine a diffusion loss value according to the sample final image and the sample image.

[0076] In the embodiment of the present specification, the denoising process performed on the random noise image by the diffusion model to be trained based on the sample generation conditions in step S706 includes multiple denoising cycles. The specific number of denoising cycles can be preset, for example, set to 10,000 cycles. (Similarly, the denoising process performed on the random noise image by the diffusion model based on the generation conditions in step S206 also includes multiple denoising cycles.) Therefore, when determining the diffusion loss value, the noise distribution of the noise to be removed in the random noise image can be determined based on the sample image and the random noise image. Based on the noise distribution, the noise distribution required for each denoising cycle is estimated. The diffusion loss value is determined based on the noise distribution actually denoised by the diffusion model to be trained based on the sample generation conditions each time the random noise image is denoised, as well as the estimated noise distribution required for each denoising cycle.

[0077] Among them, the noise distribution of the random noise image actually denoised each time the diffusion model to be trained is based on the sample generation conditions can be determined according to the difference between the images obtained before and after each actual denoising of the diffusion model to be trained, and the image obtained after the last denoising is the final image of the sample.

[0078] S714: Determine a final loss value according to the image-text alignment loss value and the diffusion loss value.

[0079] In the embodiment of the present specification, the image-text alignment loss value determined in step S710 and the diffusion loss value determined in step S712 may be weighted according to a preset weight to obtain a final loss value.

[0080] S716: Adjust the model parameters of the diffusion model to be trained with reducing the final loss value as a training goal.

[0081] With the goal of reducing the final loss value, the model parameters of the diffusion model to be trained are adjusted until the diffusion model to be trained converges, and the trained diffusion model can be obtained. Before the diffusion model to be trained converges, it can also be iterated repeatedly. Figure 7 The training process shown is carried out until the diffusion model to be trained converges.

[0082] pass Figure 7 The diffusion model obtained by training with the training method shown can be applied to Figure 2 In the image processing method shown, a final image is generated in which text to be added is added to the text area of the image to be processed, and the added text to be added is basically the same as the original image to be processed in terms of visual effects such as font, light and shadow.

[0083] It should be noted that the use of Figure 7 The device for training the diffusion model to be trained in the method shown can be the above-mentioned processing device, or other electronic devices used for training diffusion models but different from the above-mentioned processing device, and the embodiments of this specification do not limit this.

[0084] In addition, in the embodiments of this specification, Figure 2 After the final image is obtained by the method shown, the image quality of the generated final image can also be evaluated. Specifically, the processing device can determine a designated area in the final image, and the position of the designated area in the final image is the same as the position of the text area in the image to be processed; extract image features from the designated area through a pre-trained feature alignment model, and extract character features from the text to be added through a pre-trained feature alignment model, and determine the quality characterization value for characterizing the quality of the final image based on the similarity between the image features extracted from the designated area and the character features extracted from the text to be added. The higher the similarity between the image features extracted from the designated area and the character features extracted from the text to be added, the larger the quality characterization value, indicating that the quality of the final image is better; the lower the similarity between the image features extracted from the designated area and the character features extracted from the text to be added, the smaller the quality characterization value, indicating that the quality of the final image is worse. If the quality characterization value of the final image is lower than the preset threshold, the following method can be repeatedly adopted: Figure 2 The image processing method shown reprocesses the image to be processed until a final image is obtained whose quality representation value is not lower than a preset threshold.

[0085] Furthermore, a pre-trained evaluation model can be used to evaluate the quality of the final image. Specifically, the image to be processed and the final image can be input into the evaluation model, so that the evaluation model outputs the similarity between the two in semantics and visual style, and the quality of the final image is determined based on the similarity in semantics and visual style. The higher the similarity in semantics and visual style, the better the quality of the final image, and vice versa. This is because the similarity between the image features extracted from the specified area and the character features extracted from the text to be added using the above-mentioned feature alignment model can only characterize the similarity in font style between the text to be added in the final image and the text in the original image to be processed, while the evaluation model can evaluate the similarity in semantics and visual style between the final image and the original image to be processed as a whole. Therefore, the similarity obtained using the feature alignment model and the similarity obtained using the evaluation model can be combined to obtain the final similarity between the final image and the original image to be processed, and the quality of the final image can be determined based on the final similarity.

[0086] The above is an image processing method provided in an embodiment of this specification. Based on the same idea, this specification also provides corresponding devices, storage media and electronic devices.

[0087] Figure 8 A schematic diagram of an image processing device provided in an embodiment of this specification, the device comprising:

[0088] The acquisition module 801 is used to acquire the image to be processed and determine the text to be added;

[0089] A determination module 802 is configured to determine a text area that needs to be edited in the image to be processed;

[0090] A pre-processing module 803 is configured to generate a mask image corresponding to the text to be added according to the position of the text area in the image to be processed, and remove the text area from the image to be processed;

[0091] The generation module 804 is used to generate a random noise image, input the random noise image, the mask image, and the image to be processed with the text area removed into a pre-trained diffusion model, and inject the features of the text to be added as generation conditions into the diffusion model, so that the diffusion model denoises the random noise image based on the generation conditions to obtain a final image.

[0092] Optionally, the generation module 804 is specifically used to generate a text line image corresponding to the text to be added based on the text to be added; extract image features from the text line image as features of the text to be added through a pre-trained feature alignment model; and inject the features of the text to be added as generation conditions into the diffusion model.

[0093] Optionally, the device further comprises:

[0094] The training module 805 is used to pre-acquire a sample text line image and determine the sample text contained in the sample text line image; extract image features from the sample text line image through the feature alignment model to be trained, and extract character features from the sample text through the feature alignment model to be trained; and adjust the model parameters of the feature alignment model to be trained with the training goal of minimizing the difference between the image features extracted from the sample text line image and the character features extracted from the sample text.

[0095] Optionally, the training module 805 is further used to pre-acquire a sample image, determine the text contained in the sample image as the annotated text; determine the text area in the sample image where the annotated text is located as the editing area; generate a mask image corresponding to the annotated text according to the position of the editing area in the sample image, as a sample mask image, and remove the editing area from the sample image; generate a random noise image, input the random noise image, the sample mask image and the sample image with the editing area removed into the diffusion model to be trained, and inject the features of the annotated text as sample generation conditions into the diffusion model to be trained, so that the diffusion model to be trained is based on The sample generation condition denoises the random noise image to obtain a sample final image; determines a sample designated area in the sample final image, wherein the position of the sample designated area in the sample final image is the same as the position of the editing area in the sample image; extracts features of text from the sample designated area, and determines a text-image alignment loss value based on the features of the text extracted from the sample designated area and the features of the annotated text; and determines a diffusion loss value based on the sample final image and the sample image; determines a final loss value based on the text-image alignment loss value and the diffusion loss value; and adjusts the model parameters of the diffusion model to be trained with reducing the final loss value as a training goal.

[0096] Optionally, the training module 805 is specifically configured to extract image features from the sample designated area using the pre-trained feature alignment model as features of text extracted from the sample designated area.

[0097] Optionally, the denoising performed on the random noise image by the diffusion model to be trained based on the sample generation condition includes multiple denoising steps;

[0098] The training module 805 is specifically used to determine the noise distribution of the noise that needs to be removed in the random noise image based on the sample image and the random noise image; estimate the noise distribution that needs to be denoised each time based on the noise distribution; and determine the diffusion loss value based on the noise distribution of the random noise image denoised each time based on the sample generation conditions of the diffusion model to be trained and the estimated noise distribution that needs to be denoised each time.

[0099] Optionally, the device further comprises:

[0100] The evaluation module 806 is used to determine a designated area in the final image, where the position of the designated area in the final image is the same as the position of the text area in the image to be processed; extract image features from the designated area through the pre-trained feature alignment model, and extract character features from the text to be added through the pre-trained feature alignment model; and determine a quality characterization value for characterizing the quality of the final image based on the similarity between the image features extracted from the designated area and the character features extracted from the text to be added.

[0101] This specification also provides a computer-readable storage medium, wherein the storage medium stores a computer program, which can be used to perform the above-mentioned Figure 2 Provides image processing methods.

[0102] based on Figure 2 The image processing method shown in the embodiment of this specification also provides Figure 9 The structural diagram of the electronic device shown in FIG. Figure 9 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 2 The image processing method.

[0103] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. An image processing method, comprising: Get the image to be processed and determine the text to be added; Determining a text area that needs to be edited in the image to be processed; generating a mask image corresponding to the text to be added according to the position of the text area in the image to be processed, and removing the text area from the image to be processed; A random noise image is generated, and the random noise image, the mask image, and the image to be processed with the text area removed are input into a pre-trained diffusion model. The features of the text to be added are injected into the diffusion model as generation conditions, so that the diffusion model denoises the random noise image based on the generation conditions to obtain a final image.

2. The method according to claim 1, wherein the features of the text to be added are injected into the diffusion model as generation conditions, specifically comprising: Generating a text line image corresponding to the text to be added according to the text to be added; Extracting image features from the text line image using a pre-trained feature alignment model as features of the text to be added; The features of the text to be added are injected into the diffusion model as generation conditions.

3. The method according to claim 2, wherein the feature alignment model is pre-trained, specifically comprising: Acquire a sample text line image, and determine the sample text contained in the sample text line image; Extracting image features from the sample text line image using a feature alignment model to be trained, and extracting character features from the sample text using the feature alignment model to be trained; The model parameters of the feature alignment model to be trained are adjusted with minimizing the difference between the image features extracted from the sample text line image and the character features extracted from the sample text as a training goal.

4. The method according to claim 3, wherein the diffusion model is pre-trained, specifically comprising: Acquire a sample image, and determine text contained in the sample image as annotated text; Determine a text area in the sample image where the annotated text is located as an editing area; generating a mask image corresponding to the annotated text according to the position of the edit area in the sample image as a sample mask image, and removing the edit area from the sample image; generating a random noise image, inputting the random noise image, the sample mask image, and the sample image with the edited area removed into a diffusion model to be trained, and injecting the features of the annotated text into the diffusion model to be trained as sample generation conditions, so that the diffusion model to be trained denoises the random noise image based on the sample generation conditions to obtain a final sample image; determining a sample designated area in the sample final image, wherein a position of the sample designated area in the sample final image is the same as a position of the editing area in the sample image; Extracting features of text from the designated area of the sample, determining an image-text alignment loss value based on the features of the text extracted from the designated area of the sample and the features of the annotated text; and determining a diffusion loss value based on the final image of the sample and the sample image; Determine a final loss value according to the image-text alignment loss value and the diffusion loss value; Taking reducing the final loss value as a training goal, the model parameters of the diffusion model to be trained are adjusted.

5. The method according to claim 4, wherein extracting features of text from the designated area of the sample comprises: The image features are extracted from the sample designated area by using the pre-trained feature alignment model as the features of the text extracted from the sample designated area.

6. The method according to claim 4, wherein the denoising performed on the random noise image by the diffusion model to be trained based on the sample generation condition comprises multiple denoising steps; Determining a diffusion loss value according to the sample final image and the sample image specifically includes: Determining, according to the sample image and the random noise image, a noise distribution of noise to be removed in the random noise image; According to the noise distribution, estimating the noise distribution that needs to be denoised each time; The diffusion loss value is determined according to the noise distribution of each denoising operation on the random noise image by the diffusion model to be trained based on the sample generation condition and the estimated noise distribution to be denoised each time.

7. The method of claim 3, further comprising: Determine a designated area in the final image, where a position of the designated area in the final image is the same as a position of the text area in the image to be processed; Extracting image features from the designated area using the pre-trained feature alignment model, and extracting character features from the text to be added using the pre-trained feature alignment model; A quality characterization value for characterizing the quality of the final image is determined based on the similarity between the image features extracted from the designated area and the character features extracted from the text to be added.

8. An image processing device, comprising: An acquisition module is used to acquire the image to be processed and determine the text to be added; A determination module, configured to determine a text area to be edited in the image to be processed; a pre-processing module, configured to generate a mask image corresponding to the text to be added according to the position of the text area in the image to be processed, and remove the text area from the image to be processed; A generation module is configured to generate a random noise image, input the random noise image, the mask image, and the image to be processed with the text area removed into a pre-trained diffusion model, and inject the features of the text to be added as generation conditions into the diffusion model, so that the diffusion model denoises the random noise image based on the generation conditions to obtain a final image.

9. A computer-readable storage medium storing a computer program, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by a processor.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.