Image processing method and device and electronic equipment
By identifying and updating the area to be corrected for the output of the image generation model, the problem of mismatch in the text-to-image generation model is solved, and the accuracy and accuracy of the image content are improved, and the robustness of the model is enhanced.
Patent Information
- Application Number
- CN202510421783.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-08-12
AI Technical Summary
The image content output from the existing text to image generation model does not match the text description, resulting in low image content accuracy and problems such as missing, incorrect or redundant content.
Through the image processing method, the area to be corrected in the initial image and its corresponding reference text are determined, and the image content of the area to be corrected is updated using the reference text to be corrected to match the text description, including identifying and correcting typesetting errors, missing or redundant text.
It improves the rendering accuracy of image content, ensures image quality, and enhances the accuracy of text rendering, significantly improving the robustness of image generation model.
Smart Images

Figure CN120472023A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an image processing method, device and electronic equipment. Background Art
[0002] In related technologies, a text description can be input into a text-to-image generation model, which outputs an image corresponding to the text description. This text-to-image generation model can be implemented using a stable diffusion model, a flux model, or other similar models. The image output by the model may not match the text description, for example, due to missing content, incorrect content, or redundant content, resulting in low accuracy of the image content. Summary of the Invention
[0003] In view of this, an object of the present invention is to provide an image processing method, apparatus, and electronic device to improve the rendering accuracy of image content.
[0004] In a first aspect, an embodiment of the present invention provides an image processing method, which includes: inputting a text description into a preset image generation model to output an initial image; based on the text description, determining an area to be corrected in the initial image, and a reference text corresponding to the area to be corrected; wherein the image content in the area to be corrected does not match the reference text corresponding to the area to be corrected; the text description includes at least one reference text; based on the reference text, updating the image content in the area to be corrected so that the image content matches the reference text.
[0005] In a second aspect, an embodiment of the present invention provides an image processing device, which includes: a text description input module for inputting a text description into a preset image generation model and outputting an initial image; a determination module for determining an area to be corrected in the initial image and a reference text corresponding to the area to be corrected based on the text description; wherein the image content in the area to be corrected does not match the reference text corresponding to the area to be corrected; the text description includes at least one reference text; and an image content update module for updating the image content in the area to be corrected based on the reference text so that the image content matches the reference text.
[0006] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-mentioned image processing method.
[0007] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned image processing method.
[0008] The embodiments of the present invention bring the following beneficial effects:
[0009] The above-mentioned image processing method, device and electronic device input a text description into a preset image generation model and output an initial image; based on the text description, determine the area to be corrected in the initial image and the reference text corresponding to the area to be corrected; wherein the image content in the area to be corrected does not match the reference text corresponding to the area to be corrected; the text description includes at least one reference text; based on the reference text, update the image content in the area to be corrected so that the image content matches the reference text.
[0010] In this method, after the model outputs the initial image, it determines the area to be corrected and the corresponding benchmark text from the initial image based on the text description, and then updates the image content in the area to be corrected according to the benchmark text, so that the updated image content matches the benchmark text, while ensuring the image quality and improving the rendering accuracy of the image content.
[0011] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0012] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 A flowchart of an image generation method provided by an embodiment of the present invention;
[0015] Figure 2 A schematic diagram of an initial image, a region to be corrected, and an updated image provided by an embodiment of the present invention;
[0016] Figure 3 A schematic diagram of another initial image, a region to be corrected, and an updated image provided by an embodiment of the present invention;
[0017] Figure 4 A schematic structural diagram of an image generating device provided by an embodiment of the present invention;
[0018] Figure 5 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.
[0020] To facilitate understanding, the terms involved in the embodiments of the present invention are first explained.
[0021] 1. Text-to-image generation model: A text-to-image generation model is a model that can generate images based on text descriptions, such as the Stable Diffusion model and the Flux model.
[0022] 2. Typographical errors: Typographical errors refer to errors in the text in the image, such as typos, missing words, extra words, etc.
[0023] 3. Post-processing: Post-processing refers to the step of processing the image output by the model.
[0024] 4.OCR: The full name is Optical Character Recognition, which is translated as optical character recognition. It is a technology that converts text in images into text.
[0025] 5. Image restoration: Image restoration refers to the technology of repairing missing or damaged parts of an image.
[0026] 6. Text editing model: A text editing model refers to a model that can edit text, such as AnyText.
[0027] 7. Visual Language Model: A visual language model refers to a model that can understand images and text, such as GPT-4.
[0028] 8. Levenshtein distance: Levenshtein distance, also known as edit distance, is a method used to measure the difference between two strings. It is defined as the minimum number of single-character editing operations required to transform one string into another. These editing operations include inserting, deleting, or replacing characters.
[0029] 9. Pre-trained model: A pre-trained model refers to a model that has been pre-trained on a large amount of data.
[0030] In related technologies, a text description can be input into a text-to-image generation model, which then outputs an image corresponding to the text description. This text-to-image generation model can be implemented using a Stable Diffusion model, a Flux model, etc. Currently, there are several approaches to improve the accuracy and readability of the images output by the model.
[0031] In one approach, the text-to-image generation model itself is improved. For example, the model's text rendering capability is improved by increasing training data, improving the text encoder, or adding OCR loss. Some research efforts are devoted to designing more sophisticated text encoders that can learn glyph-level information from the training corpus, thereby rendering text more accurately. Alternatively, some methods add conditional inputs to the text-to-image generation model, such as text position information and glyph masks. However, such methods usually require a large amount of training data, and the visual effect of the images output by the model may be poor.
[0032] In another approach, an external model is used to improve the images output by the text-to-image generation model. For example, some researchers have proposed using a pre-trained model to process text errors in the output images. The processing steps are usually to first detect text errors in the image, then erase the erroneous text, and then generate a new text box to supplement the missing text, and finally correct spelling errors in the rendered text. However, this type of method has low accuracy and efficiency in processing images when encountering complex text errors.
[0033] In summary, existing methods cannot take into account both the accuracy and visual effects of the images output by the model. The images output by the model may not match the text description, for example, there may be missing content, incorrect content, redundant content, etc., resulting in low accuracy of the output image.
[0034] Based on the above, the embodiments of the present invention provide an image processing method, device, and electronic device that can post-process images output by a model.
[0035] To facilitate understanding of this embodiment, an image processing method disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, this method includes the following steps:
[0036] Step S102: input the text description into a preset image generation model and output an initial image;
[0037] The text description is usually a complete paragraph, but can also be multiple words or phrases, and can be in English, Chinese or other languages.
[0038] The preset image generation model can be a Stable Diffusion model, a Flux model, or other models. A text description is input into the preset image generation model, which then outputs an initial image based on the text description. At least some of the image content in this initial image matches the text description. However, due to the influence of the image generation model, the initial image may contain content that does not match the text description, or parts of the text description may not be reflected in the initial image.
[0039] It should be noted that the image processing method provided in this embodiment belongs to an image post-processing method, that is, further processing of the image output by the image generation model. This method is independent of the image generation model and does not depend on a specific type of image generation model. This method can process images output by any type of image generation model.
[0040] Step S104: determining, based on the text description, a region to be corrected in the initial image and a reference text corresponding to the region to be corrected; wherein the image content in the region to be corrected does not match the reference text corresponding to the region to be corrected; and the text description includes at least one reference text;
[0041] The reference text is included in the text description. The reference text may be one or more words in the text description, or one or more sentences. For example, one reference text corresponds to one word, such as "10%", "2014", "discount", etc.
[0042] In actual implementation, the initial image can be checked based on the text description to see if it accurately and completely reflects all the content in the text description; if not, the initial image needs to be corrected. The benchmark text can be understood as the text that is not accurately reflected in the initial image, or the text that is omitted. The area to be corrected can be understood as the area where the content corresponding to the benchmark text needs to be reflected. In the initial image, the image content in the area to be corrected does not accurately reflect the benchmark text, or does not reflect the benchmark text, that is, the benchmark text is omitted.
[0043] In the text description, there may be some text that needs to be reflected in the image. In this case, the scene text detection model can be used to detect the initial image and output at least one text area, which contains the initial text in the initial image. The initial text is then compared with the text in the text description that needs to be reflected in the image. Based on the comparison result, the area to be corrected is determined in the initial image. The text in the area to be corrected does not match the baseline text corresponding to the area to be corrected, such as text errors, text omissions, etc.
[0044] In one example, this embodiment can handle text layout errors in an image. For example, the aforementioned area to be corrected may contain typos, missing words, extra words and other layout errors. These layout errors in the area to be corrected can be corrected through this embodiment.
[0045] Step S106 : Based on the reference text, update the image content in the area to be corrected so that the image content matches the reference text.
[0046] The process of updating the image content in the area to be corrected can also be implemented using the aforementioned image generation model; the reference text and the area location identifier of the area to be corrected are input into the model, and the image content in the area to be corrected is modified based on the reference text through the model.
[0047] For example, if the text in the area to be corrected is different from the reference text, the image content of the area to be corrected is deleted, and then updated content is generated based on the image content of the adjacent area to be corrected and the reference text, and the updated content is used to fill the area to be corrected, and the updated image content matches the reference text.
[0048] If the baseline text corresponding to the area to be corrected is empty, then the image content of the area to be corrected is deleted, and then updated content is generated based on the image content of the adjacent area to be corrected, and the updated content is used to fill the area to be corrected. There is no text in the updated image content; if there is missing text in the text description corresponding to the area to be corrected relative to the text description, then the missing text is generated in the area to be corrected.
[0049] The above-mentioned image processing method, device and electronic device input a text description into a preset image generation model and output an initial image; based on the text description, determine the area to be corrected in the initial image and the reference text corresponding to the area to be corrected; wherein the image content in the area to be corrected does not match the reference text corresponding to the area to be corrected; the text description includes at least one reference text; based on the reference text, update the image content in the area to be corrected so that the image content matches the reference text.
[0050] In this method, after the model outputs the initial image, it determines the area to be corrected and the corresponding benchmark text from the initial image based on the text description, and then updates the image content in the area to be corrected according to the benchmark text, so that the updated image content matches the benchmark text, while ensuring the image quality and improving the rendering accuracy of the image content.
[0051] The following embodiments provide specific implementations of determining the area to be corrected in the initial image.
[0052] In one approach, text detection is performed on an initial image to obtain detected initial text; and based on the initial text and the text description, an area to be corrected in the initial image is determined.
[0053] The initial text is the text contained in the initial image. The initial text may include one or more. The initial text may be a word, a phrase, or a sentence. For example, each initial text represents a word, such as if, 100%, hamburger, and so on.
[0054] Text detection can be achieved by a scene text detection model, which detects text regions in the initial image. Specifically, the initial image is input into the scene text detection model, and at least one text region is output. The text region is usually a polygonal region, such as a rectangular region, a triangular region, a hexagonal region, etc. The text region contains the initial text in the initial image.
[0055] After obtaining at least one text region, the scene text recognition model is used to identify the initial text in at least one text region. Parse the text area Convert the initial text in to a text string; Where i represents the i-th text region, which is used as an index variable to traverse the initial text detected in the initial image. Indicates the total number of regions.
[0056] Furthermore, by comparing the initial text with the text description, typesetting errors in the initial image can be identified, thereby determining the area to be corrected in the initial image.
[0057] In an optional way, determine the initial text The similarity between the initial text and the reference text W in the text description is obtained, and the highest similarity is obtained from them; according to the value of the highest similarity and the threshold range where the highest similarity lies, the erroneous text, correct text or redundant text is determined in the initial text, or the missing text is determined in the reference text.
[0058] In a specific example, target text that does not match the text description is determined from the initial text, and the text area occupied by the target text in the initial image is determined as the area to be corrected.
[0059] After comparing the initial text and the text description, the target text that does not match the text description is determined from the initial text based on the text description; then the text area occupied by the target text is determined in the initial image, and the text area occupied by the target text is determined as the area to be corrected.
[0060] The above-mentioned target text includes: erroneous text relative to the text description; wherein the similarity between the erroneous text and part of the reference text in the text description is greater than a preset first similarity threshold and less than a preset second similarity threshold; the first similarity threshold is less than the second similarity threshold; or, the target text includes: redundant text relative to the text description; wherein the redundant text is text that does not exist in the text description.
[0061] When the target text specifically includes erroneous text relative to the text description, the similarity between the character string of the erroneous text and the character string of part of the reference text in the text description is greater than a preset first similarity threshold and less than a preset second similarity threshold; it can be understood that the erroneous text has a high similarity with a certain reference text in the text description, but is not exactly the same; in one example, the first similarity threshold can be 0% and the second similarity threshold can be 100%.
[0062] When the target text specifically includes redundant text relative to the text description, the redundant text string is a text string that does not exist in the text description. The redundant text can be understood as having a low similarity with all texts in the text description, such as all zeros.
[0063] In a specific example, based on the initial text, missing text is determined from the text description; wherein the missing text is text that does not exist in the initial text; based on the image content of the initial image, the area to be displayed of the missing text is determined from the initial image, and the area to be displayed is determined as the area to be corrected in the initial image.
[0064] After comparing the initial text and the text description, the missing text is determined from the text description based on the initial text, and the character string of the missing text is the text character string that does not exist in the initial text; then the layout is replanned in the initial image, and the area to be displayed of the missing text is determined based on the image content of the initial image, and the area to be displayed is determined as the area to be corrected in the initial image.
[0065] Furthermore, the layout information of the initial text in the initial image is obtained; and based on the background content and the layout information of the initial image, the area to be displayed of the missing text is determined from the initial image.
[0066] The layout information is usually expressed in the form of a bounding box. Specifically, the bounding box layout information of the initial text in the initial image can be obtained; the bounding box can be a rectangular bounding box, and the layout information can include four coordinate values of width, height, left margin and top margin. These coordinate values are used to determine the position and size of the text area in the initial image.
[0067] First, after the scene text detection model outputs at least one polygonal text region, the coordinates of the vertices of these text regions are obtained. Among all the vertices in the text region, the maximum and minimum horizontal coordinate values, as well as the maximum and minimum vertical coordinate values, are obtained. For example:
[0068] minX=min(x1,x2,...,xn)
[0069] maxX=max(x1,x2,...,xn)
[0070] minY=min(y1,y2,...,yn)
[0071] maxY=max(y1,y2,...,yn)
[0072] Using the range calculated based on all vertices of the text area, a bounding box is constructed with the coordinates of the upper left corner (minX, minY), the coordinates of the lower right corner (maxX, maxY), the width of maxX-minX, and the height of maxY-minY, ensuring that the bounding box is within the boundaries of the initial image.
[0073] The background content and layout information of the initial image can then be input into the GPT-4 model, and the visual language model can be used to plan the layout of the missing text, and the area to be displayed of the missing text can be determined from the initial image so that the area to be displayed is coordinated with the visual elements of the adjacent areas.
[0074] In a specific example, the highest similarity between the initial text and the benchmark text in the text description is obtained; if the highest similarity is greater than a preset first similarity threshold and less than a preset second similarity threshold, the initial text is determined to be an incorrect text; wherein the first similarity threshold is less than the second similarity threshold; if the highest similarity is greater than or equal to the preset second similarity threshold, the initial text is determined to be a correct text.
[0075] The similarity between the initial text and the reference text in the text description can be calculated based on the character string of the initial text and the character string of the reference text in the text description, and then the highest similarity is obtained therefrom; when the highest similarity is greater than a preset first similarity threshold and less than a preset second similarity threshold, an error type is obtained indicating that the initial text is an incorrect text; when the highest similarity is greater than or equal to the preset second similarity threshold, the initial text is determined to be a correct text.
[0076] In actual implementation, the similarity between any initial text and any benchmark text can be calculated; for each initial text, the similarity between the initial text and each benchmark text is obtained, and the highest similarity is obtained. If the highest similarity is greater than a preset first similarity threshold and less than a preset second similarity threshold, the initial text is an erroneous text. For example, the first similarity threshold can be 0%, and the second similarity threshold can be 100%. It can be understood that if the initial text has a certain similarity with one of the benchmark texts, but no similarity with the other benchmark texts; if the benchmark text with a certain similarity is not completely identical to the initial text, the initial text can be determined to be an erroneous text.
[0077] When the highest similarity is greater than or equal to a preset second similarity threshold, such as when the second similarity threshold is 100%, the highest similarity is also 100%; or, when the second similarity threshold is 90%, the highest similarity is also 95% or 100%; in this case, it means that the initial text is exactly the same as one of the reference texts. At this time, it can be determined that the initial text is the correct text and no correction is required.
[0078] In a specific example, if the highest similarity is less than a preset third similarity threshold, the initial text is determined to be redundant text; wherein the third similarity threshold is less than or equal to the first similarity threshold.
[0079] The third similarity threshold can be equal to the first similarity threshold, for example, both are 0; the third similarity threshold can also be less than the first similarity threshold, for example, the first similarity threshold is 10% and the third similarity threshold is 0. When the highest similarity is less than the third similarity threshold, it can be understood that the similarity between the initial text and any reference text in the text description is very low, and in this case, it can be determined that the initial text is redundant text. When the highest similarity is less than the preset third similarity threshold, the error type of the initial text being redundant text is obtained.
[0080] In another specific example, if there is a reference text whose highest similarity is less than a preset third similarity threshold, the reference text is determined to be a missing text; wherein the third similarity threshold is less than or equal to the first similarity threshold.
[0081] In actual implementation, the similarity between any initial text and any benchmark text can be calculated; for each benchmark text, the similarity between the benchmark text and each initial text is obtained, and the highest similarity is obtained.
[0082] The third similarity threshold may be equal to the first similarity threshold, for example, both are 0; the third similarity threshold may also be smaller than the first similarity threshold, for example, the first similarity threshold is 10% and the third similarity threshold is 0.
[0083] When there is a reference text whose highest similarity is less than the preset third similarity threshold, it can be understood that the similarity between the reference text and any initial text is very low. In this case, it can be determined that the reference text is a missing text.
[0084] In a specific example, a first amount of initial text is obtained; determining that the first amount is less than a second amount of reference text in the text description, and adding a first preset filler text to the initial text;
[0085] After the scene text recognition model completes parsing of the initial text in the text area, a first amount of the initial text is obtained; if the first amount is less than a second amount of the reference text in the text description, a first preset filler text is added to the initial text. Make the number of initial text and benchmark text match, then Wherein N is the second number. The first preset filler text may be a specified text, a specified symbol, etc.
[0086] Furthermore, if there is a reference text having the highest similarity with the first preset filling text, the reference text having the highest similarity with the first preset filling text is determined to be the missing text.
[0087] When the benchmark text has the highest similarity with the first preset filling text, it can be understood that the similarity between the benchmark text and all the initial texts is very low, the benchmark text is not reflected in the initial text, and the benchmark text is a missing text.
[0088] For example, if there is a reference text W i If the reference text has the highest similarity with the first preset filling text, then the reference text with the highest similarity with the first preset filling text is the missing text.
[0089] In a specific example, a first quantity of initial text is obtained; determining that the first quantity is greater than a second quantity of reference text in the text description, and adding a second preset filler text to the reference text;
[0090] After the scene text recognition model completes parsing of the initial text in the text area, a first amount of the initial text is obtained; if the first amount is greater than a second amount of the reference text in the text description, a second preset filler text p is added to the reference text. i , so that the number of initial text and benchmark text matches, then Wherein N is the first number. The second preset filler text may be a specified text, a specified symbol, etc.
[0091] Furthermore, if there is an initial text having the highest similarity with the second preset filling text, the initial text having the highest similarity with the second preset filling text is determined to be redundant text.
[0092] When the initial text has the highest similarity with the second preset filling text, it can be understood that the initial text has a low similarity with all the reference texts and the initial text is redundant. For example, if there is an initial text The initial text having the highest similarity with the second preset filling text is the redundant text.
[0093] In a specific example, a text distance between the initial text and a reference text in the text description is obtained; based on the text distance, the similarity between the initial text and the reference text is determined; wherein, the smaller the text distance, the higher the similarity.
[0094] The text distance can be achieved by the Levenshtein distance algorithm. Specifically, the one-to-one correspondence π between the initial text and the reference text in the text description is calculated by maximizing the matching Levenshtein distance algorithm, that is:
[0095]
[0096] Among them, w i ∈W, represents each benchmark text; Represents each initial text; the Lev() algorithm is used to calculate the similarity between the initial text and the benchmark text in the text description.
[0097] It can be understood that any pairing with a non-zero distance indicates that there is a typographical error in the initial image. The smaller the text distance, the higher the similarity. When the text distance is zero, it means that the initial text completely matches the reference text in the text description.
[0098] The above method can identify typesetting errors in the initial image and provide necessary information for subsequent updating of the area to be corrected.
[0099] In one embodiment, the following embodiments provide a specific implementation method for determining the reference text corresponding to the area to be corrected.
[0100] Specifically, it is determined that the area to be corrected contains erroneous text relative to the text description, a first text having the highest similarity to the erroneous text is obtained from the text description, and the first text is determined as a reference text corresponding to the area to be corrected.
[0101] The erroneous text corresponds to the first text. Therefore, the area to be corrected where the erroneous text is located needs to be corrected so that the erroneous text in the area to be corrected is corrected to the first text, ie, the reference text.
[0102] It is determined that the area to be corrected contains redundant text relative to the text description, and that the reference text corresponding to the area to be corrected is empty; wherein, for the area to be corrected with an empty reference text, no text exists in the updated area to be corrected.
[0103] The redundant text should not be displayed in the image originally. Therefore, the redundant text in the area to be corrected needs to be deleted so that the area to be corrected contains text. Therefore, the reference text corresponding to the area to be corrected is empty.
[0104] It is determined that the area to be corrected corresponds to missing text relative to the text description, and the missing text is determined as the reference text corresponding to the area to be corrected.
[0105] In the initial image, there is no text in the area to be corrected. In order to reflect the missing text, it is necessary to determine the missing text as the reference text of the area to be corrected so that the missing text is reflected in the image.
[0106] The above method determines the reference text corresponding to the area to be corrected according to the text error type of the area to be corrected, providing a basis for updating the area to be corrected.
[0107] The following embodiments provide specific implementation methods for updating the image content in the area to be corrected.
[0108] In one method, it is determined that the area to be corrected contains redundant text relative to the text description, and the baseline text corresponding to the area to be corrected is empty, and the image content of the area to be corrected is deleted; based on the image content in the adjacent area of the area to be corrected, updated content of the area to be corrected is generated, and the area to be corrected is filled with the updated content; wherein the updated content is continuous with the image content in the adjacent area.
[0109] First, it is determined that the area to be corrected contains redundant text relative to the text description, and the baseline text corresponding to the area to be corrected is empty, then the image content in the area to be corrected is deleted through the image restoration model; then the image restoration model can obtain the pixel information around the area to be corrected, and generate updated content of the area to be corrected based on the image content in the adjacent areas of the area to be corrected, and use the updated content to fill the area to be corrected. The updated content of the area to be corrected is continuous with the image content in the adjacent areas, thereby reconstructing the missing area of the image.
[0110] This image restoration model can be implemented by the LaMa model. In actual implementation, the mask range can be slightly expanded to ensure that the image content in the area to be corrected is covered and deleted.
[0111] The above method can effectively remove redundant text from the initial image, so as to facilitate the subsequent re-planning of the layout and reconstruction of the image content.
[0112] Furthermore, the first image obtained by filling the area to be corrected with the updated content is superimposed with the initial image after deleting the image content of the area to be corrected, to obtain a final image.
[0113] After the updated content is used to fill the area to be corrected and a first image is obtained, the first image can be superimposed with the initial image after deleting the image content of the area to be corrected to obtain a final image, thereby ensuring the fit between the updated content and the initial image after deleting the image content of the area to be corrected, and ensuring the natural integration of the updated content with the surrounding images.
[0114] In one approach, it is determined that the region to be corrected corresponds to missing text relative to the text description, and the missing text is generated in the region to be corrected.
[0115] When the area to be corrected corresponds to missing text in the text description, the image content in the area to be corrected is updated according to the missing text, and the missing text is generated in the image content of the area to be corrected.
[0116] In one approach, it is determined that the area to be corrected contains erroneous text relative to the text description, and based on a preset text editing model, the erroneous text in the area to be corrected is corrected.
[0117] When the area to be corrected contains incorrect text relative to the text description, a preset text editing model can be used to automatically correct the incorrect text in the area to be corrected, generating correct text for each character in the incorrect text that needs to be corrected. This text editing model can be implemented by the AnyText model.
[0118] Furthermore, the corrected text in the area to be corrected after correction processing is obtained, and it is determined that the corrected text is different from the baseline text. The steps of correcting the erroneous text in the area to be corrected based on the preset text editing model are continued until the corrected text is the same as the baseline text, or the number of correction processing reaches a preset threshold, such as 200 times.
[0119] After the correction process, the corrected text in the area to be corrected is obtained, and the corrected text is compared with the baseline text; if the corrected text is different from the baseline text, then continue to execute the aforementioned steps based on the preset text editing model, correct the erroneous text in the area to be corrected, and compare the corrected text with the baseline text until the corrected text is the same as the baseline text, or the preset maximum number of iterations is reached. For each iteration, the erroneous text in the area to be corrected is corrected once, and the maximum number of iterations is the aforementioned preset number threshold.
[0120] This embodiment is independent of the base model and, as a post-processing method, can be used in conjunction with any image generation model without any modification or fine-tuning of the base model, thereby maximizing the advantages of existing text-to-image generation models.
[0121] This embodiment improves the text rendering accuracy, can effectively identify and correct various text layout errors in the generated image, significantly improves the accuracy of text rendering, and improves the accuracy of text while ensuring image quality.
[0122] This embodiment improves the balance between image quality and accuracy, and is higher in terms of the balance between text accuracy and image quality than the baseline model that focuses on text generation.
[0123] This embodiment is fully automatic and easy to integrate. The steps of this embodiment are all fully automatic processes, do not require manual intervention, and are easy to integrate into existing processes, making them convenient for users to use.
[0124] This embodiment improves the robustness of the image generation model. Through post-processing, it effectively solves the inherent defects of the text-to-image generation model in text rendering, improves the robustness of the text-to-image generation model, and enables it to generate higher quality and more accurate images.
[0125] In one embodiment, if Figure 2 As shown, a schematic diagram of an initial image, an area to be corrected, and an updated image is provided, which is described in detail below:
[0126] exist Figure 2 In (a), the text description is first input into the preset image generation model, and the initial image is output; the text description is as follows:
[0127] 'If We Survive'Unabridged AudiobookDownload Narrated ByJeremy JohnsonBy Andrew 'Klavan';
[0128] Based on the benchmark texts 'If', 'We', 'Survive', and 'Klavan' in the text description, the area to be corrected containing the erroneous texts 'SURVIIVE' and 'KlAVVIN' is determined in the initial image; the image content in the area to be corrected is updated according to the benchmark texts, and the erroneous texts in the area to be corrected are corrected to 'Survive' and 'Klavan'.
[0129] exist Figure 2 In (b), the text description is first input into the preset image generation model, and the initial image is output; the text description is as follows:
[0130] 'Rescue Simulator 2014';
[0131] According to the reference texts 'Rescue', 'Simulator', and '2014' in the text description, in the initial image, a region to be corrected having missing text '2014' relative to the reference text is determined; and the missing text '2014' is generated in the region to be corrected according to the reference text.
[0132] exist Figure 2 In (c), the text description is first input into the preset image generation model, and the initial image is output; the text description is as follows:
[0133] 'Vintage Made in 1942Original'M nnerPremium T Shirt;
[0134] Based on the reference texts 'Vintage', 'Made', 'in', '1942' and 'Original' in the text description, a region to be corrected containing the erroneous text '942' is determined in the initial image. The image content in the region to be corrected is updated based on the reference texts, and the erroneous text in the region to be corrected is corrected to '1942'.
[0135] The image processing method provided in this embodiment is independent of a specific image generation model and can effectively identify and correct various text layout errors in images, significantly improving the accuracy of text rendering and enhancing text accuracy while maintaining image quality. This method can be automatically implemented and easily integrated, enhancing the robustness of the image generation model.
[0136] In one embodiment, if Figure 3 As shown in FIG, another schematic diagram of an initial image, a region to be corrected, and an updated image is provided, which is described in detail below:
[0137] First, the text description is input into the preset image generation model, and the initial image is output; the text description is as follows:
[0138] Create an advertisement for a hamburger with a pop background.Displaythe texts'10% OFF'and'HAMBURGER'with a tasty burger,fried patatoes,and aCoke.The text'Combo Meals'is required;
[0139] Among them, the benchmark text includes '10%' 'OFF' 'HAMBURGER' 'Combo' and 'Meals'.
[0140] exist Figure 3 In (a), the initial text includes '100%' 'OFF' 'HAMBURGER' 'AAA'.
[0141] exist Figure 3 In (b), the scene text detection model detects polygonal word regions in the initial image and obtains the detected initial text. The scene text recognition model parses the text strings in these polygonal word regions, then measures the similarity between the initial text and the benchmark text using the Levenshtein distance, and determines the area to be corrected in the initial image, as well as the error type.
[0142] exist Figure 3 In (c), the image restoration model deletes the redundant text 'AAA' in the area to be corrected, generates an updated image based on the image content adjacent to the area to be corrected, and fills the updated image into the area to be corrected; it also generates an overlay image without typesetting errors.
[0143] exist Figure 3 In (d), the visual language model replans the layout information of the missing text 'Combo' and 'Meals' and determines the area to be corrected in the initial image where the missing text is located.
[0144] exist Figure 3 In (e), the superimposed image is superimposed on the initial image after the text is deleted; the text editing model generates the missing texts 'Combo' and 'Meals' in the area to be corrected corresponding to the missing text based on the reference text.
[0145] The image processing method provided in this embodiment can improve the rendering accuracy of text content in an image and enhance the balance between image quality and text accuracy. The image processing method provided in this embodiment can be automatically implemented, is easy to integrate, and can be combined with any image generation model.
[0146] The embodiment of the present invention also provides a schematic diagram of an image processing device, such as Figure 4 As shown, the device includes:
[0147] A text description input module 41 is used to input the text description into a preset image generation model and output an initial image;
[0148] a determination module 42 configured to determine, based on the text description, a region to be corrected in the initial image and a reference text corresponding to the region to be corrected; wherein the image content in the region to be corrected does not match the reference text corresponding to the region to be corrected; and the text description includes at least one reference text;
[0149] The image content updating module 43 is configured to update the image content in the area to be corrected based on the reference text, so that the image content matches the reference text.
[0150] The above-mentioned image processing device inputs a text description into a preset image generation model and outputs an initial image; based on the text description, determines the area to be corrected in the initial image and the reference text corresponding to the area to be corrected; wherein the image content in the area to be corrected does not match the reference text corresponding to the area to be corrected; the text description includes at least one reference text; based on the reference text, updates the image content in the area to be corrected so that the image content matches the reference text.
[0151] In this method, after the model outputs the initial image, it determines the area to be corrected and the corresponding benchmark text from the initial image based on the text description, and then updates the image content in the area to be corrected according to the benchmark text, so that the updated image content matches the benchmark text, while ensuring the image quality and improving the rendering accuracy of the image content.
[0152] The above-mentioned determination module is further used to perform text detection on the initial image to obtain the detected initial text; and determine the area to be corrected in the initial image based on the initial text and the text description.
[0153] The above-mentioned determination module is further used to determine the target text that does not match the text description from the initial text, and determine the text area occupied by the target text in the initial image as the area to be corrected.
[0154] The above-mentioned target text includes: erroneous text relative to the text description; wherein the similarity between the erroneous text and part of the reference text in the text description is greater than a preset first similarity threshold and less than a preset second similarity threshold; the first similarity threshold is less than the second similarity threshold; or, the target text includes: redundant text relative to the text description; wherein the redundant text is text that does not exist in the text description.
[0155] The above-mentioned determination module is also used to determine the missing text from the text description based on the initial text; wherein the missing text is the text that does not exist in the initial text; based on the image content of the initial image, the area to be displayed of the missing text is determined from the initial image, and the area to be displayed is determined as the area to be corrected in the initial image.
[0156] The above-mentioned determination module is further used to obtain layout information of the initial text in the initial image; and determine the area to be displayed of the missing text from the initial image based on the background content and layout information of the initial image.
[0157] The above-mentioned device also includes a maximum similarity acquisition module, which is used to obtain the maximum similarity between the initial text and the reference text in the text description; if the maximum similarity is greater than a preset first similarity threshold and less than a preset second similarity threshold, the initial text is determined to be an incorrect text; wherein the first similarity threshold is less than the second similarity threshold; if the maximum similarity is greater than or equal to the preset second similarity threshold, the initial text is determined to be a correct text.
[0158] The above device also includes a redundant text determination module, which is used to determine that the initial text is redundant text if the highest similarity is less than a preset third similarity threshold; wherein the third similarity threshold is less than or equal to the first similarity threshold.
[0159] The above device also includes a missing text determination module, which is used to determine that the reference text is a missing text if there is a reference text with a highest similarity less than a preset third similarity threshold; wherein the third similarity threshold is less than or equal to the first similarity threshold.
[0160] The above-mentioned device also includes a first filling module, which is used to obtain a first quantity of initial text; determine that the first quantity is less than the second quantity of reference text in the text description, and add a first preset filling text to the initial text; the above-mentioned device also includes a second missing text determination module, which is used to determine that if there is a reference text with the highest similarity to the first preset filling text, the reference text with the highest similarity to the first preset filling text is the missing text.
[0161] The above-mentioned device also includes a second filling module, which is used to obtain a first quantity of initial text; determine that the first quantity is greater than the second quantity of reference text in the text description, and add a second preset filling text to the reference text; the above-mentioned device also includes a second redundant text determination module, which is used to determine that if there is an initial text with the highest similarity to the second preset filling text, the initial text with the highest similarity to the second preset filling text is redundant text.
[0162] The above device also includes a similarity determination module for obtaining a text distance between the initial text and a reference text in the text description; based on the text distance, determining the similarity between the initial text and the reference text; wherein, the smaller the text distance, the higher the similarity.
[0163] The determination module is further configured to determine whether the area to be corrected contains erroneous text relative to the text description, obtain a first text having the highest similarity to the erroneous text from the text description, and determine the first text as a reference text corresponding to the area to be corrected.
[0164] The above-mentioned determination module is also used to determine whether the area to be corrected contains redundant text relative to the text description, and determine that the reference text corresponding to the area to be corrected is empty; wherein, for the area to be corrected with an empty reference text, there is no text in the updated area to be corrected.
[0165] The above-mentioned determination module is further used to determine whether the area to be corrected corresponds to missing text relative to the text description, and determine the missing text as the reference text corresponding to the area to be corrected.
[0166] The above-mentioned image content update module is also used to determine that the area to be corrected contains redundant text relative to the text description, and the baseline text corresponding to the area to be corrected is empty, and delete the image content of the area to be corrected; based on the image content in the adjacent area of the area to be corrected, generate updated content of the area to be corrected, and use the updated content to fill the area to be corrected; wherein the updated content is continuous with the image content in the adjacent area.
[0167] The above-mentioned device also includes a superposition module for superimposing the first image obtained by filling the area to be corrected with the updated content with the initial image after deleting the image content of the area to be corrected, to obtain a final image.
[0168] The image content updating module is further configured to determine whether the area to be corrected corresponds to missing text relative to the text description, and to generate the missing text in the area to be corrected.
[0169] The image content updating module is further configured to determine whether the area to be corrected contains erroneous text relative to the text description, and to correct the erroneous text in the area to be corrected based on a preset text editing model.
[0170] The above-mentioned device also includes a corrected text acquisition module, which is used to obtain the corrected text in the area to be corrected after correction processing, determine that the corrected text is different from the baseline text, and continue to execute the steps of correcting the erroneous text in the area to be corrected based on the preset text editing model until the corrected text is the same as the baseline text, or the number of correction processing reaches a preset threshold.
[0171] This embodiment further provides an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-mentioned image processing method. The electronic device can be a server or a terminal device.
[0172] See also Figure 5 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores computer-executable instructions that can be executed by the processor 100. The processor 100 executes the computer-executable instructions to implement the above-mentioned image processing method.
[0173] Further, Figure 5 The electronic device shown further includes a bus 102 and a communication interface 103 , and the processor 100 , the communication interface 103 and the memory 101 are connected via the bus 102 .
[0174] The memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is achieved through at least one communication interface 103 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 102 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0175] The processor 100 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 100 or software instructions. The above processor 100 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as a random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or register. The storage medium is located in the memory 101. The processor 100 reads the information in the memory 101 and, in conjunction with its hardware, completes the steps of the method of the aforementioned embodiment.
[0176] The processor in the electronic device can implement the following operations in the image processing method by executing computer-executable instructions:
[0177] A text description is input into a preset image generation model to output an initial image; based on the text description, an area to be corrected in the initial image and a reference text corresponding to the area to be corrected are determined; wherein the image content in the area to be corrected does not match the reference text corresponding to the area to be corrected; the text description includes at least one reference text; based on the reference text, the image content in the area to be corrected is updated so that the image content matches the reference text.
[0178] Perform text detection on the initial image to obtain detected initial text; and determine the area to be corrected in the initial image based on the initial text and the text description.
[0179] The target text that does not match the text description is determined from the initial text, and the text area occupied by the target text in the initial image is determined as the area to be corrected.
[0180] The above-mentioned target text includes: erroneous text relative to the text description; wherein the similarity between the erroneous text and part of the reference text in the text description is greater than a preset first similarity threshold and less than a preset second similarity threshold; the first similarity threshold is less than the second similarity threshold; or, the target text includes: redundant text relative to the text description; wherein the redundant text is text that does not exist in the text description.
[0181] Based on the initial text, missing text is determined from the text description; wherein the missing text is text that does not exist in the initial text; based on the image content of the initial image, the area to be displayed of the missing text is determined from the initial image, and the area to be displayed is determined as the area to be corrected in the initial image.
[0182] Layout information of the initial text in the initial image is obtained; and based on the background content and layout information of the initial image, an area to be displayed of the missing text is determined from the initial image.
[0183] Obtain the highest similarity between the initial text and the reference text in the text description; if the highest similarity is greater than a preset first similarity threshold and less than a preset second similarity threshold, determine that the initial text is an incorrect text; wherein the first similarity threshold is less than the second similarity threshold; if the highest similarity is greater than or equal to the preset second similarity threshold, determine that the initial text is a correct text.
[0184] If the highest similarity is less than a preset third similarity threshold, the initial text is determined to be redundant text; wherein the third similarity threshold is less than or equal to the first similarity threshold.
[0185] If there is a reference text whose highest similarity is less than a preset third similarity threshold, the reference text is determined to be a missing text; wherein the third similarity threshold is less than or equal to the first similarity threshold.
[0186] Obtain a first quantity of initial text; determine that the first quantity is less than a second quantity of reference text in the text description, and add a first preset filler text to the initial text; if there is a reference text that has the highest similarity with the first preset filler text, determine that the reference text that has the highest similarity with the first preset filler text is the missing text.
[0187] Obtain a first quantity of initial text; determine that the first quantity is greater than a second quantity of reference text in the text description, and add a second preset filler text to the reference text; if there is an initial text with the highest similarity to the second preset filler text, determine that the initial text with the highest similarity to the second preset filler text is redundant text.
[0188] Obtain a text distance between the initial text and a reference text in the text description; and determine the similarity between the initial text and the reference text based on the text distance; wherein, the smaller the text distance, the higher the similarity.
[0189] It is determined that the area to be corrected contains erroneous text relative to the text description, a first text having the highest similarity to the erroneous text is obtained from the text description, and the first text is determined as a reference text corresponding to the area to be corrected.
[0190] It is determined that the area to be corrected contains redundant text relative to the text description, and that the reference text corresponding to the area to be corrected is empty; wherein, for the area to be corrected with an empty reference text, no text exists in the updated area to be corrected.
[0191] It is determined that the area to be corrected corresponds to missing text relative to the text description, and the missing text is determined as the reference text corresponding to the area to be corrected.
[0192] Determine that the area to be corrected contains redundant text relative to the text description, and that the baseline text corresponding to the area to be corrected is empty, and delete the image content of the area to be corrected; generate updated content of the area to be corrected based on the image content in the adjacent area of the area to be corrected, and fill the area to be corrected with the updated content; wherein the updated content is continuous with the image content in the adjacent area.
[0193] The first image obtained by filling the area to be corrected with the updated content is superimposed with the initial image after deleting the image content of the area to be corrected to obtain a final image.
[0194] It is determined that the area to be corrected corresponds to missing text relative to the text description, and the missing text is generated in the area to be corrected.
[0195] It is determined that the area to be corrected contains erroneous text relative to the text description, and based on a preset text editing model, the erroneous text in the area to be corrected is corrected.
[0196] Obtain the corrected text in the area to be corrected after correction processing, determine that the corrected text is different from the baseline text, and continue to execute the steps of correcting the erroneous text in the area to be corrected based on the preset text editing model until the corrected text is the same as the baseline text, or the number of correction processing reaches a preset threshold.
[0197] In this method, after the model outputs the initial image, it determines the area to be corrected and the corresponding benchmark text from the initial image based on the text description, and then updates the image content in the area to be corrected according to the benchmark text, so that the updated image content matches the benchmark text, while ensuring the image quality and improving the rendering accuracy of the image content.
[0198] This embodiment further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned image processing method.
[0199] The computer-executable instructions stored in the computer-readable storage medium can implement the following operations in the above-mentioned image processing method by executing the computer-executable instructions:
[0200] A text description is input into a preset image generation model to output an initial image; based on the text description, an area to be corrected in the initial image and a reference text corresponding to the area to be corrected are determined; wherein the image content in the area to be corrected does not match the reference text corresponding to the area to be corrected; the text description includes at least one reference text; based on the reference text, the image content in the area to be corrected is updated so that the image content matches the reference text.
[0201] Perform text detection on the initial image to obtain detected initial text; and determine the area to be corrected in the initial image based on the initial text and the text description.
[0202] The target text that does not match the text description is determined from the initial text, and the text area occupied by the target text in the initial image is determined as the area to be corrected.
[0203] The target text includes: erroneous text relative to the text description; wherein the similarity between the erroneous text and part of the reference text in the text description is greater than a preset first similarity threshold and less than a preset second similarity threshold; the first similarity threshold is less than the second similarity threshold; or, the target text includes: redundant text relative to the text description; wherein the redundant text is text that does not exist in the text description.
[0204] Based on the initial text, missing text is determined from the text description; wherein the missing text is text that does not exist in the initial text; based on the image content of the initial image, the area to be displayed of the missing text is determined from the initial image, and the area to be displayed is determined as the area to be corrected in the initial image.
[0205] Layout information of the initial text in the initial image is obtained; and based on the background content and layout information of the initial image, an area to be displayed of the missing text is determined from the initial image.
[0206] Obtain the highest similarity between the initial text and the reference text in the text description; if the highest similarity is greater than a preset first similarity threshold and less than a preset second similarity threshold, determine that the initial text is an incorrect text; wherein the first similarity threshold is less than the second similarity threshold; if the highest similarity is greater than or equal to the preset second similarity threshold, determine that the initial text is a correct text.
[0207] If the highest similarity is less than a preset third similarity threshold, the initial text is determined to be redundant text; wherein the third similarity threshold is less than or equal to the first similarity threshold.
[0208] If there is a reference text whose highest similarity is less than a preset third similarity threshold, the reference text is determined to be a missing text; wherein the third similarity threshold is less than or equal to the first similarity threshold.
[0209] Obtain a first quantity of initial text; determine that the first quantity is less than a second quantity of reference text in the text description, and add a first preset filler text to the initial text; if there is a reference text that has the highest similarity with the first preset filler text, determine that the reference text that has the highest similarity with the first preset filler text is the missing text.
[0210] Obtain a first quantity of initial text; determine that the first quantity is greater than a second quantity of reference text in the text description, and add a second preset filler text to the reference text; if there is an initial text with the highest similarity to the second preset filler text, determine that the initial text with the highest similarity to the second preset filler text is redundant text.
[0211] Obtain a text distance between the initial text and a reference text in the text description; and determine the similarity between the initial text and the reference text based on the text distance; wherein, the smaller the text distance, the higher the similarity.
[0212] It is determined that the area to be corrected contains erroneous text relative to the text description, a first text having the highest similarity to the erroneous text is obtained from the text description, and the first text is determined as a reference text corresponding to the area to be corrected.
[0213] It is determined that the area to be corrected contains redundant text relative to the text description, and that the reference text corresponding to the area to be corrected is empty; wherein, for the area to be corrected with an empty reference text, no text exists in the updated area to be corrected.
[0214] It is determined that the area to be corrected corresponds to missing text relative to the text description, and the missing text is determined as the reference text corresponding to the area to be corrected.
[0215] Determine that the area to be corrected contains redundant text relative to the text description, and that the baseline text corresponding to the area to be corrected is empty, and delete the image content of the area to be corrected; generate updated content of the area to be corrected based on the image content in the adjacent area of the area to be corrected, and fill the area to be corrected with the updated content; wherein the updated content is continuous with the image content in the adjacent area.
[0216] The first image obtained by filling the area to be corrected with the updated content is superimposed with the initial image after deleting the image content of the area to be corrected to obtain a final image.
[0217] It is determined that the area to be corrected corresponds to missing text relative to the text description, and the missing text is generated in the area to be corrected.
[0218] It is determined that the area to be corrected contains erroneous text relative to the text description, and based on a preset text editing model, the erroneous text in the area to be corrected is corrected.
[0219] Obtain the corrected text in the area to be corrected after correction processing, determine that the corrected text is different from the baseline text, and continue to execute the steps of correcting the erroneous text in the area to be corrected based on the preset text editing model until the corrected text is the same as the baseline text, or the number of correction processing reaches a preset threshold.
[0220] In this method, after the model outputs the initial image, it determines the area to be corrected and the corresponding benchmark text from the initial image based on the text description, and then updates the image content in the area to be corrected according to the benchmark text, so that the updated image content matches the benchmark text, while ensuring the image quality and improving the rendering accuracy of the image content.
[0221] The computer program products of the image processing methods, devices, and electronic devices provided in the embodiments of the present invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the previous method embodiments. For specific implementation, please refer to the method embodiments and will not be repeated here.
[0222] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0223] In addition, in the description of the embodiments of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0224] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0225] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0226] Finally, it should be noted that the above embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. An image processing method, characterized in that: The method comprises: Input the text description into the preset image generation model and output the initial image; Based on the text description, determining a region to be corrected in the initial image and a reference text corresponding to the region to be corrected; wherein the image content in the region to be corrected does not match the reference text corresponding to the region to be corrected; and the text description includes at least one reference text; Based on the reference text, the image content in the area to be corrected is updated so that the image content matches the reference text.
2. The method according to claim 1, characterized in that The step of determining the area to be corrected in the initial image based on the text description includes: Performing text detection on the initial image to obtain detected initial text; Based on the initial text and the text description, a region to be corrected in the initial image is determined.
3. The method according to claim 2, characterized in that The step of determining the area to be corrected in the initial image based on the initial text and the text description includes: Target text that does not match the text description is determined from the initial text, and a text area occupied by the target text in the initial image is determined as an area to be corrected.
4. The method according to claim 3, characterized in that The target text includes: an error text relative to the text description; wherein the similarity between the error text and a portion of the reference text in the text description is greater than a preset first similarity threshold and less than a preset second similarity threshold; and the first similarity threshold is less than the second similarity threshold; Alternatively, the target text includes: redundant text relative to the text description; wherein the redundant text is text that does not exist in the text description.
5. The method according to claim 2, characterized in that The step of determining the area to be corrected in the initial image based on the initial text and the text description includes: Determining missing text from the text description based on the initial text; wherein the missing text is text that does not exist in the initial text; Based on the image content of the initial image, a region to be displayed of the missing text is determined from the initial image, and the region to be displayed is determined as a region to be corrected in the initial image.
6. The method according to claim 5, characterized in that The step of determining the area to be displayed of the missing text from the initial image based on the image content of the initial image comprises: Obtaining layout information of the initial text in the initial image; Based on the background content of the initial image and the layout information, an area to be displayed of the missing text is determined from the initial image.
7. The method according to claim 2, characterized in that After performing text detection on the initial image to obtain the detected initial text, the method further includes: Obtaining the highest similarity between the initial text and the reference text in the text description; If the highest similarity is greater than a preset first similarity threshold and less than a preset second similarity threshold, determining that the initial text is an erroneous text; wherein the first similarity threshold is less than the second similarity threshold; If the highest similarity is greater than or equal to a preset second similarity threshold, the initial text is determined to be a correct text.
8. The method according to claim 7, characterized in that The method further comprises: If the highest similarity is less than a preset third similarity threshold, the initial text is determined to be redundant text; wherein the third similarity threshold is less than or equal to the first similarity threshold.
9. The method according to claim 7, characterized in that The method further comprises: If there is a reference text whose highest similarity is less than a preset third similarity threshold, the reference text is determined to be a missing text; wherein the third similarity threshold is less than or equal to the first similarity threshold.
10. The method according to claim 7, characterized in that Before the step of obtaining the highest similarity between the initial text and the reference text in the text description, the method further includes: Acquire a first quantity of the initial text; determine that the first quantity is less than a second quantity of the reference text in the text description, and add a first preset filler text to the initial text; After the step of obtaining the highest similarity between the initial text and the reference text in the text description, the method further includes: If there is a reference text having the highest similarity with the first preset filling text, the reference text having the highest similarity with the first preset filling text is determined to be the missing text.
11. The method according to claim 7, characterized in that Before the step of obtaining the highest similarity between the initial text and the reference text in the text description, the method further includes: Acquire a first quantity of the initial text; determine that the first quantity is greater than a second quantity of reference text in the text description, and add a second preset filler text to the reference text; After the step of obtaining the highest similarity between the initial text and the reference text in the text description, the method further includes: If there is the initial text having the highest similarity with the second preset filling text, the initial text having the highest similarity with the second preset filling text is determined to be redundant text.
12. The method according to claim 7, characterized in that Before the step of obtaining the highest similarity between the initial text and the reference text in the text description, the method further includes: Obtaining a text distance between the initial text and a reference text in the text description; Based on the text distance, the similarity between the initial text and the reference text is determined; wherein, the smaller the text distance, the higher the similarity.
13. The method according to claim 1, wherein The step of determining the reference text corresponding to the area to be corrected based on the text description includes: It is determined that the area to be corrected contains erroneous text relative to the text description, a first text having the highest similarity to the erroneous text is obtained from the text description, and the first text is determined as a reference text corresponding to the area to be corrected.
14. The method according to claim 1, wherein The step of determining the reference text corresponding to the area to be corrected based on the text description includes: Determine whether the area to be corrected contains redundant text relative to the text description, and determine whether the reference text corresponding to the area to be corrected is empty; wherein, for the area to be corrected where the reference text is empty, no text exists in the updated area to be corrected.
15. The method according to claim 1, wherein The step of determining the reference text corresponding to the area to be corrected based on the text description includes: It is determined that the area to be corrected corresponds to missing text relative to the text description, and the missing text is determined as a reference text corresponding to the area to be corrected.
16. The method according to claim 1, characterized in that The step of updating the image content in the area to be corrected based on the reference text comprises: Determining that the area to be corrected contains redundant text relative to the text description and that the reference text corresponding to the area to be corrected is empty, and deleting the image content of the area to be corrected; Based on the image content in the adjacent area of the area to be corrected, updated content of the area to be corrected is generated, and the area to be corrected is filled with the updated content; wherein the updated content is continuous with the image content in the adjacent area.
17. The method according to claim 16, characterized in that After the steps of generating updated content of the area to be corrected based on image content in an adjacent area of the area to be corrected, and filling the area to be corrected with the updated content, the method further comprises: The first image obtained by filling the area to be corrected with the updated content is superimposed on the initial image after deleting the image content of the area to be corrected to obtain a final image.
18. The method according to claim 1, wherein The step of updating the image content in the area to be corrected based on the reference text comprises: It is determined that the area to be corrected corresponds to missing text relative to the text description, and the missing text is generated in the area to be corrected.
19. The method according to claim 1, wherein The step of updating the image content in the area to be corrected based on the reference text comprises: It is determined that the area to be corrected contains erroneous text relative to the text description, and based on a preset text editing model, the erroneous text in the area to be corrected is corrected.
20. The method according to claim 19, characterized in that After the step of correcting the erroneous text in the area to be corrected, the method further includes: Obtain the corrected text in the area to be corrected after correction processing, determine that the corrected text is different from the baseline text, and continue to execute the steps of correcting the erroneous text in the area to be corrected based on the preset text editing model until the corrected text is the same as the baseline text, or the number of correction processing reaches a preset threshold number.
21. An image processing device, characterized in that: The device comprises: A text description input module is used to input the text description into a preset image generation model and output an initial image; a determination module configured to determine, based on the text description, a region to be corrected in the initial image and a reference text corresponding to the region to be corrected; wherein the image content in the region to be corrected does not match the reference text corresponding to the region to be corrected; and the text description includes at least one reference text; An image content updating module is used to update the image content in the area to be corrected based on the reference text so that the image content matches the reference text.
22. An electronic device, characterized in that: The image processing method comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the image processing method according to any one of claims 1 to 20.
23. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the image processing method according to any one of claims 1 to 20.