A method and related apparatus for generating images
By extracting features from template images and overlaying elements on the underlying images for verification, the problem of AI models struggling to replicate template images is solved, achieving efficient and accurate image and text generation that meets visual consistency and brand style requirements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CTRIP TRAVEL NETWORK TECH SHANGHAI0
- Filing Date
- 2026-03-27
- Publication Date
- 2026-06-23
AI Technical Summary
Existing AI-based image and text generation solutions struggle to replicate template images, failing to meet the requirements for visual consistency and brand style in image and text design.
By extracting the text style, layout, and background features of the template image, elements are overlaid on the original underlying image based on these features to generate initial text and images. The integrity of the underlying image is then verified to ensure that the generated text and images are consistent with the template image.
It achieves accurate replication of template images, and the generated graphics and text have complete background images, meeting the design requirements of graphic products and improving generation efficiency.
Smart Images

Figure CN122265462A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method and apparatus for generating images and text. Background Technology
[0002] In the digital age, posters, illustrations, and other graphic products have become important information delivery carriers for corporate promotion and event publicity. In marketing material creation and new media content publishing, it is often necessary to combine underlying images and text information to form graphic products that possess both visual appeal and information delivery functionality. Traditional graphic design methods rely on designers manually creating graphic products using image editing software, which is inefficient. To improve the efficiency of graphic generation, graphic generation solutions based on AI (Artificial Intelligence) models have emerged.
[0003] However, in order to maintain brand visual consistency and reduce trial-and-error costs, the designed graphic and text products often need to replicate the style of specific template images. The original intention of AI-based graphic and text generation solutions is to generate new content that matches user prompts, making them unsuitable for non-open-ended graphic and text generation tasks such as template replication.
[0004] Therefore, there is an urgent need for a graphic and text generation solution to address the problem of difficulty in replicating template images in graphic and text design tasks. Summary of the Invention
[0005] In view of the above problems, this application provides a method and related apparatus for generating text and images by accurately replicating template images. The specific solution is as follows:
[0006] The first aspect of this application provides a method for generating images and text, including:
[0007] Obtain the template image and the graphic and text materials to be generated, wherein the graphic and text materials include the original underlying image and the text content, wherein the text content includes the title text;
[0008] Extract template features from the template image, including: text style features, text layout features, and text background features;
[0009] Based on the template features and the text content, elements are overlaid on the original underlying image to generate an initial image and text; the overlaid elements include: text elements containing the title text;
[0010] Perform a base map integrity check on the initial image and text. If the base map integrity check passes, output the initial image and text as the target image and text.
[0011] A second aspect of this application provides a graphic generation apparatus, comprising:
[0012] The parameter input unit is used to obtain the template image and the graphic and text materials to be generated. The graphic and text materials include the original underlying image and the text content, and the text content includes the title text.
[0013] The feature extraction unit is used to extract template features from the template image, the template features including: text style features, text layout features and text background features;
[0014] The image and text generation unit is used to overlay elements on the original underlying image based on the template features and the text content to generate an initial image and text; the overlaid elements include: text elements containing the title text;
[0015] An integrity verification unit is used to perform a base map integrity verification on the initial image and text.
[0016] The image and text output unit is used to output the initial image and text as the target image and text if the base map integrity verification passes.
[0017] A third aspect of this application provides a graphic generation device, comprising at least one processor and a memory connected to the processor, wherein:
[0018] The memory is used to store computer programs;
[0019] The processor is used to execute the computer program so that the image and text generation device can implement the image and text generation method of the first aspect described above.
[0020] A fourth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to perform the graphic generation method of the first aspect described above.
[0021] By employing the above technical solution, this application overlays elements onto the original underlying image based on the template features of the template image to generate initial text and graphics. Since the template features include the text style, layout, and background features of the template image, the text elements in the initial text and graphics can replicate the template style. Furthermore, the generated initial text and graphics undergo background image integrity verification. If the verification passes, the initial text and graphics are output as the target text and graphics, ensuring that the final generated text and graphics possess background image integrity. Therefore, applying this application can achieve text and graphics generation tasks that meet the design requirements of text and graphics products, and can also achieve accurate replication of template images. Attached Figure Description
[0022] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0023] Figure 1 A schematic diagram of an implementation system architecture for the image and text generation method provided in this application embodiment;
[0024] Figure 2 A flowchart illustrating a text and image generation method provided in this application;
[0025] Figure 3 An example is provided illustrating the process of generating text and images;
[0026] Figure 4 This application provides a schematic diagram of the structure of a graphic generation device;
[0027] Figure 5 This is a schematic diagram of the structure of a graphic generation device provided in this application. Detailed Implementation
[0028] The embodiments of this application are described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. The terminology used in the implementation section of this application is only used to explain specific embodiments of this application and is not intended to limit this application.
[0029] This application provides a method for generating images and text, which can be applied to, for example... Figure 1 The system architecture shown may include a terminal 100 and a server 200. The server 200 may include one or more servers (…). Figure 1 (This example uses a server as an illustration).
[0030] Either terminal 100 or server 200 can be used independently to execute the image and text generation method provided in this application embodiment. Alternatively, terminal 100 and server 200 can also be used collaboratively to execute the image and text generation method provided in this application embodiment. Terminal 100 in this application embodiment can be a mobile phone, computer, etc., and this application embodiment does not impose any limitations on this.
[0031] This application provides a method for generating images and text, illustrated by applying the method to a computer device. Specifically, the computer device may be... Figure 1 The system consists of terminal 100 or a combination of terminal 100 and server 200. (Refer to...)Figure 2 The image and text generation method specifically includes the following steps:
[0032] Step S101: Obtain the template image and the graphic materials to be generated.
[0033] The template image is the image and text template to be replicated in this image and text generation task, which includes certain text styles and layout features. The image and text materials include the original underlying image and the text content. The original underlying image is a clean image without a text layer or watermark; in this task, it is used as the canvas background. The text content can include title text, which is the text element content added to the original underlying image. The aforementioned title text can include the main title; in some application scenarios, such as when the template image includes a subtitle, the title text can also include a subtitle.
[0034] In one possible implementation, the process of obtaining the title text may include:
[0035] Obtain the original text of the image and text to be generated, and extract the title text from the original text.
[0036] The original text can be an article containing text content, etc. The extracted title text can be the title in the original text. By identifying the structure of the original text, the title and body text are distinguished to determine the title text. Alternatively, the extracted title text can be the title text determined after content recognition analysis of the original text, such as calling a large model to instruct the large model to extract the title text of the original text.
[0037] For example, the graphic design to be generated can be a poster or other graphic design product.
[0038] Step S102: Extract template features from the template image.
[0039] The template features may include: text style features, text layout features, and text background features.
[0040] For example, text style features may include: text size, shadow effect, stroke parameters, etc.; text layout features may include the relative position of the main title within the entire canvas, etc. If the template image contains a subtitle, the text layout features may also include the relative position of the subtitle within the entire canvas. In another possible implementation, the subtitle layout features may represent the relative positional relationship between the main title and the subtitle, the spacing between the main title and the subtitle, etc.; text background features may include: features of the text background layer in the template image, such as background color, text box border features, etc. Text layout features and text background features can also be collectively referred to as text area features.
[0041] It should be noted that the extracted template features are features relative to a template image of a fixed size. For example, the size feature of text can be represented as: c-point font size on a template image with length a and width b. In one possible implementation, step S102 can be achieved by performing visual analysis on the template image, such as calling a large model to instruct it to extract template features from the template image.
[0042] Optionally, after extracting the template features, the text style features can be modified to ensure that the text style meets preset text constraints. These constraints limit font type, number of lines, alignment, etc. For example, the main title can be centered and have a maximum of two lines, while the subtitle can be displayed on a single line. By limiting the font type, the misuse of commercial fonts can be avoided, thus reducing the risk of infringement. By modifying the text style through these constraints, a graphic product adapted to the actual text can be generated while replicating the template image, thereby ensuring the quality of the generated graphic product.
[0043] Step S103: Based on the template features and the text content, overlay elements on the original underlying image to generate an initial image and text.
[0044] The superimposed elements include text elements containing the title text. By superimposing elements based on template features, the text elements in the generated initial image and text can be made visually consistent with the text elements in the template image, thus achieving the template replication task adapted to the original underlying image. In other words, when superimposing elements, the following objectives can be followed: if the original underlying image and the template image have the same size, then the text features in the generated initial image and text should be consistent with the text features of the template image; if the original underlying image and the template image have different sizes, then the text features in the scaled initial image and text should be consistent with the text features of the template image. The scaled initial image and text are obtained by scaling the generated initial image and text to match the size of the template image.
[0045] Step S104: Perform a base map integrity check on the initial image and text. If the base map integrity check passes, output the initial image and text as the target image and text.
[0046] Base image integrity verification refers to checking whether the base image area displayed in the initial graphic and text output matches the corresponding area in the original underlying image; in other words, it verifies whether the initial graphic and text output can guarantee the pixel-level integrity of the base image. By performing base image integrity verification before outputting the graphic and text output, the problem of original base image information loss caused by pixel-level modifications such as scaling, cropping, and color adjustment of the base image by the AI model can be solved, thus meeting the base image integrity requirements in graphic and text product design.
[0047] This embodiment overlays elements onto the original underlying image based on the template features of the template image to generate initial text and graphics. Since the template features include the text style, layout, and background features of the template image, the text elements in the initial text and graphics can replicate the template style. Furthermore, the generated initial text and graphics undergo a base image integrity check. If the check passes, the initial text and graphics are output as the target text and graphics, ensuring that the final generated text and graphics possess base image integrity. Therefore, this embodiment can achieve text and graphics generation tasks that meet the design requirements of text and graphics products, and can also achieve accurate replication of template images.
[0048] In one or more embodiments provided in this application, where the text content also includes address information, the superimposed element may further include: a positioning container.
[0049] It should be noted that posters and other graphic products typically include address information to indicate the specific location of the event venue; to improve information delivery efficiency, location icons corresponding to the address information can also be included in the graphic product. Location icons, also known as location markers, are visual address information, such as map images containing location indicator icons.
[0050] The location container described in this application refers to a combination of elements in a graphic product used to implement a location indication function. The location container may include: the address information and / or the location illustration corresponding to the address information. Where the template features also include location container features, the style of the superimposed location container matches the location container features. Where the template features do not include container features, the style of the superimposed location container may be a preset style.
[0051] Based on the above, graphic and textual products containing address indication information can be generated, ensuring the quality of graphic and textual generation.
[0052] In one or more embodiments provided in this application, the process of obtaining the text content may include:
[0053] Step S201: Obtain the original text of the image and text to be generated.
[0054] Step S202: Extract the title text and address information from the original document to generate the text content.
[0055] In one or more embodiments provided in this application, step S103, which involves overlaying elements onto the original underlying image based on the template features and the text content to generate an initial image and text, may include:
[0056] Invoke the large model to instruct it to generate initial images and text based on the image and text materials and the template features.
[0057] The prompts used can be used to: instruct the large model to overlay elements on the original underlying image so that the generated initial text and image have the integrity of the underlying image; instruct the large model to overlay elements according to the template features so that the text and text regions in the generated initial text and image are visually consistent with the text and text regions in the template image; and instruct the large model to prioritize ensuring that the overlaid elements are size-appropriate for the original underlying image and do not obscure the core visual content of the original underlying image when generating text and image.
[0058] By instructing the large model to overlay elements onto the base image, it's possible to ensure, to a certain extent, that the generated text and images match the original base image's dimensions. This means the width and height pixel values of the generated text and images are identical to the original base image. This avoids the large model scaling, cropping, stretching, or performing pixel-level modifications such as color correction or noise reduction on the base image during initial text and image generation. Furthermore, by instructing the large model to overlay elements according to template features, the characteristics of the template image can be replicated. Prioritize ensuring that the dimensions of the superimposed elements are adapted to the original underlying image; this can also be called the element adaptation rule. Based on this rule, if the template features are not suitable for the base image size, such as when the template ratio conflicts with the base image ratio, the template ratio can be abandoned. Prioritize ensuring that the superimposed elements do not obscure the core visual content of the original underlying image; this can also be called the occlusion avoidance rule. When possible, when superimposing elements, a certain safe distance can be maintained between the superimposed elements and the core visual content, such as 3% of the width of the original underlying image. Based on the aforementioned rules, the large model may not completely replicate the template when generating the initial text and images, in order to ensure that the superimposed elements are adapted to the original underlying image and do not obscure the core visual elements of the base image, thereby ensuring the quality of the generated text and images.
[0059] In one possible implementation, the prompt words used can also be used to instruct the large model to overlay elements in the following order: original underlying image, text background layer, main title, subtitle (if any), positioning container (if any), positioning icon (if any), and address information.
[0060] Compared to the approach of writing custom code to implement step S103, this embodiment reduces the difficulty of implementing the image and text generation task by inputting image and text elements and template features into the large artificial intelligence model and calling the large model to generate images and text; furthermore, the quality of image and text generation is guaranteed to a certain extent by using relevant prompts.
[0061] In one or more embodiments provided in this application, step S104, performing a base map integrity check on the initial image and text, may include:
[0062] Step S301: Compare whether the initial image and text are the same size as the original underlying image.
[0063] Step S302: Identify the overlay elements in the initial image and text, determine the area in the initial image and text other than the overlay elements as the target area, and compare whether the pixel values of the pixels in the target area are consistent with the corresponding pixels in the original underlying image.
[0064] The process of identifying overlapping elements in the initial image and text and defining the area outside the overlapping elements as the target area can be called interference layer stripping. Interference layers include text layers, positioning container layers, etc. The aforementioned pixel values specifically refer to RGB values.
[0065] Step S303: Compare whether the detailed features of the target region are consistent with those of the corresponding region in the original underlying image.
[0066] The detailed features may include at least one of the following: background texture direction, noise distribution, and color threshold range.
[0067] Step S304: Determine the base map integrity score based on each comparison result; if the base map integrity score is within the preset threshold, the base map integrity check is determined to be passed; otherwise, the base map integrity check is determined to be failed.
[0068] The base map integrity score can be represented as a value from 0 to 100. A base map integrity score greater than a preset threshold (e.g., 95) indicates that the base map integrity check has passed, while a score not greater than the preset threshold indicates that the base map integrity check has failed. Furthermore, the base map integrity check result can be represented by 0 or 1, where 0 indicates that the base map integrity check has failed, and 1 indicates that the base map integrity check has passed.
[0069] It should be noted that when the initial image and text are not the same size as the original underlying image, the base image integrity score can be determined to be 0, and the verification fails. When the initial image and text are the same size as the original underlying image, but the difference in pixel values or detail features exceeds the corresponding preset limit, the base image integrity score can be determined to be no more than the preset threshold, and the verification fails. When the initial image and text are the same size as the original underlying image, and the difference in pixel values or detail features is within the corresponding preset limit, the base image integrity score can be determined to be more than the preset threshold, and the verification passes. This application does not limit the specific method for determining the score.
[0070] By performing the above-mentioned base map integrity verification, we can ensure to a certain extent that the base map is not tampered with, providing a basis for outputting graphics and text that meet design requirements.
[0071] Optionally, the aforementioned base map integrity verification can be implemented by calling a model, which can be called an image comparison model. The output format of this model can be JSON format. For example, the output of this model can be represented as:
[0072] json
[0073] {
[0074] "result": <Base map integrity verification result, 1 / 0>,
[0075] "desc": "<Initial image description>",
[0076] "score": "<Base image completeness score, 0-100>"
[0077] }
[0078] In one or more embodiments provided in this application, the step of performing base map integrity verification on the initial image and text may further include:
[0079] Step S401: Perform image recognition on the initial image and text to generate a first image description, and perform image recognition on the original underlying image to generate a second image description.
[0080] Step S402: Determine the difference information between the first image description and the second image description and the text content.
[0081] Step S403: Determine whether the initial image and text contain unnecessary elements based on the difference information. If so, determine that the base map integrity check fails; otherwise, determine the base map integrity check result based on each comparison result.
[0082] In another possible implementation, image recognition can be performed on the initial image and text excluding the text layer to generate a third image description. Subsequently, the difference information between the third image description and the second image description can be determined to determine whether the initial image and text contain unnecessary elements.
[0083] By combining image descriptions with integrity checks, we can prevent the model from adding unnecessary elements or tampering with the base image, thus helping to ensure the quality of the final image and text generation.
[0084] In another possible implementation, the base map integrity score can be determined by combining the image description, and then the pass / fail status can be determined based on the base map integrity score.
[0085] Optionally, when calling the large model to generate the initial text and images, the output of the large model may include, in addition to the initial text and images, an image description of the initial text and images, such as an image description excluding the text layer. For example, the output format of the large model can be JSON format, and its specific structure can be represented as follows:
[0086] json
[0087] {
[0088] "img": "<Initial image and text base64 encoding>",
[0089] "imgDesc": "<Image description of the initial image and text>"
[0090] }
[0091] In one or more embodiments provided in this application, the method may further include:
[0092] If the base image integrity check fails, return to the step of extracting template features from the template image and execute sequentially until the generated initial image and text pass the base image integrity check or the preset maximum number of retries is reached.
[0093] For example, the maximum number of retries can be 3. Optionally, if the preset maximum number of retries is reached, a prompt message indicating that the image and text generation failed can also be output. By automatically retrying, the success rate of image and text generation is improved.
[0094] The graphic and text generation solution provided in this application can achieve high-quality graphic and text generation tasks by accurately replicating template images without manual intervention. Compared with the manual graphic and text product design method, the graphic and text generation efficiency of this application is higher and can be widely applied to scenarios such as marketing material generation, new media content production, and advertising graphic and text generation, especially in scenarios with high requirements for the integrity of the base image and the replicability of the style.
[0095] The following section illustrates the image and text generation scheme provided in this application with examples. See also... Figure 3 The specific process of generating images and text includes:
[0096] The first step is to obtain the template image and the text / image materials to be generated. Specifically, this may include obtaining the template image used to provide the style template. The template image includes the main title "Summer Resort Hotel", the subtitle "Encountering the Slow Time by the Sea", and a location container carrying the address information "Address: xx City xxxxxxxxxx". Next, obtain the original underlying image and original text of the text / image materials to be generated. Extract the title text "Water Town Night Charm: Embark on a Romantic Encounter with Lights and Stars" from the original text, and use it as the text content of the text / image materials to be generated. This completes the acquisition of the text / image materials for the text / image materials to be generated.
[0097] The second step is to extract template features. Template features can specifically include text style features, text layout features, text background features, and positioning container features. For example, text style features can include: font / size, shadow color / blur radius, stroke color / size, etc.; text layout features can include: the distance between the main title and the top of the canvas, the spacing between the main and subtitles, etc.; text background features can include: border style / color, etc.; positioning container features can include: container position, border style / color, and text style features of address information, etc. Furthermore, after identifying the features, the size information in the identified features can be converted into relative size information relative to the template image size to obtain the final template features.
[0098] The third step is to generate the initial text and image. Specifically, this may include: determining the main title and subtitle elements that need to be superimposed on the original underlying image according to the aforementioned template features, and superimposing the elements on the original underlying image in the following order: main title background (if any), main title, subtitle background (if any), subtitle, and positioning container (if any), following the principle of occlusion avoidance.
[0099] Step 4: Perform base map integrity verification. This may include: stripping interfering elements from the generated initial image and text, removing the main title background (if any), main title, subtitle background (if any), subtitle, and positioning container (if any), resulting in an image containing only the target area, called the initial base map; then comparing the dimensions of the initial image and text with the original base image, comparing the pixel values at the same coordinates in the initial base map and the original base image, and comparing the background texture, noise distribution, and color range of the initial base map and the original base image to determine the base map integrity score based on the comparison results, and outputting the initial image and text that pass the integrity verification in base64 encoded form, completing this image and text generation task.
[0100] The image and text generation apparatus provided in the embodiments of this application is described below. The image and text generation apparatus described below can be referred to in correspondence with the image and text generation method described above.
[0101] See Figure 4 , Figure 4 This is a schematic diagram of the structure of a graphic generation device disclosed in an embodiment of this application.
[0102] like Figure 4 As shown, the device may include:
[0103] The parameter input unit 11 is used to obtain the template image and the graphic and text materials to be generated. The graphic and text materials include the original underlying image and the text content, and the text content includes the title text.
[0104] Feature extraction unit 12 is used to extract template features from the template image, the template features including: text style features, text layout features and text background features;
[0105] The image and text generation unit 13 is used to overlay elements on the original underlying image according to the template features and the text content to generate an initial image and text; the overlaid elements include: text elements containing the title text;
[0106] Integrity verification unit 14 is used to perform base map integrity verification on the initial image and text;
[0107] The image and text output unit 15 is used to output the initial image and text as the target image and text if the base map integrity verification passes.
[0108] In one or more embodiments provided in this application, when the text content further includes address information, the superimposed element further includes: a positioning container, the positioning container including: the address information and / or the positioning illustration corresponding to the address information; wherein, when the template feature further includes a positioning container feature, the style of the superimposed positioning container matches the positioning container feature.
[0109] In one or more embodiments provided in this application, the process of parameter input unit 11 acquiring text content includes:
[0110] Obtain the original text of the image and text to be generated;
[0111] The title text and address information are extracted from the original text to generate the text content.
[0112] In one or more embodiments provided in this application, the process by which the image and text generation unit 13 overlays elements on the original underlying image to generate initial images and text based on the template features and the text content may include:
[0113] The large model is invoked to instruct it to generate initial text and images based on the text and image materials and the template features. The prompts used are: to instruct the large model to overlay elements on the original underlying image so that the generated initial text and images have the integrity of the underlying image; to instruct the large model to overlay elements according to the template features so that the text and text regions in the generated initial text and images are visually consistent with the text and text regions in the template image; and to instruct the large model to prioritize ensuring that the overlaid elements are size-appropriate for the original underlying image and do not obscure the core visual content of the original underlying image when generating text and images.
[0114] In one or more embodiments provided in this application, the process of the integrity verification unit 14 performing a base map integrity verification on the initial image and text may include:
[0115] Compare whether the dimensions of the initial image and text are consistent with those of the original underlying image;
[0116] Identify the overlay elements in the initial image and text, determine the area in the initial image and text other than the overlay elements as the target area, and compare whether the pixel values of the pixels in the target area are consistent with the corresponding pixels in the original underlying image.
[0117] Compare the target region with the corresponding region in the original underlying image to see if their detailed features are consistent; the detailed features include at least one of the following: background texture direction, noise distribution, and color threshold range;
[0118] The base map integrity score is determined based on the comparison results; if the base map integrity score is within a preset threshold, the base map integrity check is determined to be passed; otherwise, the base map integrity check is determined to be failed.
[0119] In one or more embodiments provided in this application, the process of the integrity verification unit 14 performing base map integrity verification on the initial image and text may further include:
[0120] The initial image and text are subjected to image recognition to generate a first image description;
[0121] Perform image recognition on the original underlying image to generate a second image description;
[0122] Determine the difference information between the first image description and the second image description and the text content;
[0123] Based on the difference information, determine whether the initial image and text contain unnecessary elements. If so, determine that the base map integrity check fails; otherwise, determine the base map integrity check result based on each comparison result.
[0124] In one or more embodiments provided in this application, the device may further include: a retry control unit, used to retry image and text generation if the base map integrity verification fails;
[0125] Image and text generation retries may include returning to the step of extracting template features from the template image and executing them sequentially. When the preset maximum number of retries is reached, the retry control unit stops retrying image and text generation.
[0126] Each unit in the image and text generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These units can be embedded in the processor of a computer device in hardware form or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each unit.
[0127] This application also provides a graphic generation device in its embodiments. (See reference) Figure 5 The diagram illustrates a structure suitable for implementing the image and text generation device in the embodiments of this application. The image and text generation device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, tablets, etc. Figure 5 The illustrated image and text generation device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0128] like Figure 5 As shown, the image and text generation device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 2 or a program loaded from a storage device 8 into a random access memory (RAM) 3, to implement the image and text generation method of the foregoing embodiments of this application. When the image and text generation device is powered on, the RAM 3 also stores various programs and data required for the operation of the image and text generation device. The processing unit 1, ROM 2, and RAM 3 are interconnected via a bus 4. An input / output (I / O) interface 5 is also connected to the bus 4.
[0129] Typically, the following devices can be connected to I / O interface 5: input devices 6 including, for example, touchscreens, touchpads, keyboards, mice, etc.; output devices 7 including, for example, liquid crystal displays (LCDs); storage devices 8 including, for example, memory cards, hard drives, etc.; and communication devices 9. Communication device 9 allows the graphic generation device to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 A graphic generation apparatus with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0130] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the graphic generation methods provided in this application.
[0131] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the image and text generation methods provided in this application.
[0132] Finally, it should be noted that in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0133] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0134] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating images and text, characterized in that, include: Obtain the template image and the graphic and text materials to be generated, wherein the graphic and text materials include the original underlying image and the text content, wherein the text content includes the title text; Extract template features from the template image, including: text style features, text layout features, and text background features; Based on the template features and the text content, elements are overlaid on the original underlying image to generate an initial image and text; the overlaid elements include: text elements containing the title text; Perform a base map integrity check on the initial image and text. If the base map integrity check passes, output the initial image and text as the target image and text.
2. The image and text generation method according to claim 1, characterized in that, If the text content also includes address information, the superimposed element also includes a positioning container, which includes the address information and / or the positioning icon corresponding to the address information; wherein, if the template feature also includes a positioning container feature, the style of the superimposed positioning container matches the positioning container feature.
3. The image and text generation method according to claim 2, characterized in that, The process of obtaining the text content includes: Obtain the original text of the image and text to be generated; The title text and address information are extracted from the original text to generate the text content.
4. The image and text generation method according to any one of claims 1-3, characterized in that, Based on the template features and the text content, elements are overlaid on the original underlying image to generate an initial image and text, including: The large model is invoked to instruct it to generate initial text and images based on the text and image materials and the template features. The prompts used are: to instruct the large model to overlay elements on the original underlying image so that the generated initial text and images have the integrity of the underlying image; to instruct the large model to overlay elements according to the template features so that the text and text regions in the generated initial text and images are visually consistent with the text and text regions in the template image; and to instruct the large model to prioritize ensuring that the overlaid elements are size-appropriate for the original underlying image and do not obscure the core visual content of the original underlying image when generating text and images.
5. The image and text generation method according to any one of claims 1-3, characterized in that, Perform a base map integrity check on the initial image and text, including: Compare whether the dimensions of the initial image and text are consistent with those of the original underlying image; Identify the overlay elements in the initial image and text, determine the area in the initial image and text other than the overlay elements as the target area, and compare whether the pixel values of the pixels in the target area are consistent with the corresponding pixels in the original underlying image. Compare the target region with the corresponding region in the original underlying image to see if their detailed features are consistent; the detailed features include at least one of the following: background texture direction, noise distribution, and color threshold range; The base map integrity score is determined based on the comparison results; if the base map integrity score is within a preset threshold, the base map integrity check is determined to be passed; otherwise, the base map integrity check is determined to be failed.
6. The image and text generation method according to claim 5, characterized in that, The process of verifying the integrity of the initial image and text also includes: The initial image and text are subjected to image recognition to generate a first image description; Perform image recognition on the original underlying image to generate a second image description; Determine the difference information between the first image description and the second image description and the text content; Based on the difference information, determine whether the initial image and text contain unnecessary elements. If so, determine that the base map integrity check fails; otherwise, determine the base map integrity check result based on each comparison result.
7. The image and text generation method according to any one of claims 1-3, characterized in that, The method also includes: If the base image integrity check fails, return to the step of extracting template features from the template image and execute sequentially until the generated initial image and text pass the base image integrity check or the preset maximum number of retries is reached.
8. A graphic generation device, characterized in that, include: The parameter input unit is used to obtain the template image and the graphic and text materials to be generated. The graphic and text materials include the original underlying image and the text content, and the text content includes the title text. The feature extraction unit is used to extract template features from the template image, the template features including: text style features, text layout features and text background features; The image and text generation unit is used to overlay elements on the original underlying image based on the template features and the text content to generate an initial image and text; the overlaid elements include: text elements containing the title text; An integrity verification unit is used to perform a base map integrity verification on the initial image and text. The image and text output unit is used to output the initial image and text as the target image and text if the base map integrity verification passes.
9. A graphic generation device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the image and text generation device can implement the image and text generation method as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the image and text generation method as described in any one of claims 1 to 7.