Image-text generation method, image-text generation device, electronic equipment and storage medium

Through the combination of copywriting generation model, content integration model and image generation model, the problem of the inability to generate multiple copywriting and images in the existing technology that the theme is closely matched and the style is consistent, and efficient and excellent quality graphic content generation is achieved, improving the user experience.

CN120147445APending Publication Date: 2025-06-13MICRO DREAM TECHTRONIC NETWORK TECH CHINACO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510131638.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art cannot generate multiple copywriting that is closely aligned with a specific topic and consistent in style, and cannot generate multiple images with consistent creative style and unified themes for each copywriting, resulting in the generated content being relatively monotonous and insufficient accuracy.

Method used

Through the combination of copywriting generation model, content integration model and image generation model, multiple copywriting matching with a specific graphic and text theme are generated, and images with consistent styles are generated based on the key visual information of each copywriting and preset image style prompt words.

Benefits of technology

It realizes the generation of multiple copy and image series that are closely matched with a specific topic, ensuring the style consistency of the graphic and text content, improving the efficiency and quality of content generation, and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147445A_ABST
    Figure CN120147445A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image-text generation method, an image-text generation device, electronic equipment and a storage medium, and the method comprises the steps: generating a plurality of target copywriting matched with an image-text theme through a copywriting generation model according to the predetermined image-text theme and copywriting prompt information customized based on the image-text theme; analyzing each target copywriting generated by the copywriting generation model through a content integration model, and extracting key visual information corresponding to each target copywriting; respectively generating at least one target image matched with each target copywriting through an image generation model according to the key visual information corresponding to each target copywriting and a preset image style prompt word; the image style cue word represents a unified target style type of the generated target image; and determining each target copywriting and the at least one target image matched with each target copywriting as each target image-text generation result corresponding to the image-text theme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular, to a method and device for generating text and images, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of artificial intelligence and deep learning technologies, the intelligent content generation (AIGC) technology has received significant attention and progress, and the AIGC technology has opened up new possibilities for content creation. For example, social media platforms can use the AIGC technology to automatically generate content that attracts users and reduce the dependence of users on professional content creators. In the prior art, corresponding text or images can be generated using the AIGC technology according to the key information provided by users. However, the existing AIGC technology can only generate single text or image content that is close in meaning to a specific keyword among the key information provided by users, and cannot generate multiple copies that are closely matched to a specific theme and have a consistent style, nor can it generate multiple images with a consistent creative style and a unified theme for each copy, resulting in relatively monotonous and inaccurate generated content, lacking sufficient creativity and attractiveness, and thus greatly weakening the user experience. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide a method and device for generating text and images, an electronic device, and a storage medium, so as to solve the problem in the prior art that multiple copies that are closely matched to a specific theme and have a consistent style cannot be generated, and multiple images with a consistent creative style and a unified theme cannot be generated for each copy.

[0004] To solve the above technical problems, the embodiments of the present application are implemented as follows: In a first aspect, the embodiments of the present application provide a method for generating text and images, including: Generating multiple target copies that match the text and image theme according to a pre-determined text and image theme and copy prompt information customized based on the text and image theme through a copy generation model; Analyzing each target copy generated by the copy generation model through a content integration model, and extracting key visual information corresponding to each target copy; Generating at least one target image that matches each target copy through an image generation model according to the key visual information corresponding to each target copy and a preset image style prompt word; the image style prompt word represents the target style type that the generated target images uniformly have; Determining each target copy and at least one target image that matches each target copy as each target text and image generation result corresponding to the text and image theme.

[0005] In a second aspect, an embodiment of the present application provides a graphic and text generation device, including: A copywriting generation module, configured to generate a plurality of target copywritings matching the graphic and text theme through a copywriting generation model according to a pre-determined graphic and text theme and copywriting prompt information customized based on the graphic and text theme; A content integration module, configured to analyze each target copywriting generated by the copywriting generation model through a content integration model, and extract key visual information corresponding to each target copywriting; An image generation module, configured to generate at least one target image matching each target copywriting respectively through an image generation model according to the key visual information corresponding to each target copywriting and a preset image style prompt word; the image style prompt word represents a target style type uniformly possessed by the generated target images; A graphic and text generation module, configured to determine each target copywriting and at least one target image matching each target copywriting as each target graphic and text generation result corresponding to the graphic and text theme.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the graphic and text generation method described in the first aspect above are implemented.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the graphic and text generation method described in the first aspect above are implemented.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the graphic and text generation method described in the first aspect above are implemented.

[0009] In the embodiments of the present application, first, through a copywriting generation model, multiple target copywritings matching the graphic and text theme are generated according to a pre-determined graphic and text theme and copywriting prompt information customized based on the graphic and text theme; then, through a content integration model, each target copywriting generated by the copywriting generation model is analyzed to extract the key visual information corresponding to each target copywriting; then, through an image generation model, at least one target image matching each target copywriting is generated respectively according to the key visual information corresponding to each target copywriting and a preset image style prompt word; the image style prompt word represents the target style type that the generated target images uniformly have. Finally, each target copywriting and at least one target image matching each target copywriting are determined as each target graphic generation result corresponding to the graphic and text theme. It can be seen that through the embodiments of the present application, the generated copywriting and images are no longer monotonous, but a graphic series closely matching a certain specific graphic and text theme, and the style of each generated target copywriting and the target image corresponding to each target copywriting is consistent and the theme is unified. In this way, not only the style consistency of the graphic and text content is ensured, but also the efficiency and quality of the graphic and text content generation are greatly improved, thereby enhancing the user experience to a great extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following described drawings are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a schematic flowchart of a graphic generation method provided by an embodiment of the present application; Figure 2a It is a schematic diagram of a target image provided by an embodiment of the present application; Figure 2b It is a schematic diagram of another target image provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of the target graphic generation process provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of a graphic generation device provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The embodiments of the present application provide a graphic generation method, a graphic generation device, an electronic device, and a storage medium.

[0013] To enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0014] The application scenarios of the graphic and text generation method provided in the embodiments of the present application include, but are not limited to, the scenarios of creative posting activities in Internet platforms or application programs. Among them, the Internet platform can be an e-commerce platform, a streaming service platform, a social platform, an advertising platform, etc., and the application program can be a social software or a game software, etc. In an exemplary application scenario, a user can enter the theme page of "If Returning to Ancient Times" through a specific link in the Internet platform and obtain their own ancient identity through an interactive method similar to card drawing. Each card-drawing result includes a detailed personal background introduction, which covers occupation, family, partner, achievements, and corresponding poems, as well as a classical portrait that matches it. Through the graphic and text generation method provided in the embodiments of the present application, multiple different copywriters that match the theme of "If Returning to Ancient Times" and multiple images that match each copywriter can be generated, greatly enriching the interaction of users on the Internet platform.

[0015] The following takes the above exemplary application scenario as an example to detail the graphic and text generation method provided in the embodiments of the present application.

[0016] Figure 1 It is a schematic flowchart of a graphic and text generation method provided in the embodiments of the present application. As Figure 1 shown, the embodiments of the present application provide a graphic and text generation method, which can be executed by a server. The server can be an independent server, a server cluster composed of multiple servers, or a cloud server performing cloud computing processing. The graphic and text generation method can specifically include the following steps: S102, through a copywriting generation model, generate multiple target copywriters that match the graphic and text theme according to the pre-determined graphic and text theme and the copywriting prompt information customized based on the graphic and text theme.

[0017] In the embodiments of the present application, the pre-determined graphic and text theme and the copywriting prompt information customized according to the graphic and text theme can be input into the copywriting generation model, and through the copywriting generation model, multiple target copywriters that match the graphic and text theme are generated. Among them, the style types of the generated multiple target copywriters are unified.

[0018] In one embodiment, the copywriting generation model can be a large language model. The large language model is, for example, the ChatGPT (Chat Generative Pre-trained Transformer) model.

[0019] In one embodiment, the copywriting prompt information customized based on the graphic and text theme includes copywriting generation prompt information and copywriting extension prompt information; Through the copywriting generation model, according to the pre-determined graphic and text theme and the copywriting prompt information customized based on the graphic and text theme, multiple target copywritings matching the graphic and text theme are generated, including: Through the copywriting generation model, according to the pre-determined graphic and text theme and the copywriting generation prompt information, multiple initial copywritings matching the graphic and text theme are generated; Through the copywriting generation model, according to each initial copywriting and the copywriting extension prompt information, at least one target copywriting corresponding to each initial copywriting is generated respectively.

[0020] In an exemplary implementation, assume that the above copywriting generation prompt information is: "Suppose you have traveled back to ancient China. Please generate a life story of a character based on the following information: Gender: [gender], Occupation: [occupation], Family background: [family background], Partner situation: [partner situation], Main achievements or special experiences: [main achievements or special experiences]. Please include the details of this information in the generated story and ensure that the story has cultural depth and creativity." Input the above copywriting generation prompt information into the copywriting generation model. Through the copywriting generation model, multiple initial copywritings matching the graphic and text theme and having the same style type can be obtained. For example, Initial copywriting 1 is: Gender: male, Family origin: Escort family, Occupation: Escort head, Partner: The empire's chief songstress, Achievement: Becoming a legend of border escorts with superb martial arts skills. Initial copywriting 2 is: Gender: male, Family origin: Fishing village, Occupation: Fisherman, Partner: A kind-hearted fisherman's daughter, Achievement: Catching a mysterious big fish and enriching the entire village. Initial copywriting 3 is: Gender: female, Family origin: The richest family in the south of the Yangtze River, Occupation: Cloth store manager, Partner: A handsome young man, Achievement: Helping the family become the richest in the south of the Yangtze River. Initial copywriting 4 is: Gender: female, Family origin: A scholarly family, Occupation: Painter, Partner: General, Achievement: One of the four talented women.

[0021] After obtaining the above multiple initial copywritings, through the copywriting generation model, according to each initial copywriting and the copywriting extension prompt information, at least one target copywriting corresponding to each initial copywriting can be generated respectively, that is, on the basis of the already generated initial copywritings, one or more target copywritings corresponding to each initial copywriting are generated.

[0022] In an exemplary implementation, it can be based on each of the above-generated initial copywriting texts, and through a copywriting generation model, generate corresponding poems, articles, etc. for each initial copywriting text. For example, the copywriting expansion prompt information can be: "You are an ancient poet, good at writing seven-character quatrains. A seven-character quatrain means there are a total of four lines of poetry, with seven characters in each line, and it should rhyme. Please refer to the following background and write a seven-character quatrain. The format of the background is: Gender: {}, Family: {}, Occupation: {}, Partner: {}, Achievements: {}." Taking the above initial copywriting text 2: "Gender: male, Family origin: Fishing Village, Occupation: Fisherman, Partner: Kind-hearted fisherwoman, Achievements: Caught a mysterious big fish, enriching the entire village." as an example, input the above copywriting expansion information and initial copywriting text 2 into the copywriting generation model. Through the copywriting generation model, the poem generated corresponding to this initial copywriting text 2 is: "The wave-breaking fisherman catches the divine fish, and the kind-hearted fisherwoman warms the boat. Together, they make their family name known far and wide, and the village is rich and the people are at peace, with laughter filling the corridors." The poem generated for the initial copywriting text 2 above is the target copywriting corresponding to the initial copywriting text 2.

[0023] In one implementation, the initial copywriting text can also be determined as the target copywriting.

[0024] It can be understood that the styles of the above initial copywriting text 1, initial copywriting text 2, initial copywriting text 3, and initial copywriting text 4 are consistent. They all adopt a concise and clear narrative method to introduce the gender, family origin, occupation, partner, and achievements of the characters, etc. Each copywriting text highlights a certain characteristic or achievement of the character and has a certain story. Based on this, the styles of each target copywriting corresponding to the initial copywriting text 1, initial copywriting text 2, initial copywriting text 3, and initial copywriting text 4 are also consistent.

[0025] S104: Through the content integration model, analyze each target copywriting generated by the copywriting generation model, and extract the key visual information corresponding to each target copywriting.

[0026] In the embodiments of the present application, after generating multiple target copywriting texts, through the content integration model, each target copywriting generated by the copywriting generation model can be analyzed and processed to extract the key visual information corresponding to each target copywriting.

[0027] In the embodiments of the present application, the content integration model can be a large language model.

[0028] In one embodiment, through the content integration model, analyze each target copywriting generated by the copywriting generation model, and extract the key visual information corresponding to each target copywriting, including: Through the content integration model, respectively extract the object keywords and theme keywords included in each target copywriting; the object keywords are used to describe the target objects included in the target copywriting, and the theme keywords are used to describe the graphic and text theme corresponding to the target copywriting. Through the content integration model, according to the object keywords and theme keywords included in each target copywriting, the key visual information corresponding to each target copywriting is generated respectively.

[0029] In the embodiments of the present application, the object keywords used to describe the target objects included in each target copywriting and the theme keywords used to describe the graphic theme corresponding to each target copywriting can be extracted respectively through the content integration model. After the object keywords and theme keywords included in each target copywriting are obtained, through the content integration model, according to the object keywords and theme keywords included in each target copywriting, the key visual information corresponding to each target copywriting is generated respectively.

[0030] In an exemplary implementation manner, when extracting the object keywords used to describe the target objects included in each target copywriting and the theme keywords used to describe the graphic theme corresponding to each target copywriting through the content integration model, the extraction can be performed according to the pre-customized key visual information extraction prompt information. The key visual information extraction prompt information is, for example: "Suppose you are an artist who is good at creating pictures according to text descriptions. Please describe the scene of a picture according to the following information: Gender: {}, Occupation: {}, Era: Ancient China. Please transform this information into specific visual elements, describe the appearance, posture and background environment of this occupation, so as to create a vivid picture."

[0031] In one embodiment, through the content integration model, according to the object keywords and theme keywords included in each target copywriting, the key visual information corresponding to each target copywriting is generated respectively, including: Through the content integration model, according to the object keywords included in each target copywriting, the object key visual information matching the target objects included in each target copywriting is generated respectively; Through the content integration model, according to the theme keywords included in each target copywriting, the scene key visual information corresponding to the object key visual information is generated respectively; Through the content integration model, for each target copywriting, the object key visual information and the scene key visual information corresponding to the target copywriting are determined as the key visual information corresponding to the target copywriting.

[0032] In an exemplary implementation, it is assumed that the target copy 1 is: gender: male, family background: escort family, occupation: escort head, partner: empire's chief singer, achievement: relying on superb martial arts, becoming a frontier bodyguard legend. Target copy 2 is: gender: male, family background: fishing village, occupation: fisherman, partner: a kind-hearted fisherman's daughter, achievement: caught a mysterious big fish and enriched the entire village. First, taking target copy 1 as an example, through the content integration model, according to the above key visual information, the prompt information is extracted, and the object keywords contained in target copy 1 are extracted, such as: male, bodyguard, superb martial arts, frontier bodyguard. The subject keywords contained in target copy 1 are extracted, such as: escort family, empire, superb martial arts. Next, taking target copy 2 as an example, through the content integration model, according to the above key visual information, the prompt information is extracted, and the object keywords contained in target copy 2 are extracted, such as: male, fisherman. The subject keywords contained in target copy 2 are extracted, such as: fishing village.

[0033] Object keywords may include attribute keywords and feature keywords of the target object. The attribute keywords are, for example, the gender of the target object, and the feature keywords are used to describe the appearance features, expression features, identity features, etc. of the target object. For example, the object keywords contained in the above-extracted target copywriting 1 are: male, bodyguard, superb martial arts, border bodyguard. Among them, the attribute keyword is: male, and the feature keywords are: bodyguard, superb martial arts, border bodyguard. Based on the object keywords contained in the target copywriting 1, the object key visual information matching the target object contained in the target copywriting 1 can be generated as: "Bodyguard, strong build, serious and alert, holding a spear, wearing leather armor, adopting a defensive posture.". Further, based on the above-extracted theme keywords in the target copywriting 1: bodyguard family, empire, superb martial arts, through the content integration model, the scene key visual information corresponding to the above object key visual information is generated as: "Standing against the background of an ancient Chinese street.". Based on this, the key visual information corresponding to the target copywriting 1 can be determined as: "Bodyguard, strong build, serious and alert, holding a spear, wearing leather armor, adopting a defensive posture, standing against the background of an ancient Chinese street.". Similarly, the object keywords contained in the above-extracted target copywriting 2 are: male, fisherman. Among them, the attribute keyword is: male, and the feature keyword is: fisherman. Based on the object keywords contained in the target copywriting 2, the object key visual information matching the target object contained in the target copywriting 2 can be generated as: "Fisherman, rugged and weathered face, holding a fishing net.". Further, based on the above-extracted theme keyword in the target copywriting 2: fishing village, through the content integration model, the scene key visual information corresponding to the object key visual information contained in the above target copywriting 2 is generated as: "Standing on a wooden boat, surrounded by water, fish, and reeds, with a peaceful and contented expression.". Based on this, the key visual information corresponding to the target copywriting 2 can be determined as: "Fisherman, rugged and weathered face, holding a fishing net, standing on a wooden boat, surrounded by water, fish, and reeds, with a peaceful and contented expression.".

[0034] S106: Through the image generation model, at least one target image matching each target copywriting is respectively generated according to the key visual information corresponding to each target copywriting and the preset image style prompt words; the image style prompt words represent the target style type that the generated target images uniformly have.

[0035] In the embodiment of the present application, after obtaining the key visual information corresponding to each target copywriting as above, through the image generation model, at least one target image matching each target copywriting can be respectively generated according to the key visual information corresponding to each target copywriting and the preset image style prompt words. Among them, the image style prompt words represent the target style type that the generated target images uniformly have.

[0036] In one embodiment, through an image generation model, at least one target image matching each target copywriting is generated respectively according to the key visual information corresponding to each target copywriting and a preset image style prompt, including: Through an image generation model, at least one target image matching each target copywriting is generated respectively according to the key visual information corresponding to each target copywriting, a preset image style prompt, and the image parameters corresponding to each target copywriting. The image parameters include an object style parameter and an image size parameter, and the image parameters are determined according to the attribute information of the target object included in the target copywriting and the graphic theme corresponding to the target copywriting.

[0037] In the embodiment of the present application, the image parameters corresponding to the target copywriting include an object style parameter and an image size parameter. Among them, the object style parameter can be used to describe the appearance characteristics of the target object, etc., and the image size parameter is used to describe the size of the generated target image. The attribute information of the target object can be the gender of the target object, and the image parameters of the target copywriting can be determined according to the attribute information of the target object and the graphic theme corresponding to the target copywriting.

[0038] In one example, the target copywriting 1 is: Gender: male, Family background: Escort family, Occupation: Escort head, Partner: Empire's chief songstress, Achievement: Becoming a legend of border bodyguards with superb martial arts skills. The target copywriting 4 is: Gender: female, Family background: Intellectual family, Occupation: Painter, Partner: General, Achievement: One of the four talented women. The genders of the target objects included in the target copywriting 1 and the target copywriting 4 are different. Then, the image parameters corresponding to the target copywriting 1 and the target copywriting 4 are different. For example, the image parameters corresponding to the target copywriting 1 include: Handsome male, Male, Long robe, etc. The image parameters corresponding to the target copywriting 4 include: Beautiful female, Female, Detailed eye description, Detailed facial description, etc.

[0039] In one embodiment, through an image generation model, at least one target image matching each target copywriting is generated respectively according to the key visual information corresponding to each target copywriting, a preset image style prompt, and the image parameters corresponding to each target copywriting, including: Performing weight parameter configuration processing on the preset image style prompt and the image parameters corresponding to each target copywriting respectively to obtain the configured image style prompt and the configured image parameters corresponding to each target copywriting; Through the image generation model, at least one target image matching each target copywriting is generated respectively according to the key visual information corresponding to each target copywriting, the configured image style prompt, and the configured image parameters corresponding to each target copywriting.

[0040] In the embodiments of the present application, when generating at least one target image that matches each target copywriting through an image generation model, first, weight parameter configuration processing can be performed on the preset image style prompt words and the image parameters corresponding to each target copywriting respectively to obtain the configured image style prompt words and the configured image parameters corresponding to each target copywriting; then, through the image generation model, at least one target image that matches each target copywriting is generated respectively according to the key visual information corresponding to each target copywriting, the configured image style prompt words, and the configured image parameters corresponding to each target copywriting. Among them, the target object included in the target copywriting is located in the target image corresponding to the target copywriting.

[0041] In the embodiments of the present application, performing weight parameter configuration processing on the preset image style prompt words and image parameters can be to configure the weight scaling factors for the preset image style prompt words and the image parameters corresponding to each target copywriting respectively. In an exemplary implementation manner, the preset image style prompt words are: ancient China, masterpiece, best quality, Chinese ink painting style, ultra-high detail. Continuing with the above target copywriting 1: Gender: male, Family background: Escort family, Occupation: Escort head, Partner: Chief imperial songstress, Achievement: Became a legend of border bodyguards with superb martial arts skills. The attribute information of the target object included in this target copywriting 1 is male, and the image parameters corresponding to target copywriting 1 are, for example: Handsome male, Male, Male focus, Robe. The above target copywriting 4: Gender: female, Family background: Literati family, Occupation: Painter, Partner: General, Achievement: One of the four talented women. The attribute information of the target object included in this target copywriting 4 is female, and the image parameters corresponding to target copywriting 4 are, for example: Highest picture quality, Detailed eye description, 4K resolution, Detailed facial description, Beautiful female, Female.

[0042] Further, when performing weight parameter configuration processing on the preset image style prompt words and the image parameters corresponding to the above target copywriting respectively, a weight scaling factor of 1 can be configured for the preset image style prompt word "masterpiece", and a weight scaling factor of 1 can be configured for the preset image style prompt word "best quality". The configured image style prompt words obtained are: (masterpiece: 1), (best quality: 1), Chinese ink painting style, ultra-high detail, ancient China. A weight scaling factor of 1.5 is configured for the image parameter "handsome male" corresponding to target copywriting 1, and a weight scaling factor of 1.2 is configured for the image parameter "male focus". Based on this, for target copywriting 1, the configured image parameters obtained are: (handsome male: 1.5), male, (male focus: 1.2), robe. When performing weight parameter configuration processing on the image parameters corresponding to target copywriting 4, a weight scaling factor of 1.5 can be configured for the image parameter "beautiful female". Based on this, for target copywriting 4, the configured image parameters obtained are: highest image quality, detailed eye description, 4K resolution, detailed facial description, (beautiful female: 1.5), female. Among them, the numbers corresponding to the above configured image style prompt words represent the weight scaling factors of the preset image style prompt words when generating images. Similarly, the numbers corresponding to the above configured image parameters represent the weight scaling factors of the image parameters when generating images.

[0043] After obtaining the configured image style prompt words, the configured image parameters corresponding to target copywriting 1, and the configured image parameters corresponding to target copywriting 4, the key visual information corresponding to target copywriting 1, the configured image parameters, and the configured image style prompt words, and the key visual information corresponding to target copywriting 4, the configured image parameters, and the configured image style prompt words are respectively input into the image generation model. Through the image generation model, at least one target image matching target copywriting 1 and at least one target image matching target copywriting 4 are respectively output. The target objects included in each target copywriting are respectively located in the target images matching the target copywriting. The target images output by the image generation model and matching target copywriting 1 are as Figure 2a shown, Figure 2a and the male in the above target copywriting 1 is included in the shown target image. The target images output by the image generation model and matching target copywriting 4 are as Figure 2b shown, Figure 2b and the female in the above target copywriting 4 is included in the shown target image.

[0044] S108: Determine each target copywriting and at least one target image matching each target copywriting as each target graphic generation result corresponding to the graphic theme.

[0045] In the embodiments of the present application, after generating at least one target image corresponding to each target copywriting, each target copywriting and at least one target image matching each target copywriting can be determined as each target graphic generation result corresponding to the graphic theme.

[0046] Continuing with the above example, target copywriting 1 and the target image as shown in Figure 2a can be determined as a target graphic generation result, and target copywriting 4 and the target image as shown in Figure 2b can be determined as a target graphic generation result.

[0047] In one embodiment, the method further includes: Obtaining a sample copywriting, a standard sample image corresponding to the sample copywriting, and label information; the label information includes the sample key visual information corresponding to the sample copywriting and a preset sample image style prompt; the sample style prompt represents the target style type that the generated target sample images uniformly have; Generating a target sample image matching the sample copywriting through a pre-built neural network according to the sample copywriting, the sample key visual information corresponding to the sample copywriting, and the sample image style prompt; Training the neural network according to the difference between the standard sample image corresponding to the sample copywriting and the target sample image generated by the neural network, and using the trained neural network as an image generation model.

[0048] In the embodiments of the present application, it is also possible to obtain a sample copywriting, a standard sample image corresponding to the sample copywriting, the sample key visual information corresponding to the sample copywriting, and a preset sample image style prompt. Then, input the sample copywriting, the standard sample image corresponding to the sample copywriting, the sample key visual information corresponding to the sample copywriting, and the preset sample image style prompt into the pre-built neural network, and output a target sample image matching the sample copywriting through the pre-built neural network. Next, perform iterative training on the pre-built neural network according to the standard sample image corresponding to the sample copywriting, the target sample image matching the sample copywriting, and the loss function until the difference between the target sample image matching the sample copywriting and the standard sample image corresponding to the sample copywriting is less than a preset difference threshold, and use the trained neural network as an image generation model.

[0049] In one implementation, the image generation model can be a trained Stable Diffusion model.

[0050] In the embodiments of the present application, first, through a copywriting generation model, multiple target copywritings matching the graphic and text theme are generated according to the pre-determined graphic and text theme and the copywriting prompt information customized based on the graphic and text theme; then, through a content integration model, each target copywriting generated by the copywriting generation model is analyzed to extract the key visual information corresponding to each target copywriting; then, through an image generation model, at least one target image matching each target copywriting is generated respectively according to the key visual information corresponding to each target copywriting and the preset image style prompt words; the image style prompt words represent the target style type that the generated target images uniformly have. Finally, each target copywriting and at least one target image matching each target copywriting are determined as each target graphic and text generation result corresponding to the graphic and text theme. It can be seen that through the embodiments of the present application, the generated copywriting and images are no longer monotonous, but a graphic and text series closely matching a certain specific graphic and text theme, and the style of each generated target copywriting and the target image corresponding to each target copywriting are consistent and the theme is unified. In this way, not only the style consistency of the graphic and text content is ensured, but also the efficiency and quality of the graphic and text content generation are greatly improved, thereby enhancing the user experience to a great extent. Figure 3 is a schematic flowchart of the target graphic and text generation process provided by the embodiments of the present application, as Figure 3 shown, the target graphic and text generation process specifically includes the following steps: S301: Determine the graphic and text theme and the copywriting prompt information matching the graphic and text theme.

[0051] S302: Input the copywriting prompt information matching the graphic and text theme into the large language model, and output multiple copywritings matching the graphic and text theme.

[0052] S303: For each copywriting, expand the copywriting through the large language model according to the copywriting expansion prompt information to obtain the target copywriting.

[0053] S304: For each target copywriting, extract the object keywords and theme keywords included in the target copywriting through the large language model.

[0054] S305: For each target copywriting, generate the key visual information corresponding to the target copywriting through the large language model according to the object keywords and theme keywords.

[0055] S306: Obtain the preset image style prompt words.

[0056] S307: For each target copywriting, input the key visual information corresponding to the target copywriting and the preset image style prompt words into the image generation model to generate at least one target image matching the target copywriting.

[0057] S308: For each target text, determine the target text and at least one target image that matches the target text as each target text-image generation result corresponding to the text-image theme.

[0058] The specific implementation process of the above text-image generation process can refer to the implementation processes of the above Figure 1 each embodiment, which will not be elaborated in this application.

[0059] The above is the text-image generation method provided by the embodiments of this application. Based on the same idea, the embodiments of this application also provide a text-image generation device. Figure 4 It is a schematic structural diagram of a text-image generation device provided by the embodiments of this application. As Figure 4 shown, the text-image generation device includes: A text generation module 401, configured to generate multiple target texts that match the text-image theme through a text generation model according to a pre-determined text-image theme and text prompt information customized based on the text-image theme; A content integration module 402, configured to analyze each target text generated by the text generation model through a content integration model, and extract key visual information corresponding to each target text; An image generation module 403, configured to generate at least one target image that matches each target text through an image generation model according to the key visual information corresponding to each target text and a preset image style prompt word; the image style prompt word represents the target style type that the generated target images uniformly have; A text-image generation module 404, configured to determine each target text and at least one target image that matches each target text as each target text-image generation result corresponding to the text-image theme.

[0060] In the embodiments of this application, the text prompt information customized based on the text-image theme includes text generation prompt information and text extension prompt information; the text generation module 401 is specifically configured to generate multiple initial texts that match the text-image theme through the text generation model according to the pre-determined text-image theme and the text generation prompt information; generate at least one target text corresponding to each initial text through the text generation model according to each initial text and the text extension prompt information.

[0061] In the embodiments of this application, the content integration module 402 is specifically configured to Through the content integration model, object keywords and theme keywords included in each of the target texts are extracted respectively; the object keywords are used to describe the target objects included in the target texts, and the theme keywords are used to describe the graphic themes corresponding to the target texts. Through the content integration model, according to the object keywords and the theme keywords included in each of the target texts, key visual information corresponding to each of the target texts is generated respectively.

[0062] In the embodiment of the present application, the content integration module 402 is further specifically configured to Through the content integration model, according to the object keywords included in each of the target texts, object key visual information matching the target objects included in each of the target texts is generated respectively. Through the content integration model, according to the theme keywords included in each of the target texts, scene key visual information corresponding to the object key visual information is generated respectively. Through the content integration model, for each of the target texts, the object key visual information and the scene key visual information corresponding to the target text are determined as the key visual information corresponding to the target text.

[0063] In the embodiment of the present application, the image generation module 403 is specifically configured to Through the image generation model, according to the key visual information corresponding to each of the target texts, the preset image style prompt words, and the image parameters corresponding to each of the target texts, at least one target image matching each of the target texts is generated respectively. The image parameters include object style parameters and image size parameters, and the image parameters are determined according to the attribute information of the target objects included in the target texts and the graphic themes corresponding to the target texts.

[0064] In the embodiment of the present application, the image generation module 403 is further specifically configured to Perform weight parameter configuration processing on the preset image style prompt words and the image parameters corresponding to each of the target texts respectively, to obtain the configured image style prompt words and the configured image parameters corresponding to each of the target texts. Through the image generation model, according to the key visual information corresponding to each of the target texts, the configured image style prompt words, and the configured image parameters corresponding to each of the target texts, at least one target image matching each of the target texts is generated respectively.

[0065] In the embodiment of the present application, a model training module is further included. The model training module is used to obtain sample texts, the standard sample images corresponding to the sample texts, and label information; the label information includes the sample key visual information corresponding to the sample texts, and preset sample image style prompt words; the sample style prompt words represent the target style types that the generated target sample images uniformly have. Through a pre-built neural network, generate a target sample image that matches the sample text according to the sample text, the sample key visual information corresponding to the sample text, and the sample image style prompt words. According to the difference between the standard sample image corresponding to the sample text and the target sample image generated by the neural network, train the neural network, and use the trained neural network as the image generation model.

[0066] The above text-image generation device provided by the embodiments of the present application can execute the text-image generation methods provided in the above method embodiments. For the specific process, please refer to the description in the above text-image generation method embodiments, and details are not described herein again.

[0067] In the embodiments of the present application, first, through a text generation model, generate multiple target texts that match the text-image theme according to a pre-determined text-image theme and text prompt information customized based on the text-image theme; then, through a content integration model, analyze each target text generated by the text generation model, and extract the key visual information corresponding to each target text; then, through an image generation model, generate at least one target image that matches each target text according to the key visual information corresponding to each target text and preset image style prompt words; the image style prompt words represent the target style types that the generated target images uniformly have. Finally, determine each target text and at least one target image that matches each target text as each target text-image generation result corresponding to the text-image theme. It can be seen that through the embodiments of the present application, the generated texts and images are no longer monotonous, but a series of texts and images that are closely matched with a specific text-image theme, and the style of each generated target text and the target image corresponding to each target text is consistent and the theme is unified. In this way, not only the style consistency of the text-image content is ensured, but also the efficiency and quality of the text-image content generation are greatly improved, thereby enhancing the user experience to a great extent.

[0068] The embodiments of the present application also provide an electronic device. Figure 5 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application. As Figure 5As shown in the figure, the device includes: a memory 501, a processor 502, a bus 503, and a communication interface 504. The memory 501, the processor 502, and the communication interface 504 communicate through the bus 503. The communication interface 504 may include an input / output interface, and the input / output interface includes, but is not limited to, a keyboard, a mouse, a display, a microphone, a loudspeaker, etc.

[0069] Figure 5 In the memory 501, there are computer-executable instructions that can run on the processor 502. When the computer-executable instructions are executed by the processor 502, the following processes are implemented: Through a copywriting generation model, according to a pre-determined graphic and text theme and copywriting prompt information customized based on the graphic and text theme, generate multiple target copywritings that match the graphic and text theme; Through a content integration model, analyze each target copywriting generated by the copywriting generation model, and extract the key visual information corresponding to each target copywriting; Through an image generation model, according to the key visual information corresponding to each target copywriting and a preset image style prompt word, respectively generate at least one target image that matches each target copywriting; the image style prompt word represents the target style type that the generated target images uniformly have; Determine each target copywriting and at least one target image that matches each target copywriting as each target graphic and text generation result corresponding to the graphic and text theme.

[0070] As can be seen from the technical solutions provided by the embodiments of the present application above, in the embodiments of the present application, first, through a copywriting generation model, according to a pre-determined graphic and text theme and copywriting prompt information customized based on the graphic and text theme, generate multiple target copywritings that match the graphic and text theme; then, through a content integration model, analyze each target copywriting generated by the copywriting generation model, and extract the key visual information corresponding to each target copywriting; then, through an image generation model, according to the key visual information corresponding to each target copywriting and a preset image style prompt word, respectively generate at least one target image that matches each target copywriting; the image style prompt word represents the target style type that the generated target images uniformly have. Finally, determine each target copywriting and at least one target image that matches each target copywriting as each target graphic and text generation result corresponding to the graphic and text theme. It can be seen that through the embodiments of the present application, the generated copywriting and images are no longer monotonous, but a graphic and text series that is closely matched to a certain specific graphic and text theme, and the style of each generated target copywriting and the target image corresponding to each target copywriting is consistent and the theme is unified. In this way, not only the style consistency of the graphic and text content is ensured, but also the efficiency and quality of the graphic and text content generation are greatly improved, thereby enhancing the user experience to a great extent.

[0071] The electronic device provided by the embodiment of the present application can implement each process in the foregoing embodiment of the graphic and text generation method, and achieve the same functions and effects, which will not be repeated here.

[0072] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process in the foregoing embodiment of the graphic and text generation method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0073] The embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements each process in the foregoing embodiment of the graphic and text generation method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0074] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.

[0075] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0076] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the process in Figure 1One or more processes and / or blocks Figure 1 The functions specified in one or more blocks.

[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 One or more processes and / or blocks Figure 1 The steps of the functions specified in one or more blocks.

[0078] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0079] Memory may include non-permanent memory in a computer-readable storage medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable storage medium.

[0080] Computer-readable storage media include permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media do not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0081] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0082] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0083] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and changes can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A method for generating images and texts, characterized in that: The method comprises: Generate multiple target texts matching the image and text theme through a text generation model according to a predetermined image and text theme and text prompt information customized based on the image and text theme; Analyzing each target copy generated by the copy generation model through a content integration model, and extracting key visual information corresponding to each target copy; At least one target image matching each target copy is generated respectively according to the key visual information corresponding to each target copy and the preset image style prompt word through the image generation model; the image style prompt word represents the target style type uniformly possessed by the generated target images; Each target text and at least one target image matching each target text are determined as each target graphic-text generation result corresponding to the graphic-text theme.

2. The method according to claim 1, characterized in that The text prompt information customized based on the graphic theme includes text generation prompt information and text extension prompt information; The copy generation model generates a plurality of target copies matching the image and text theme according to a predetermined image and text theme and copy prompt information customized based on the image and text theme, including: Generate multiple initial copies matching the image and text themes through the copy generation model according to the predetermined image and text themes and the copy generation prompt information; At least one target text corresponding to each of the initial texts is generated respectively through the text generation model according to each of the initial texts and the text extension prompt information.

3. The method according to claim 1, characterized in that The content integration model is used to analyze each target copy generated by the copy generation model, and key visual information corresponding to each target copy is extracted, including: Through the content integration model, object keywords and subject keywords contained in each target copy are extracted respectively; the object keywords are used to describe the target object contained in the target copy, and the subject keywords are used to describe the graphic subject corresponding to the target copy; Through the content integration model, key visual information corresponding to each target copy is generated according to the object keyword and the subject keyword contained in each target copy.

4. The method according to claim 3, characterized in that The key visual information corresponding to each target copy is generated respectively according to the object keyword and the subject keyword contained in each target copy by using the content integration model, including: By means of the content integration model, according to the object keywords contained in each target copy, object key visual information matching the target object contained in each target copy is generated respectively; Generate scene key visual information corresponding to the object key visual information respectively according to the subject keywords contained in each target copy by using the content integration model; By using the content integration model, for each target text, the object key visual information and the scene key visual information corresponding to the target text are determined as the key visual information corresponding to the target text.

5. The method according to claim 1, characterized in that The method of generating at least one target image matching each target copy respectively according to the key visual information corresponding to each target copy and the preset image style prompt words through the image generation model includes: Through the image generation model, at least one target image matching each target copy is generated according to the key visual information corresponding to each target copy, the preset image style prompt words, and the image parameters corresponding to each target copy. The image parameters include object style parameters and image size parameters, and the image parameters are determined according to the attribute information of the target object contained in the target copy and the graphic theme corresponding to the target copy.

6. The method according to claim 5, characterized in that The method of generating at least one target image matching each target copy respectively according to the key visual information corresponding to each target copy, the preset image style prompt words, and the image parameters corresponding to each target copy by using the image generation model includes: Performing weight parameter configuration processing on the preset image style prompt words and the image parameters corresponding to each target copy respectively, to obtain the configured image style prompt words and the configured image parameters corresponding to each target copy; Through the image generation model, at least one target image matching each target copy is generated according to the key visual information corresponding to each target copy, the configured image style prompt words, and the configured image parameters corresponding to each target copy.

7. The method according to claim 1, characterized in that The method also includes: Acquire a sample text, a standard sample image corresponding to the sample text, and label information; the label information includes sample key visual information corresponding to the sample text, and a preset sample image style prompt word; the sample style prompt word represents a target style type uniformly possessed by the generated target sample image; Generate a target sample image matching the sample text according to the sample text, the sample key visual information corresponding to the sample text, and the sample image style prompt words through a pre-built neural network; The neural network is trained according to the difference between the standard sample image corresponding to the sample text and the target sample image generated by the neural network, and the trained neural network is used as the image generation model.

8. A graphic and text generating device, characterized in that: The device comprises: A copywriting generation module, used to generate a plurality of target copies matching the image and text theme through a copywriting generation model according to a predetermined image and text theme and copywriting prompt information customized based on the image and text theme; A content integration module, used to analyze each target copy generated by the copy generation model through a content integration model, and extract key visual information corresponding to each target copy; An image generation module, configured to generate at least one target image matching each target copy respectively according to the key visual information corresponding to each target copy and a preset image style prompt word through an image generation model; the image style prompt word represents a target style type uniformly possessed by the generated target images; The image-text generation module is used to determine each target text and at least one target image matching each target text as each target image-text generation result corresponding to the image-text theme.

9. An electronic device, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the image and text generation method as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium having computer-executable instructions stored therein, characterized in that: When the computer executable instructions are executed by a processor, the steps of the image and text generation method according to any one of claims 1 to 7 can be implemented.