Story text-based cover image generation method and apparatus, and related product
By obtaining story text to determine the frame and using a generative model to generate cover images, the problem of inefficient selection of cover images is solved, and automated and efficient cover image generation and recommendation are achieved.
Patent Information
- Application Number
- CN202510466622.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the method of manually selecting the story cover image is inefficient, resulting in a long time to determine the cover image, which affects the story recommendation efficiency.
By obtaining the story text, determining the story frame, and using a generative model to generate the cover basemap description information, finally generating a matching cover image, and generating a cover image in combination with the cover text.
It realizes automatic generation of matching cover images, improving the efficiency of cover image generation and story recommendation.
Smart Images

Figure CN120339437A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, and related products for generating a cover image based on a story text. Background Art
[0002] Currently, users can publish stories written by themselves in reading applications for other users to read. A story needs to be equipped with a cover image, and the cover image of the story can be displayed on the story recommendation page to recommend the story to other users. In one case, the user selects a matching cover image for the story locally according to the content of the story and uploads it to the reading application. In another case, the operator of the reading application selects a matching cover image for the story in the image library according to the content of the story. However, selecting a cover image for a story manually takes a long time to select a matching cover image, and there is a problem of low efficiency in determining the cover image. Summary of the Invention
[0003] Embodiments of the present disclosure provide a method, an apparatus, and related products for generating a cover image based on a story text, which can automatically generate a matching cover image for a story, improve the generation efficiency of the cover image, and improve the recommendation efficiency of recommending a story based on the cover image.
[0004] In a first aspect, an embodiment of the present disclosure provides a method for generating a cover image based on a story text, including: Obtaining the story text of the story, and determining the story framework of the story according to the story text; wherein, the story framework is used to constrain the element content of the generation dimension of the cover base map of the story; Generating, by a generative model, description information of the cover base map of the story according to the story framework, and generating the cover base map of the story according to the description information of the cover base map; Obtaining the cover text of the story, and generating the cover image of the story based on the cover text and the cover base map; the cover image is used to recommend the story.
[0005] In a second aspect, an embodiment of the present disclosure provides an apparatus for generating a cover image based on a story text, including: An information generation unit, configured to obtain the story text of the story, and determine the story framework of the story according to the story text; wherein, the story framework is used to constrain the element content of the generation dimension of the cover base map of the story; A base map generation unit, configured to generate, by a generative model, description information of the cover base map of the story according to the story framework, and generate the cover base map of the story according to the description information of the cover base map; A cover generation unit, configured to obtain the cover text of the story and generate a cover image of the story based on the cover text and the cover base map; the cover image is used to recommend the story.
[0006] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor; and a memory configured to store computer-executable instructions, which when executed cause the processor to implement the method described in the first aspect above.
[0007] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, which is used to store computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the method described in the first aspect above.
[0008] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, which includes a computer program that, when executed by a processor, implements the method described in the first aspect above.
[0009] In this embodiment, first, obtain the story text of the story, determine the story framework of the story according to the story text, and the story framework is used to constrain the element content of the generation dimension of the cover base map of the story. Then, through a generative model, generate the description information of the cover base map of the story according to the story framework, and generate the cover base map of the story according to the description information of the cover base map. Finally, obtain the cover text of the story and generate a cover image of the story based on the cover text and the cover base map, and the cover image is used to recommend the story. It can be seen that through this embodiment, a matching cover image can be automatically generated for the story according to the story text of the story, the generation efficiency of the cover image of the story is improved, and the recommendation efficiency of recommending the story based on the cover image of the story is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in one or more embodiments of the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments recorded in the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts; Figure 1 It is a schematic flowchart of a method for generating a cover image based on a story text provided by an embodiment of the present disclosure; Figure 2 It is a schematic diagram of generating description information of a cover base map through a generative model provided by an embodiment of the present disclosure; Figure 3 It is a schematic diagram of a cover image of a story provided by an embodiment of the present disclosure; Figure 4 Schematic diagram of the cover image of the story provided by another embodiment of the present disclosure; Figure 5 Schematic diagram of the recommended card of the story provided by an embodiment of the present disclosure; Figure 6 Schematic diagram of the structure of the cover image generation device based on the story text provided by an embodiment of the present disclosure; Figure 7 Schematic diagram of the structure of the electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0011] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present disclosure, the following will clearly and completely describe the technical solutions in one or more embodiments of the present disclosure in conjunction with the accompanying drawings in one or more embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Based on one or more embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0012] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the relevant parties shall be informed of the type, scope of use, usage scenarios, etc. of the information involved in the present disclosure in an appropriate manner and obtain the authorization of the relevant parties in accordance with relevant laws and regulations.
[0013] For example, when receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0014] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0015] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0016] Embodiments of the present disclosure provide a method, apparatus, and related products for generating a cover image based on a story text, which can automatically generate a matching cover image for a story, improve the generation efficiency of the cover image, and improve the recommendation efficiency of recommending stories based on the cover image. This image generation method can be applied to the server side and executed by the server side, which includes both a single server form and a server cluster form.
[0017] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained, and the nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations.
[0018] Large Language Model (LLM) generally refers to a pre-trained language model with a large number of parameters and layers, such as the Generative Pre-trained Transformer (GPT) series and the Bidirectional Encoder Representations (BERT) series, etc. These models can learn the grammar, semantics, and context information of the language through pre-training on a large-scale corpus, thus providing strong support for various natural language processing tasks.
[0019] Artificial Intelligence refers to the ability of a computer system to perform tasks that typically require human intelligence, such as learning, reasoning, problem-solving, understanding language, recognizing images, making decisions, etc. It involves technologies and theories in multiple fields, including machine learning, deep learning, natural language processing, computer vision, etc., aiming to enable machines to have intelligent behaviors and capabilities similar to humans to solve various complex practical problems.
[0020] An AI drawing model refers to an image generation model based on artificial intelligence technology. By learning a large amount of image data and art style information, it can automatically generate art works or perform operations such as image style transfer. In the field of AI drawing, the text used to generate an image is called a prompt.
[0021] Prompt, in the fields of artificial intelligence and computer science, "prompt" usually refers to the input text or instruction provided to the model to guide the model to generate an output. For example, in a language model, a prompt is the text input by the user, and the model will generate a relevant response based on this text. In an AI drawing model, a prompt is a text description used to describe the content of an image, and the model creates a corresponding image according to this description.
[0022] GPT is a pre-trained language model based on the Transformer architecture. It can perform self-supervised learning on large-scale text corpora, learning the grammar, semantics, and context information of the language, thus providing strong support for various natural language processing tasks.
[0023] Figure 1 The flowchart of the cover image generation method based on the story text provided by an embodiment of the present disclosure is shown as Figure 1 shown. The process includes: Step S102, obtaining the story text of the story, and determining the story framework of the story according to the story text; wherein, the story framework is used to constrain the element content of the generation dimension of the cover background image of the story. Step S104, through a generative model, generating the description information of the cover background image of the story according to the story framework, and generating the cover background image of the story according to the description information of the cover background image. Step S106, obtaining the cover text of the story, and generating the cover image of the story based on the cover text and the cover background image; the cover image is used to recommend the story.
[0024] In this embodiment, first, obtain the story text of the story, determine the story framework of the story according to the story text, where the story framework is used to constrain the element content of the generation dimension of the cover background image of the story. Then, through a generative model, generate the description information of the cover background image of the story according to the story framework, and generate the cover background image of the story according to the description information of the cover background image. Finally, obtain the cover text of the story, and generate the cover image of the story based on the cover text and the cover background image, where the cover image is used to recommend the story. It can be seen that through this embodiment, it is possible to automatically generate a matching cover image for the story according to the story text of the story, improve the generation efficiency of the cover image of the story, and improve the recommendation efficiency of recommending the story based on the cover image of the story.
[0025] In this embodiment, the story can be a story of any length and any content, which is not limited here. In one example, the story is a short story of several thousand to tens of thousands of words uploaded by the author of the story in a reading application. When the author of the story uploads the story to the reading application, the cover image of the story may not be uploaded. When the server detects that the story does not have a cover image, it can generate a cover image for the story through the Figure 1 method flow in.
[0026] In the above step S102, first, the story text of the story is obtained. In one case, the story text of the story includes the title of the story and the entire text of the story. In another case, the story text of the story includes the title of the story and the first several words in the text of the story. For example, the story text of the story includes the title of the story and the first 4000 words in the text of the story. In the above two cases, the story text of the story may or may not include the category tag of the story. The category tag of the story can be a category tag set for the story by the author of the story, or it can be a category tag set for the story by the operator of the reading application. The category tags of the story can be exemplified by: Xianxia, ancient romance, modern workplace, etc.
[0027] In the above step S102, the story frame of the story is also determined according to the story text of the story. The story frame of the story is used to represent the general content and basic structure of the story. The story frame is also used to constrain the element content of the generation dimension of the cover background of the story. That is, the story frame in the present embodiment has the function of constraining the image content of the cover background. The cover background of the story is a part of the cover image that constitutes the story, and the cover image of the story at least includes the cover background and the cover text added to the cover background. For example, the cover image of the story is an image of a girl holding a rose and showing the title of the story "Road of Roses", then the image of the girl holding a rose is the cover background, and the title "Road of Roses" is the cover text.
[0028] In this embodiment, the story framework of the story can be determined according to the story text of the story through a generative model or not. The generative model can be a model with multimodal content generation function including a large language model. For example, the generative model includes a large language model and an AI drawing model, and has the ability to generate text and images. The generative model used to generate the story framework and the generative model used to generate the cover background description information can be the same model or different models.
[0029] In one embodiment, determining the story framework of a story according to the story text of the story includes: Determine the story dimensions used to represent the story framework, and determine the elements and contents of the story dimensions based on the story text; the story dimensions include subject object dimension, era background dimension, regional background dimension, story category dimension, and story emotional atmosphere dimension; Combine the elements of each story dimension to get the story framework.
[0030] Taking the generation of a story framework by a generative model as an example. In this embodiment, the story text of the story is input into the generative model. The generative model first determines each story dimension for representing the story framework of the story. Each story dimension can be pre-configured in the generative model. The following five dimensions are included in each story dimension: the main object dimension, the era background dimension, the geographical background dimension, the story category dimension, and the story emotional atmosphere dimension.
[0031] Then, through the generative model, according to the story text of the story, the element content of each story dimension is determined. The element content of the above five story dimensions is all learned and generated by the generative model based on the story text of the story. The element content of the story category dimension is not the category label pre-input into the generative model. Among them, the element content of the main object dimension includes the relevant information of the main object in the story. When the main object in the story is a person, the element content of the main object dimension includes but is not limited to the names, ages, genders, appearances, etc. of each main character in the story. When the main object in the story is not a person but an item, the element content of the main object dimension includes but is not limited to the names, external characteristics, functions, roles, etc. of each main item in the story. Of course, the main object in the story may also be an animal, and the element content of the main object dimension includes but is not limited to the names, ages, genders, personality characteristics, etc. of each main animal in the story. The element content of the era background dimension is used to represent the era characteristics when the story occurs, and can be exemplified as ancient times, modern times, etc. The element content of the geographical background dimension is used to represent the geographical characteristics when the story occurs, and can be exemplified as XX region, etc. The element content of the story category dimension is used to represent the category of the story, and can be exemplified as fantasy, science fiction, xianxia, romance, etc. The element content of the story emotional atmosphere dimension is used to represent the emotional type of the story, and can be exemplified as comedy, tragedy, etc.
[0032] When there are terms directly describing the element content of each story dimension in the story text of the story, the element content of each story dimension can be extracted from the story text of the story through the generative model. When there are no terms directly describing the element content of each story dimension in the story text of the story, the element content of each story dimension can be summarized by the generative model according to the story text of the story. Finally, through the generative model, the element content of each story dimension is combined to obtain the story framework of the story.
[0033] Appropriate prompt words can be input into the generative model to prompt the generative model to accurately determine the elements of each story dimension based on the story text and then generate the story framework of the story. In one example, the prompt words include: "You are an excellent content summary expert. You can summarize the relevant information of the main object in the story based on the story text, such as name, age, gender, appearance, etc. The main object of the story may also be an object or an animal. You need to determine the relevant information of the main object based on the actual situation. This information should be as consistent as possible with the original text. You can also determine the era background of the story, such as ancient or modern. The era background can accurately reflect the era characteristics of the story. You can also determine the regional background of the story, such as XX region, which can accurately reflect the regional characteristics of the story. You can also determine the story category, such as fantasy and fairy tales, and determine the emotional atmosphere of the story, such as comedy. The story category and emotional atmosphere must be consistent with the content and theme of the story. It should be noted that the above determined content needs to be highly consistent with the original text. Based on the above content, determine the story framework of the story."
[0034] In a specific example, the story text of the story is: Title: "The Road to Academic Excellence", part of the text: Zhao Yanyan originally worked as a security guard at University A. While working, she prepared for the entrance examination to University A with the help of Teacher Wang from University A. She was successfully admitted to University A as a college student, and later she was admitted to University A as a graduate student in a popular major.
[0035] Through the generative model, the story dimensions are determined to include the subject object dimension, the era background dimension, the regional background dimension, the story category dimension, and the story emotional atmosphere dimension. According to the story text, the elements of the subject object dimension are determined to include: Name: Zhao Yanyan, Age: 25-28 years old, Gender: Female, Appearance: Long hair, delicate features, decent clothes, Relationship: The protagonist of the story, who was admitted to university and then to graduate school from security guard; Name: Teacher Wang, Age: 45-50 years old, Gender: Male, Appearance: Tall figure, three-dimensional features, Relationship: Zhao Yanyan's teacher at University A. The elements of the era background dimension are determined to include: Modern. The elements of the regional background dimension are determined to include: University A in XX region. The elements of the story category dimension are determined to include: Modern city. The elements of the story emotional atmosphere dimension are determined to include: Inspirational struggle. Finally, through the generative model, the elements of the above story dimensions are combined to obtain the story framework of the story.
[0036] It can be seen that through this embodiment, it is possible to determine the various story dimensions used to represent the story framework of the story, determine the element content of each story dimension according to the story text of the story, combine the element content of each story dimension to obtain the story framework of the story, and generate the story framework of the story from each story dimension, so that the story framework of the story can accurately reflect the core content of the story, thereby improving the accuracy of generating the cover image of the story based on the story framework of the story.
[0037] In one embodiment, determining the story framework of a story based on the story text includes: Determining the target character of the story based on the story text; the cover image includes the target character; Generating a story framework based on the content in the story text used to describe the target character; this content is used to describe the appearance information and clothing information of the target character.
[0038] Taking the generation of the story framework by a generative model as an example. In this embodiment, the story text of the story is input into the generative model. The generative model first determines the target character of the story based on the story text. The target character can be the main character in the story. The target character is used to be displayed in the cover image, and the cover background image can be the character image of the target character.
[0039] Then, through the generative model, obtain the content in the story text used to describe the target character. This content is used to describe the appearance information and clothing information of the target character. For example, obtain the text in the story text used to describe the appearance information of the target character, and obtain the text in the story text used to describe the clothing information of the target character. The appearance information includes but is not limited to information such as the appearance characteristics and expressions of the target character.
[0040] Finally, through the generative model, generate a story framework based on the obtained content. For example, use the obtained content as the story framework, or combine the obtained content with the element content of the era background dimension, the element content of the regional background dimension, the element content of the story category dimension, and the element content of the story emotional atmosphere dimension of the story to form a story framework.
[0041] It can be seen that through this embodiment, it is also possible to start from the target character in the story, obtain the content in the story text used to describe the appearance information and clothing information of the target character, and generate a story framework based on this content, so that the cover background image generated based on the story framework can accurately depict the appearance and clothing of the target character, and improve the matching degree between the cover background image and the story.
[0042] In one embodiment, generating a story framework based on the content in the story text used to describe the target character includes: If the number of target characters is one, generate a story framework based on the appearance information and clothing information of the target character described in the story text; If the number of target characters is multiple, generate a story framework based on the first content used to describe each target character itself and the second content used to describe the target characters; the first content includes the appearance information and clothing information of the target character; the second content includes the pose information between the target characters.
[0043] In this embodiment, first, the number of target characters is determined. If the number of target characters is one, a story framework is generated according to the appearance information and clothing information of the target character described in the story text. For example, the appearance information and clothing information of the target character described in the story text are used as the story framework, or the appearance information and clothing information of the target character described in the story text, as well as the element content of the story's era background dimension, the element content of the regional background dimension, the element content of the story category dimension, and the element content of the story's emotional atmosphere dimension are combined together as the story framework.
[0044] If the number of target characters is multiple, a story framework is generated according to the appearance information and clothing information of each target character itself and the pose information between each of the described target characters. The story framework includes the appearance information and clothing information of each target character itself, and also includes the pose information between each target character. The story framework can also include at least one of the above-mentioned element content of the story's era background dimension, the element content of the regional background dimension, the element content of the story category dimension, and the element content of the story's emotional atmosphere dimension.
[0045] It can be seen that in this embodiment, the situation where the main characters of the story are single or multiple is also considered. When it is single, a story framework is generated based on the appearance information and clothing information of the single character. When it is multiple, a story framework is generated according to the appearance information and clothing information of each person among the multiple characters, as well as the pose information between the multiple characters, so that the cover bottom map generated based on the story framework can accurately depict the appearance and clothing of each target character, and can also depict the pose position relationship between each target character, improving the matching degree between the cover bottom map and the story.
[0046] In one embodiment, a story framework is generated according to the content used to describe the target character in the story text, including: Determine the story category of the story according to the story text, and determine the image style of the cover bottom map of the story according to the story category; Generate guiding information for the image style, and generate a story framework according to the guiding information; the guiding information is used to guide the generative model to generate style information of the image style as part of the cover bottom map description information.
[0047] The story framework can be generated with or without a generative model. It can be understood that the generative model used to generate the story framework and the generative model used to generate the cover bottom map description information can be the same model or different models.
[0048] In this embodiment, first, the story category of the story is determined according to the story text. The story category can be, for example, a romance category, a modern workplace category, etc. The story category can be the story category in the above-mentioned story category dimension. Then, according to the story category, the image style of the cover background image of the story is determined. The image style can be, for example, ink painting style, comic style, cartoon style, realistic style, etc. Next, guiding information for the image style is generated, and a story framework is generated according to the guiding information. The guiding information is carried in the story framework, and the element content of one or more of the above-mentioned story dimensions is also carried. The guiding information can be a guiding statement, which is used to guide the generative model to generate the style information of the above-determined image style when generating the cover background image description information, and use the style information as part of the cover background image description information. For example, the guiding information can be "This story is suitable for a cover in ancient style", so as to guide the generative model to generate "Image style - ancient style" as part of the cover background image description information.
[0049] It can be seen that through this embodiment, the image style of the cover background image can also be determined according to the story category, and a story framework is generated according to the image style, so that when the generative model generates the cover background image description information according to the story framework, the style information of the above-determined image style can be carried in the cover background image description information, so that the image style of the generated cover background image matches the story category.
[0050] In the above step S104, through the generative model, the cover background image description information of the story is generated according to the story framework, and the cover background image of the story is generated according to the cover background image description information. The generative model for generating the cover background image description information and the generative model for generating the story framework can be the same model or different models. In this embodiment, when the generative model for generating the cover background image description information has an image generation function, the cover background image of the story can be generated through this generative model according to the cover background image description information. Of course, the cover background image of the story can also be generated through other AI drawing models according to the cover background image description information.
[0051] In this embodiment, the cover background image description information of the story, as a prompt for generating the cover background image of the story, can accurately describe and depict the cover background image of the story. Continuing with the above example, the cover background image description information of the story can be, for example: "A young female student in a certain XX area, dressed as a college student, with a firm face, standing sideways in front of the teaching building of the university. The background is the teaching building of the university. She is holding a textbook in her hand, with a confident smile on her face, looking into the distance. The overall tone is mainly fresh and beautiful natural colors and outdoor sunny light, reflecting the relaxed and pleasant learning atmosphere of the university. The female protagonist is located in the center of the picture, with a front parallel view for the audience, and a modern urban university campus background, in the style of anime".
[0052] In one embodiment, through a generative model, the description information of the cover background image of the story is generated according to the story framework, including: Through the generative model, the information content requirements corresponding to the description information of the cover background image are obtained, and the example information corresponding to the description information of the cover background image is obtained; the information content requirements are used to indicate the contents that the description information of the cover background image needs to describe and the description requirements for each content. Through the generative model, content expansion is performed based on the story framework to obtain the description information of the cover background image that meets the information content requirements and matches the example information.
[0053] In this embodiment, when the story framework and the description information of the cover background image are generated by the same generative model, after determining the story framework of the story, the generative model may not output the story framework of the story, but directly output the description information of the cover background image after generating the description information of the cover background image. When generating the description information of the cover background image, the generative model obtains the information content requirements corresponding to the description information of the cover background image, and the information content requirements can be pre-configured in the generative model. The information content requirements are used to indicate the contents that the description information of the cover background image needs to describe and the description requirements for each content. For example, the information content requirements are used to indicate that the description information of the cover background image needs to describe the foreground content and the background environment of the cover background image, and the main features of the foreground content need to be described, and the specific scene and the environmental atmosphere of the background environment need to be described. Among them, the main features of the foreground content include but are not limited to the appearance, clothing, posture, and expression of the character, etc., the specific scene of the background environment includes but is not limited to elements such as buildings involved in the background, and the environmental atmosphere of the background environment can be exemplified as relaxed and pleasant, sad, etc.
[0054] The generative model also obtains the example information corresponding to the description information of the cover background image that is pre-input or pre-configured. The example information is a standardized example of the description information of the cover background image. For example, the example information is: "A male in a certain XX area stands on campus, and there is a teaching building behind him. The boy has a complex expression, and there is a little confusion and determination in his eyes. Mainly in blue and gray, with a slightly gloomy light atmosphere. The male protagonist is located in the upper right central position of the picture, and the viewer's front parallel perspective, and the background is a blurred campus building. Modern XX area campus scene, anime style".
[0055] Then, through the generative model, content expansion is performed based on the story framework to obtain the description information of the cover background image that meets the information content requirements and matches the example information. The generative model has excellent text understanding and text writing capabilities. After learning and understanding the contents indicated by the information content requirements and the description requirements for each content, content expansion can be performed based on the story framework, and by imitating the expression mode of the example information, the description information of the cover background image that meets the information content requirements and matches the example information is generated.
[0056] It can be understood that the content of the story framework is relatively small, and it does not necessarily describe all the items indicated by the information content requirements. For example, it may not describe the background environment. Therefore, it is necessary for the generative model to expand based on the content of the story framework under the condition of conforming to the story theme represented by the story framework, so as to obtain the cover bottom map description information that meets the information content requirements and matches the example information.
[0057] Referring to the previous example, the specific scene and environmental atmosphere of the background environment where the female protagonist is located are not described in the story framework, and the posture and expression of the female protagonist are not described either. Under the condition of conforming to the story theme represented by the story framework, based on the content of the story framework, it can be expanded that the female protagonist can stand in front of the teaching building of the university, holding textbooks in her hand, with a confident smile on her face, looking into the distance, indicating the female protagonist's love for learning, and the specific scene of the background environment is the background of a modern urban university campus, and the environmental atmosphere of the background environment is relaxed and pleasant.
[0058] It can be seen that through this embodiment, through the generative model, the information content requirements corresponding to the cover bottom map description information are obtained, and the example information corresponding to the cover bottom map description information is obtained. Based on the story framework, content expansion is carried out to obtain the cover bottom map description information that meets the information content requirements and matches the example information. By providing the example information, the accuracy of the generated cover bottom map description information is improved, and the effect of accurately generating the cover bottom map description information is achieved.
[0059] In one embodiment, through the generative model, content expansion is carried out based on the story framework to obtain the cover bottom map description information that meets the information content requirements and matches the example information, including: Through the generative model, each sub-requirement included in the information content requirements is determined, and the information segment corresponding to the sub-requirement in the example information is determined; the sub-requirement is used to indicate a piece of content and the description requirements for this content; Through the generative model, content expansion is carried out based on the story framework to obtain the sub-description information that meets the sub-requirement and matches the information segment corresponding to the sub-requirement; Through the generative model, according to each sub-description information, the cover bottom map description information is obtained.
[0060] In this embodiment, the information content requirements are used to indicate the various items that the cover bottom map description information needs to describe and the description requirements for each item. Based on this, the information content requirements can be divided into multiple sub-requirements, and each sub-requirement is used to indicate an item that the cover bottom map description information needs to describe and the description requirements for this item. In one example, the specific sub-requirements included in the information content requirements are: (1)Foreground content of the cover background image: including the main features of the foreground content, such as the appearance, clothing, posture, and expression of the character, etc.; (2)Background environment of the cover background image: including the specific scene of the background environment, such as elements like buildings involved in the background, and also including the environmental atmosphere of the background environment, such as relaxed and pleasant, sad, etc.; (3)Image style of the cover background image: used to represent the artistic style of the image, such as anime, watercolor painting, ancient style, stick figure, light and shadow, etc.; (4)Image lighting effect of the cover background image: used to represent the lighting of the image, such as side lighting, top lighting, warm-colored light, dim light, natural light, etc.; (5)Image tone effect of the cover background image: used to represent the color information of the image, such as metallic tone, cold tone, warm tone, Maillard color system, macaron color system, etc.; (6)Image composition of the cover background image: used to represent the aspect ratio of the image and the position of the main body of the image, such as the picture format is 3:4 ratio, and the character is located in the middle of the picture; (7)Image perspective of the cover background image: used to represent the perspective that the user feels when viewing the image, such as the viewer's perspective is looking up, the viewer's perspective is looking down, etc.; (8)Image era background of the cover background image: used to represent the era background in which the image is located, such as the image is a modern image; (9)Image regional background of the cover background image: used to represent the regional background in which the image is located, such as the buildings in the image are in the style of XX region.
[0061] By setting each sub-requirement in the information content requirements, the generated description information of the cover background image can meet the above-mentioned each sub-requirement, enabling the description information of the cover background image to accurately depict the cover background image and improving the accuracy of the generated cover background image.
[0062] In this embodiment, after determining the above-mentioned each sub-requirement, the information segments corresponding to each sub-requirement in the example information are also determined, that is, the information segments that meet each sub-requirement are determined. In the example information: "A male in XX region stands on campus, and there is a teaching building behind him. The boy has a complex expression, with a bit of confusion and determination in his eyes. Mainly in blue and gray, with a slightly gloomy light atmosphere. The male protagonist is located in the middle right position of the picture, and the viewer's front parallel perspective, and the background is blurred campus buildings. Modern XX region campus scene, anime style", in this example: The information segment that meets sub-requirement (1) is: A male in XX region stands on campus, and the boy has a complex expression, with a bit of confusion and determination in his eyes; The information segment that meets sub-requirement (2) is: There is a teaching building behind him, and the background is blurred campus buildings; The information segments that meet sub-requirements (3), (4), and (5) are: anime style, mainly in blue and gray, with a slightly gloomy light atmosphere; The information segments that meet sub-requirements (6) and (7) are: The male protagonist is located slightly to the right of the center of the picture, with a frontal parallel view for the audience; The information segments that meet sub-requirements (8) and (9) are: A campus scene in a modern XX area.
[0063] Next, through the generative model, content expansion is carried out based on the story framework to obtain sub-description information that meets each sub-requirement and matches the information segments corresponding to each sub-requirement. In this step, starting from each sub-requirement and the information segment corresponding to each sub-requirement, the work of generating a large long text is simplified to generating individual short texts. The generative model only needs to, under the condition of conforming to the story theme represented by the story framework, generate sub-description information that meets each sub-requirement and matches the information segments corresponding to each sub-requirement by understanding each sub-requirement and imitating the expression way of each information segment.
[0064] It can be understood that the content in the sub-description information may come from the original meaning of the story framework or may need to be expanded based on the story framework around the story theme. For example, in the previous example, the part "The overall color tone is mainly fresh and bright natural colors and outdoor sunny light, reflecting the relaxed and happy atmosphere of university life and study. The female protagonist is located in the center of the picture, with a frontal parallel view for the audience, against the background of a modern urban university campus, in anime style" is the sub-description information generated by the generative model based on the story framework of the story, the above sub-requirements (2)-(9), and the corresponding information segments, by imitating the expression way of the information segments and expanding the content based on the story framework.
[0065] In an example, for each sub-requirement and the corresponding information segment, content expansion can be first carried out based on the story framework of the story to generate content that meets the sub-requirement, and this content matches the story framework of the story. Then, according to the expression way of the corresponding information segment, the expression way of the generated content is determined. The expression way refers to the writing style, word-using habits, etc. Based on the determined expression way, the generated content is expressed to obtain the corresponding sub-description information.
[0066] In one example, when generating sub-description information that meets the above sub-requirement (3) and matches the corresponding information segment, first, according to the above sub-requirement (3) and the story framework of the story, determine the image style identifier that matches the story framework among various preset image style identifiers. Then, the generative model determines the above-matched image style identifier and the mapped image style from the mapping relationship between the preset image style identifiers and the image styles. Based on the determined image style, imitate the expression way of the information segment to generate sub-description information that meets the above sub-requirement (3) and matches the corresponding information segment.
[0067] In this example, the image style identifiers include: general, realistic emotion, brainstorm, historical martial arts, and the corresponding mapped image styles are: anime, semi-thick painting method; anime, semi-thick painting method, lighting style; anime, semi-thick painting method, lighting style; anime, ancient style, watercolor painting, lighting style.
[0068] Finally, through the generative model, according to each sub-description information, obtain the cover bottom image description information. For example, combine each sub-description information to obtain the cover bottom image description information, or combine each sub-description information and perform modification and arrangement in text expression, such as removing duplicate semantics, to obtain the cover bottom image description information.
[0069] It can be seen that through this embodiment, by setting each sub-requirement in the information content requirement, the generated cover bottom image description information can meet the above sub-requirements, enabling the cover bottom image description information to accurately depict the cover bottom image and improving the accuracy of the generated cover bottom image. Moreover, starting from each sub-requirement and the information segment corresponding to each sub-requirement, the work of generating a large long text is simplified into the work of generating short texts one by one, improving the coherence and generation efficiency of the generative model in generating the cover bottom image description information and the accuracy of the generated cover bottom image.
[0070] In one embodiment, through the generative model, based on the story framework for content expansion, to obtain sub-description information that meets the sub-requirements and matches the information segments corresponding to the sub-requirements, including: Through the generative model, obtain the information inference order represented by the thinking chain strategy for generating the cover bottom image description information; the information inference order includes the order of reasoning for each sub-requirement. Through the generative model, according to the information inference order, based on the story framework for content expansion, sequentially generate sub-description information that meets each sub-requirement and matches the information segments corresponding to the sub-requirements.
[0071] In this embodiment, a prompt can be input to the generative model, prompting the generative model to gradually think and generate the description information of the cover bottom image based on the Chain of Thought (COT) strategy. In this embodiment, the recommended information reasoning order can also be set in the prompt, and the information reasoning order includes the order of reasoning for each sub-requirement. For example, the information reasoning order is: in the order of the above sub-requirements (1)-(9), the sub-description information that meets each sub-requirement is inferred. Based on this, the prompt can be: "Please output the description information of the cover bottom image of the story according to the set order and story framework, think step by step and output the thinking process, and finally output the description information of the cover bottom image. The set order includes: think about the foreground content of the cover bottom image, the background environment of the cover bottom image, the image style of the cover bottom image, the image lighting effect of the cover bottom image, the image color tone effect of the cover bottom image, the image composition of the cover bottom image, the image perspective of the cover bottom image, the image era background of the cover bottom image, and the image regional background of the cover bottom image in sequence."
[0072] Based on this, the generative model can obtain the above information reasoning order, and through the way of gradually thinking with the Chain of Thought strategy, according to the above information reasoning order, expand the content based on the story framework, and sequentially generate sub-description information that meets each sub-requirement and matches the information segment corresponding to the sub-requirement.
[0073] It can be seen that through this embodiment, the Chain of Thought strategy can also be adopted, and the reasoning order of each sub-requirement can be set in the Chain of Thought strategy, so that the generative model can reason each sub-requirement and the corresponding information segment according to a certain order, sequentially generate each sub-description information, improve the thinking depth of the generative model, improve the generation coherence of the sub-description information, avoid hallucinations caused by the disordered reasoning of the generative model, and improve the accuracy of the generated sub-description information.
[0074] In one embodiment, through the generative model, according to the information reasoning order, expand the content based on the story framework, and sequentially generate sub-description information that meets each sub-requirement and matches the information segment corresponding to the sub-requirement, including: Through the generative model, according to the information reasoning order, determine the sub-requirement to be inferred, and determine the target content corresponding to the sub-requirement in the story framework; Through the generative model, expand the content based on the target content, and generate sub-description information that meets the sub-requirement and matches the information segment corresponding to the sub-requirement.
[0075] In this embodiment, when generating each sub-description information in sequence according to the above information reasoning order through the generative model, the generative model can determine the current sub-requirement to be inferred according to the above information reasoning order, and determine the target content corresponding to the sub-requirement in the story framework. The target content is a content segment in the story framework that can be expanded to obtain content that meets the sub-requirement. Then, based on the target content, content expansion is performed to generate content that meets the sub-requirement. This content matches the target content. Then, according to the expression mode of the corresponding information segment, the expression mode of the generated content is determined, and the generated content is expressed based on the expression mode of the generated content to obtain the corresponding sub-description information. The generative model can reason about each sub-requirement one or more times in sequence according to the information reasoning order to improve the accuracy of generating sub-description information.
[0076] Taking the above sub-requirement (1) as an example, for sub-requirement (1), the generative model can determine in the previous story framework example that the target content is: "Name: Zhao Yanyan, Age: 25 - 28 years old, Gender: Female, Appearance characteristics: Long hair, delicate features, well-dressed, Relationship: The protagonist of the story, admitted to university from a security guard and then admitted to graduate school". Then, according to sub-requirement (1), the generative model expands based on the target content to determine the content that meets sub-requirement (1) as: A young female college student, dressed like a college student, with a firm expression, standing sideways in front of the university teaching building, holding textbooks in her hands, smiling, and looking into the distance. Finally, by imitating the expression mode of the information segment that meets sub-requirement (1), the sub-description information "A young female student in a certain area, dressed like a college student, with a firm expression, standing sideways in front of the university teaching building, holding textbooks in her hands, with a confident smile on her face, looking into the distance" is generated. No more examples will be repeated here for other sub-requirements.
[0077] It can be seen that through this embodiment, when generating each sub-description information according to the set information reasoning order, the current sub-requirement to be inferred can be determined, the target content corresponding to the sub-requirement can be determined in the story framework, content expansion is performed based on the target content, and sub-description information that meets the sub-requirement and matches the information segment corresponding to the sub-requirement is generated. By determining the target content in the story framework, the text range that the generative model needs to process is narrowed, so that the generative model can carefully analyze based on the target content and accurately generate sub-description information.
[0078] Figure 2 Schematic diagram of generating cover bottom map description information through a generative model provided by an embodiment of the present disclosure. As Figure 2 shown, the above-mentioned various story dimensions, information content requirements, and example information corresponding to the cover bottom map description information can be pre-configured in the generative model, and the story text, the first prompt word, and the second prompt word of the story are all input into the generative model.
[0079] The first prompt word is used to generate the story framework of the story. An example is "You are an excellent content summary expert. You can summarize the relevant information of the main object in the story based on the story text, such as name, age, gender, appearance, etc. The main object of the story may also be an object or an animal. You need to determine the relevant information of the main object based on the actual situation. This information should be as consistent as possible with the original text. You can also determine the historical background of the story, such as ancient or modern times. The historical background can accurately reflect the historical characteristics of the story. You can also determine the geographical background of the story, such as XX area. The geographical background can accurately reflect the geographical characteristics of the story. You can also determine the story category, such as fantasy and fairy tales, and determine the emotional atmosphere of the story, such as comedy. The story category and emotional atmosphere must be consistent with the content and theme of the story. It should be noted that the content determined above needs to be highly consistent with the original text. Based on the above content, determine the story framework of the story."
[0080] The second prompt word is used to generate the cover background description information of the story. An example is "Please output the cover background description information of the story according to the set order and story framework. Please think step by step and output the thinking process, and finally output the cover background description information. The set order includes: think about the foreground content of the cover background, the background environment of the cover background, the image style of the cover background, the image lighting effect of the cover background, the image tone effect of the cover background, the image composition of the cover background, the image perspective of the cover background, the image era background of the cover background, and the image regional background of the cover background."
[0081] Figure 2 In the example, the first prompt word and the second prompt word can also be combined into a prompt word input. The generative model generates and outputs the cover and background image description information of the story according to the pre-configured story dimensions, information content requirements and sample information corresponding to the cover and background image description information through the method flow described above.
[0082] After obtaining the cover background description information of the story, in the above step S104, the cover background description information is used as a prompt word for generating the cover background of the story and input into the image generation model. The cover background of the story is generated according to the cover background description information through the image generation model. The image generation model can be integrated into a generative model for generating cover background description information. The image generation model can be an AI drawing model that can generate a corresponding image based on the input prompt word.
[0083] In one embodiment, multiple image generation models may be provided, each image generation model being used to generate an image belonging to an image style. According to the image style of the cover base image in the cover base image description information, the cover base image description information is input into the corresponding image generation model to obtain the cover base image.
[0084] To improve the accuracy of the generated cover background image and avoid incorrect model output, in one embodiment, after obtaining the cover background image description information, the cover background image description information can also be compared with the above-mentioned information content requirements and example information. If it is determined through comparison that the cover background image description information meets the above-mentioned information content requirements and the expression manner matches the above-mentioned example information, the cover background image description information is used as a prompt word and input into the image generation model. If it is determined through comparison that the cover background image description information does not meet the above-mentioned information content requirements or the expression manner does not match the above-mentioned example information, the cover background image description information is adjusted, such as adjusting the language to be consistent with the example information, supplementing the missing sub-description information, etc. The adjusted cover background image description information is used as a prompt word and input into the image generation model.
[0085] As mentioned above, the cover background image of the story is a part of the cover image of the story. The cover image of the story at least includes the cover background image and the cover text added to the cover background image. Based on this, after obtaining the cover background image, in the above step S106, the cover text of the story is also obtained. The cover text of the story includes but is not limited to at least one of the title of the story, classic famous quotes, content summary, and the protagonist's catchphrase. Then, based on the cover text and the cover background image, the cover image of the story is generated.
[0086] In one embodiment, generating the cover image of the story based on the cover text and the cover background image includes: Determining the combination manner of the cover text and the cover background image according to the distribution scenario of the cover image of the story; Combining the cover text and the cover background image according to the combination manner to obtain the cover image.
[0087] In this embodiment, the distribution scenario of the cover image is determined. The distribution scenario includes the display manner of the cover image on the page, such as distributing the cover image in a single column on the page, or distributing the cover image in multiple columns on the page. Among them, single-column distribution means that the page is used as one column, and one cover image is displayed in one column. Multiple-column distribution means that the page is divided into multiple columns, and each column displays the cover image of its respective column. Multiple-column distribution can be exemplified by double-column distribution.
[0088] The distribution scenario also includes the distribution position of the cover image, such as distributing it to the bookshelf, or distributing it to the book promotion page, or distributing it in the book promotion data, and the book promotion data includes any one of book promotion posts, book promotion graphic and text data, and book promotion video data.
[0089] Determine the combination method of the cover text and the cover background image according to the distribution scenario of the cover image. A mapping relationship between the distribution scenario of the cover image and this combination method can be established in advance, so as to determine the combination method of the cover text and the cover background image according to this mapping relationship and the distribution scenario of the cover image. For example, when the cover image is distributed in two columns, or when the cover image is distributed on the book promotion page or in the book promotion data, set the combination method such that the cover text is located outside and below the cover background image, and set a common background image for the cover text and the cover background image. When the cover image is distributed in a single column, or when the cover image is distributed on the bookshelf, set the combination method such that the cover text is located on the cover background image.
[0090] Furthermore, combine the cover text and the cover background image according to the combination method to obtain the cover image. For example, set the cover text outside and below the cover background image, and set a common background image for the cover text and the cover background image to obtain the cover image. Another example is to set the cover text on the cover background image, and the set cover background image is the cover image.
[0091] It can be seen that through this embodiment, it is possible to determine the combination method of the cover text and the cover background image according to the distribution scenario of the cover image of the story, combine the cover text and the cover background image according to the combination method to obtain the cover image, so that the cover image matches the distribution scenario, and improve the experience of viewing the cover image in different distribution scenarios.
[0092] In one embodiment, generating a cover image of a story based on the cover text and the cover background image includes: Determine the second display position of the cover text in the cover background image according to the first display position of the main object in the cover background image, and determine the display style of the cover text; Add the cover text to the cover background image based on the second display position and the display style, and the cover background image after adding is the cover image.
[0093] In this embodiment, first, determine the second display position of the cover text in the cover background image according to the first display position of the main object in the cover background image. For example, first, identify the coordinates of the minimum bounding rectangle of the face area in the cover background image through a specific algorithm, and determine whether the face area is in the upper half or the lower half of the cover background image according to these coordinates. The first display position includes the upper half or the lower half of the cover background image. Then, determine the second display position among other positions outside the face area in the cover background image, and the second display position does not coincide with the first display position. The second display position includes the upper half or the lower half of the cover background image. The second display position and the first display position can be opposite regions.
[0094] The above specific algorithm can be the OpenPose algorithm. The OpenPose algorithm is a human pose estimation algorithm that can detect the key points of the human body, such as the head, shoulders, elbows, hands, etc., from images or videos, and estimate the positions and directions of these key points.
[0095] Then, determine the display style of the cover text. The display style of the cover text can be determined according to at least one of the story category dimension, the story emotional atmosphere dimension in the above story dimensions, and the preset category labels of the story. For example, when the story category in the story category dimension is ancient romance, the display style of the cover text is determined to be official script in black; when the story category in the story category dimension is modern workplace, the display style of the cover text is determined to be cursive script in black.
[0096] Finally, at the second display position of the cover background image, add the cover text to the cover background image according to the display style of the cover text, and use the cover background image with the added cover text as the cover image. When the number of characters included in the cover text is large, the cover text can be line-wrapped. The line-wrapping requirements include: the cover text is displayed in at most two lines, with at most 9 Chinese characters in one line. When line-wrapping, the number of characters in the first line is less than that in the second line. When line-wrapping and splitting, the Jieba algorithm is used for Chinese word segmentation, and try to avoid splitting a word into two lines. Delete the punctuation marks at the end of the first line and the beginning of the second line after splitting.
[0097] Among them, the Jieba algorithm has a corresponding Chinese word segmentation library. Using this library, Chinese text can be segmented into individual words. The Jieba library uses a variety of word segmentation algorithms, can accurately segment Chinese text, and provides a variety of word segmentation modes and part-of-speech tagging functions.
[0098] In an example, the process of determining the second display position of the cover text in the cover background image according to the first display position of the main object in the cover background image, and determining the display style of the cover text, and adding the cover text to the cover background image based on the second display position and the display style, and using the added cover background image as the cover image can be processed by an AI drawing model.
[0099] It can be seen that through this embodiment, it is possible to determine the second display position of the cover text in the cover background image according to the first display position of the main object in the cover background image, avoid the cover text from covering the important parts of the main object, and determine the display style of the cover text, and add the cover text to the cover background image based on the second display position and the display style, and use the added cover background image as the cover image, so as to achieve the effect of accurately generating the cover image based on the cover text and the cover background image.
[0100] Figure 3Schematic diagram of the cover image of the story provided by an embodiment of the present disclosure, as Figure 3 shown, the cover background image of the story can be generated first, and cover text such as the story title "The Road of Top Students" is added to the cover background image to obtain the cover image of the story. Figure 3 The cover image in can be generated by an AI drawing model based on the previous example.
[0101] In one embodiment, based on the cover text and the cover background image, generating the cover image of the story includes: Determining the cover template corresponding to the story according to the story category of the story, and determining the display style of the cover text; the third display position of the cover background image and the fourth display position of the cover text are marked in the cover template; Based on the third display position, the fourth display position and the display style, adding the cover text and the cover background image to the cover template, and the added cover template is the cover image.
[0102] In this embodiment, first determine the story category of the story. The story category of the story is determined based on at least one of the story category dimension, the story emotional atmosphere dimension in the above story dimensions and the preset category label of the story. The story category of the story can be exemplified as "Ancient - Palace - Palace Infighting - Big Female Lead - Promotion - Positive Ending". Then, obtain the mapping relationship between the story category of the story and the cover template configured in advance, and determine the cover template corresponding to the story according to this mapping relationship and the story category of the story. The third display position of the cover background image and the fourth display position of the cover text are set in the cover template.
[0103] Next, determine the display style of the cover text according to at least one of the story category dimension, the story emotional atmosphere dimension in the above story dimensions and the preset category label of the story. For example, when the story category in the story category dimension is ancient romance, determine the display style of the cover text as official script in black; when the story category in the story category dimension is modern workplace, determine the display style of the cover text as cursive script in black.
[0104] Then, add the cover background image at the third display position in the cover template, and add the cover text at the fourth display position in the cover template according to the display style of the cover text. Determine the cover template after adding the cover background image and the cover text as the cover image.
[0105] It can be seen that through this embodiment, the cover template corresponding to the story can be determined according to the story category of the story, and the cover image of the story can be accurately generated by using the cover template, improving the matching degree between the cover image and the story.
[0106] In one example, the above-introduced method of determining the cover template corresponding to a story based on the story category of the story, and determining the display style of the cover text, and adding the cover text and the cover background image to the cover template based on the third display position, the fourth display position, and the display style, and the resulting cover template is the cover image. This process can be obtained by processing through an AI drawing model.
[0107] Figure 4 The following is a schematic diagram of the cover image of a story provided by another embodiment of the present disclosure, as Figure 4 shown. First, the cover template of the story can be determined, and the cover background image and the cover text are added to the cover template. The cover text includes the story title "The Road of a Top Student" and the content summary "This story tells an inspiring story of a security guard's transformation into a graduate student at University A". The cover image of the story is obtained. Figure 4 The cover image in [] can be generated through an AI drawing model based on the previous example.
[0108] In one embodiment, Figure 1 the method flow in [] further includes: Generating a recommendation card for the story based on the cover image, and recommending the story through the recommendation card.
[0109] In this embodiment, after obtaining the cover image of the story, a recommendation card for the story can be generated based on the cover image of the story, and the story is recommended through the recommendation card.
[0110] Figure 5 The following is a schematic diagram of the recommendation card of a story provided by an embodiment of the present disclosure, as Figure 5 shown. Continuing with Figure 3 the example in [], the story lead such as "My name is Zhao Yanyan. I was originally just an ordinary female security guard. An accident..." can be obtained. Based on the cover image of the story and the story lead of the story, a recommendation card for the story is generated. The recommendation card for the story includes the cover image of the story and the story lead of the story. The recommendation cards for the story are displayed in a two-column form on the story recommendation page to achieve the effect of recommending the story. Two-column means the way of displaying the recommendation cards for the story in two columns.
[0111] It can be seen that through this embodiment, it is also possible to generate a recommendation card for the story based on the cover image of the story, recommend the story through the recommendation card, and use the cover image of the story to improve the recommendation efficiency of the recommended story.
[0112] In summary, through this embodiment, it is possible to automatically generate a matching cover image for the story based on the story text of the story, improve the generation efficiency of the cover image of the story, and improve the recommendation efficiency of the recommended story based on the cover image of the story.
[0113] Figure 6The structural schematic diagram of a cover image generation device based on a story text provided by an embodiment of the present disclosure is as follows Figure 6 As shown, the device includes: An information generation unit 61, configured to obtain the story text of the story, and determine the story framework of the story according to the story text; wherein, the story framework is used to constrain the element content of the generation dimension of the cover background image of the story; A background image generation unit 62, configured to generate description information of the cover background image of the story through a generative model, and generate the cover background image of the story according to the description information of the cover background image; A cover generation unit 63, configured to obtain the cover text of the story, and generate the cover image of the story based on the cover text and the cover background image; the cover image is used to recommend the story.
[0114] Optionally, the information generation unit 61 is specifically configured to: determine each story dimension for representing the story framework, and determine the element content of the story dimension according to the story text; the story dimension includes a main object dimension, an era background dimension, a regional background dimension, a story category dimension, and a story emotional atmosphere dimension; combine the element content of each story dimension to obtain the story framework of the story.
[0115] Optionally, the information generation unit 61 is specifically configured to: determine the target character of the story according to the story text; the cover image includes the target character; generate the story framework according to the content used to describe the target character in the story text; the content is used to describe the appearance information and clothing information of the target character.
[0116] Optionally, the information generation unit 61 is further specifically configured to: if the number of target characters is one, generate the story framework according to the appearance information and clothing information of the target character described in the story text; if the number of target characters is multiple, generate the story framework according to the first content used to describe each target character itself and the second content used to describe the relationship between the target characters in the story text; the first content includes the appearance information and clothing information of the target character; the second content includes the pose information between the target characters.
[0117] Optionally, the information generation unit 61 is specifically configured to: determine the story category of the story according to the story text, and determine the image style of the cover background image of the story according to the story category; generate guiding information for the image style, and generate the story framework according to the guiding information; the guiding information is used to guide the generative model to generate style information of the image style as part of the description information of the cover background image.
[0118] Optionally, the background image generation unit 62 is specifically configured to: obtain, through the generative model, the information content requirements corresponding to the cover background image description information, and obtain the example information corresponding to the cover background image description information; the information content requirements are used to indicate the contents required to be described by the cover background image description information and the description requirements for each content; through the generative model, perform content expansion based on the story framework to obtain the cover background image description information that meets the information content requirements and matches the example information.
[0119] Optionally, the background image generation unit 62 is further specifically configured to: determine, through the generative model, each sub-requirement included in the information content requirements, and determine the information fragment corresponding to the sub-requirement in the example information; the sub-requirement is used to indicate one content and the description requirements for the content; through the generative model, perform content expansion based on the story framework to obtain sub-description information that meets the sub-requirement and matches the information fragment corresponding to the sub-requirement; through the generative model, obtain the cover background image description information according to each sub-description information.
[0120] Optionally, the background image generation unit 62 is further specifically configured to: obtain, through the generative model, the information reasoning order represented by the thinking chain strategy for generating the cover background image description information; the information reasoning order includes the order of reasoning for each sub-requirement; through the generative model, in accordance with the information reasoning order, perform content expansion based on the story framework, and sequentially generate sub-description information that meets each sub-requirement and matches the information fragment corresponding to the sub-requirement.
[0121] Optionally, the background image generation unit 62 is further specifically configured to: determine, through the generative model, in accordance with the information reasoning order, the sub-requirement to be reasoned, and determine the target content corresponding to the sub-requirement in the story framework; through the generative model, perform content expansion based on the target content to generate sub-description information that meets the sub-requirement and matches the information fragment corresponding to the sub-requirement.
[0122] Optionally, the cover generation unit 63 is specifically configured to: determine the combination method of the cover text and the cover background image according to the distribution scenario of the cover image of the story; combine the cover text and the cover background image according to the combination method to obtain the cover image.
[0123] Optionally, the cover generation unit 63 is specifically configured to: determine a second display position of the cover text in the cover background image according to a first display position of a main object in the cover background image, and determine a display style of the cover text; and add the cover text to the cover background image based on the second display position and the display style, and the cover background image after the addition is the cover image.
[0124] Optionally, the cover generation unit 63 is specifically configured to: determine a cover template corresponding to the story according to the story category of the story, and determine a display style of the cover text; the third display position of the cover background image and the fourth display position of the cover text are marked in the cover template; and add the cover text and the cover background image to the cover template based on the third display position, the fourth display position, and the display style, and the cover template after the addition is the cover image.
[0125] Optionally, it further includes a story recommendation unit, configured to: generate a recommendation card for the story according to the cover image, and recommend the story through the recommendation card.
[0126] In this embodiment, first, obtain the story text of the story, determine the story framework of the story according to the story text, where the story framework is used to constrain the element content of the generation dimension of the cover background image of the story, then, through a generative model, generate the description information of the cover background image of the story according to the story framework, generate the cover background image of the story according to the description information of the cover background image, and finally, obtain the cover text of the story, and generate the cover image of the story based on the cover text and the cover background image, where the cover image is used to recommend the story. It can be seen that through this embodiment, a matching cover image can be automatically generated for the story according to the story text of the story, the generation efficiency of the cover image of the story is improved, and the recommendation efficiency of recommending the story based on the cover image of the story is improved.
[0127] The cover image generation device in the embodiments of the present disclosure can implement each process of the above cover image generation method embodiment, and achieve the same effects and functions, which will not be repeated here.
[0128] An embodiment of the present disclosure further provides an electronic device Figure 7 is a schematic structural diagram of the electronic device provided in an embodiment of the present disclosure, as Figure 7As shown, electronic devices can vary significantly due to different configurations or performances. They can include one or more processors 701 and a memory 702. One or more applications or data can be stored in the memory 702. Among them, the memory 702 can be for transient storage or persistent storage. The applications stored in the memory 702 can include one or more modules (not shown in the figure), and each module can include a series of computer-executable instructions in the electronic device. Further, the processor 701 can be set to communicate with the memory 702 and execute a series of computer-executable instructions in the memory 702 on the electronic device. The electronic device can also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input or output interfaces 705, one or more keyboards 706, etc.
[0129] In a specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, and when the computer-executable instructions are executed, the processor implements the following processes: Obtain the story text of the story, and determine the story framework of the story according to the story text; wherein, the story framework is used to constrain the element content of the generation dimension of the cover background image of the story; Through a generative model, generate the cover background image description information of the story according to the story framework, and generate the cover background image of the story according to the cover background image description information; Obtain the cover text of the story, and generate the cover image of the story based on the cover text and the cover background image; the cover image is used to recommend the story.
[0130] In this embodiment, first, obtain the story text of the story, determine the story framework of the story according to the story text, and the story framework is used to constrain the element content of the generation dimension of the cover background image of the story. Then, through a generative model, generate the cover background image description information of the story according to the story framework, and generate the cover background image of the story according to the cover background image description information. Finally, obtain the cover text of the story, and generate the cover image of the story based on the cover text and the cover background image. The cover image is used to recommend the story. It can be seen that through this embodiment, it is possible to automatically generate a matching cover image for the story according to the story text of the story, improve the generation efficiency of the cover image of the story, and improve the recommendation efficiency of recommending the story based on the cover image of the story.
[0131] The electronic device in the embodiments of the present disclosure can implement each process of the above-mentioned cover image generation method embodiment and achieve the same effects and functions, which will not be repeated here.
[0132] Another embodiment of the present disclosure also provides a computer-readable storage medium for storing computer-executable instructions, and when the computer-executable instructions are executed by a processor, the following processes are implemented: Obtain the story text of the story, and determine the story framework of the story according to the story text; wherein, the story framework is used to constrain the element content of the generation dimension of the cover background image of the story; Through a generative model, generate the cover background image description information of the story according to the story framework, and generate the cover background image of the story according to the cover background image description information; Obtain the cover text of the story, and generate the cover image of the story based on the cover text and the cover background image; the cover image is used to recommend the story.
[0133] In this embodiment, first, obtain the story text of the story, determine the story framework of the story according to the story text, and the story framework is used to constrain the element content of the generation dimension of the cover background image of the story. Then, through a generative model, generate the cover background image description information of the story according to the story framework, and generate the cover background image of the story according to the cover background image description information. Finally, obtain the cover text of the story, and generate the cover image of the story based on the cover text and the cover background image. The cover image is used to recommend the story. It can be seen that through this embodiment, a matching cover image can be automatically generated for the story according to the story text of the story, the generation efficiency of the cover image of the story is improved, and the recommendation efficiency of recommending the story based on the cover image of the story is improved.
[0134] The computer-readable storage medium in the embodiments of the present disclosure can implement each process of the above-mentioned cover image generation method embodiment, and achieve the same effects and functions, which will not be repeated here.
[0135] Another embodiment of the present disclosure also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the following processes are implemented: Obtain the story text of the story, and determine the story framework of the story according to the story text; wherein, the story framework is used to constrain the element content of the generation dimension of the cover background image of the story; Through a generative model, generate the cover background image description information of the story according to the story framework, and generate the cover background image of the story according to the cover background image description information; Obtain the cover text of the story, and generate the cover image of the story based on the cover text and the cover background image; the cover image is used to recommend the story.
[0136] In this embodiment, first, the story text of the story is obtained, and the story framework of the story is determined according to the story text. The story framework is used to constrain the element content of the generation dimension of the cover bottom image of the story. Then, through the generative model, the cover bottom image description information of the story is generated according to the story framework. According to the cover bottom image description information, the cover bottom image of the story is generated. Finally, the cover text of the story is obtained, and based on the cover text and the cover bottom image, the cover image of the story is generated. The cover image is used to recommend the story. It can be seen that through this embodiment, a matching cover image can be automatically generated for the story according to the story text of the story, the generation efficiency of the cover image of the story is improved, and the recommendation efficiency of recommending the story based on the cover image of the story is improved.
[0137] The computer program product in the embodiments of the present disclosure can implement each process of the above cover image generation method embodiment, and achieve the same effects and functions, which will not be repeated here.
[0138] In each embodiment of the present disclosure, the computer-readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.
[0139] In the 1990s, it was obvious to distinguish whether an improvement in a technology was a hardware improvement (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement in method flows). However, with the development of technology, many improvements in method flows today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by a user's programming of the device. Designers can program by themselves to "integrate" a digital system on a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow with the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0140] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0141] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0142] For the convenience of description, the above devices are described by dividing them into various units according to their functions. Of course, when implementing the embodiments of the present disclosure, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0143] Those skilled in the art should understand that one or more embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0144] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0145] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0147] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the said element.
[0148] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.
[0149] Each embodiment in the present disclosure is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content.
[0150] The above are only the embodiments of the present disclosure and are not intended to limit the present disclosure. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the scope of the claims of the present disclosure.
Claims
1. A method for generating a cover image based on a story text, characterized in that, Including: Obtain the story text of the story, and determine the story framework of the story according to the story text; wherein, the story framework is used to constrain the element content of the generation dimension of the cover background image of the story; Through a generative model, generate the description information of the cover background image of the story according to the story framework, and generate the cover background image of the story according to the description information of the cover background image; Obtain the cover text of the story, and generate the cover image of the story based on the cover text and the cover background image; the cover image is used to recommend the story.
2. The method according to claim 1, characterized in that The determining the story framework of the story according to the story text includes: Determine each story dimension used to represent the story framework, and determine the element content of the story dimension according to the story text; the story dimensions include the main object dimension, the era background dimension, the regional background dimension, the story category dimension, and the story emotional atmosphere dimension; Combine the element content of each story dimension to obtain the story framework of the story.
3. The method according to claim 1, wherein The determining the story framework of the story according to the story text includes: Determine the target character of the story according to the story text; the cover image includes the target character; Generate the story framework according to the content used to describe the target character in the story text; the content is used to describe the appearance information and clothing information of the target character.
4. The method according to claim 3, wherein The generating the story framework according to the content used to describe the target character in the story text includes: If the number of target characters is one, generate the story framework according to the appearance information and clothing information of the target character described in the story text; If the number of target characters is multiple, generate the story framework according to the first content used to describe each target character itself and the second content used to describe the target characters; the first content includes the appearance information and clothing information of the target character; the second content includes the pose information between the target characters.
5. The method according to claim 1, characterized in that The determining the story framework of the story according to the story text includes: Determine the story category of the story according to the story text, and determine the image style of the cover background image of the story according to the story category; Generate guiding information for the image style, and generate the story framework according to the guiding information; the guiding information is used to guide the generative model to generate the style information of the image style as part of the description information of the cover background image.
6. The method according to claim 1, characterized in that, The generating the description information of the cover background image of the story through the generative model according to the story framework includes: Through the generative model, obtain the information content requirements corresponding to the description information of the cover background image, and obtain the example information corresponding to the description information of the cover background image; the information content requirements are used to indicate the various contents that the description information of the cover background image needs to describe and the description requirements for the various contents; Based on the story framework, content expansion is performed through the generative model to obtain the description information of the cover background image that meets the information content requirements and matches the example information.
7. The method according to claim 6, wherein The method of performing content expansion based on the story framework through the generative model to obtain the description information of the cover background image that meets the information content requirements and matches the example information includes: Through the generative model, determine each sub-requirement included in the information content requirements, and determine the information fragment corresponding to the sub-requirement in the example information; the sub-requirement is used to indicate a piece of content and the description requirements for the content. Through the generative model, perform content expansion based on the story framework to obtain sub-description information that meets the sub-requirement and matches the information fragment corresponding to the sub-requirement. Through the generative model, obtain the description information of the cover background image according to each sub-description information.
8. The method according to claim 7, wherein The method of performing content expansion based on the story framework through the generative model to obtain sub-description information that meets the sub-requirement and matches the information fragment corresponding to the sub-requirement includes: Through the generative model, obtain the information reasoning order represented by the chain of thought strategy for generating the description information of the cover background image; the information reasoning order includes the order of reasoning for each sub-requirement. Through the generative model, according to the information reasoning order, perform content expansion based on the story framework, and sequentially generate sub-description information that meets each sub-requirement and matches the information fragment corresponding to the sub-requirement.
9. The method according to claim 8, wherein The method of performing content expansion based on the story framework through the generative model according to the information reasoning order to sequentially generate sub-description information that meets each sub-requirement and matches the information fragment corresponding to the sub-requirement includes: Through the generative model, according to the information reasoning order, determine the sub-requirement to be reasoned, and determine the target content corresponding to the sub-requirement in the story framework. Through the generative model, perform content expansion based on the target content to generate sub-description information that meets the sub-requirement and matches the information fragment corresponding to the sub-requirement.
10. The method according to claim 1, wherein Generating the cover image of the story based on the cover text and the cover background image includes: Determine the combination method of the cover text and the cover background image according to the distribution scenario of the cover image of the story. According to the combination method, combine the cover text and the cover background image to obtain the cover image.
11. The method according to claim 1, characterized in that, Generating the cover image of the story based on the cover text and the cover background image includes: According to the first display position of the main object in the cover background image, determine the second display position of the cover text in the cover background image, and determine the display style of the cover text. Based on the second display position and the display style, add the cover text to the cover background image, and the cover background image after addition is the cover image.
12. The method according to claim 1, wherein Generating the cover image of the story based on the cover text and the cover background image, including: Determining the cover template corresponding to the story according to the story category of the story, and determining the display style of the cover text; the third display position of the cover background image and the fourth display position of the cover text are marked in the cover template; Adding the cover text and the cover background image to the cover template based on the third display position, the fourth display position and the display style, and the added cover template is the cover image.
13. The method according to claim 1, wherein Further including: Generating a recommendation card for the story according to the cover image, and recommending the story through the recommendation card.
14. A cover image generation device based on a story text, characterized in that, Including: An information generation unit, configured to obtain the story text of the story, and determine the story framework of the story according to the story text; wherein, the story framework is used to constrain the element content of the generation dimension of the cover background image of the story; A background image generation unit, configured to generate the description information of the cover background image of the story through a generative model, and generate the cover background image of the story according to the description information of the cover background image; A cover generation unit, configured to obtain the cover text of the story, and generate the cover image of the story based on the cover text and the cover background image; the cover image is used to recommend the story.
15. An electronic device, characterized in that, Including: A processor; And A memory configured to store computer-executable instructions, and the computer-executable instructions, when executed, cause the processor to implement the method according to any one of claims 1-13 above.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1-13 above.
17. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program, when executed by a processor, implements the method according to any one of claims 1-13 above.
Citation Information
Cited By
Image generation method, image generation device, and computer readable storage medium
CN121600117A