A method, device, and storage medium for generating a cartoon image
By using patented language and text-based image generation technology, the problem of inconsistency between characters and text descriptions in comic creation has been solved, achieving consistency between characters and text descriptions and improving visual effects.
Patent Information
- Application Number
- CN202411637339.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-11-15
AI Technical Summary
Existing technologies suffer from inconsistencies between characters and text descriptions when generating comic book images featuring multiple characters.
By using a pre-defined language model and text-based image model, character information and storyboard cues from the novel text are obtained, the distribution and posture of the characters are determined, and storyboard images are generated.
Ensuring that the characters in the generated comic images match the text descriptions improves the efficiency of comic creation and the visual presentation.
Smart Images

Figure CN119810223B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to a comic image generation method and device, equipment and a storage medium. BACKGROUND
[0002] The traditional comic creation process generally includes a large number of and complex steps, such as scriptwriting, character design, shot script, draft drawing, and coloring. Therefore, the production of a comic generally requires a large number of workers to work together. In recent years, with the breakthrough of artificial intelligence technology in the field of natural language processing and computer vision, large language models and image generation technology have made continuous progress, making it possible to convert novel content into comic images, greatly improving the efficiency of comic creation and the visual presentation effect of novel content.
[0003] Large language models perform well in understanding and generating natural language texts. They can analyze complex plot content, character dialogues, and narrative styles in novels and convert these information into shot descriptions suitable for comics and prompt words for text-to-image models, without the need for significant human intervention, reducing the dependence on manual drawing in traditional comic creation.
[0004] However, the related art still faces the problem of inconsistency between the characters in the image and the text description when generating images including multiple characters. SUMMARY
[0005] The main purpose of the present application is to provide a comic image generation method, device, equipment and storage medium, which generates a shot image consistent with the text description of the characters through a preset language model and a text-to-image model.
[0006] In a first aspect, the present application provides a comic image generation method, comprising:
[0007] Obtaining initial character images corresponding to each character in a novel text;
[0008] Determining a shot text corresponding to at least one shot in the novel text based on a first language submodel of a preset language model;
[0009] Determining a shot prompt word, a character prompt word, and a character distribution position of each character corresponding to the shot based on a second language submodel of the language model according to the shot text;
[0010] Determining the posture of the characters in the shot based on a first image submodel of a preset text-to-image model according to the initial character images, the character prompt word, and the character distribution position corresponding to the shot of each character;
[0011] determine, based on the initial role image corresponding to each of the roles in the screenplay, the posture of the role, and the screenplay prompt word, a screenplay image corresponding to the screenplay based on a second image submodel of the text-to-image model.
[0012] In a second aspect, the present application further provides a comic image generation device, comprising:
[0013] an acquisition module configured to acquire an initial role image corresponding to each role in a novel text;
[0014] a text determination module configured to determine, based on a first language submodel of a preset language model, a screenplay text corresponding to at least one screenplay in the novel text;
[0015] a position determination module configured to determine, based on a second language submodel of the language model, a screenplay prompt word, a role prompt word, and a role distribution position of each of the roles corresponding to the screenplay based on the screenplay text;
[0016] a posture determination module configured to determine, based on a first image submodel of a preset text-to-image model, the posture of the role in the screenplay based on the initial role image, the role prompt word corresponding to the screenplay, and the role distribution position;
[0017] an image determination module configured to determine, based on a second image submodel of the text-to-image model, a screenplay image corresponding to the screenplay based on the initial role image corresponding to each of the roles in the screenplay, the posture of the role, and the screenplay prompt word.
[0018] In a third aspect, the present application further provides a computer device, comprising a memory and a processor;
[0019] the memory is configured to store a computer program;
[0020] the processor is configured to execute the computer program and implement the comic image generation method as described above when the computer program is executed.
[0021] In a fourth aspect, the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the comic image generation method as described above.
[0022] This application provides a method, apparatus, device, and storage medium for generating comic images. The method includes: acquiring initial character images corresponding to each character in a novel text; determining storyboard text corresponding to at least one storyboard in the novel text based on a first language sub-model of a preset language model; determining storyboard prompts, character prompts, and the character distribution positions of each character in the storyboard text based on a second language sub-model of the language model; determining the posture of the character in the storyboard based on the initial character images, the character prompts corresponding to the storyboard, and the character distribution positions based on a first image sub-model of a preset text-to-image model; and determining the storyboard image corresponding to the storyboard based on the second image sub-model of the text-to-image model, the character posture, and the storyboard prompts. This application first obtains the initial character images corresponding to each character in the novel text; then, based on a preset language model, it determines the storyboard prompts, character prompts, and the character distribution positions of each character in each storyboard; then, based on a preset text-to-image model, it determines the posture of each character in the storyboard, and generates storyboard images including multiple characters and corresponding prompts according to the initial character images, character postures, and storyboard prompts, thereby ensuring that the characters in each storyboard image are consistent with the text description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart illustrating a method for generating comic images provided in an embodiment of this application;
[0025] Figure 2 This is a schematic diagram illustrating the connection between the server and the terminal device provided in an embodiment of this application;
[0026] Figure 3 This is a partial schematic diagram of the framework of the text image model provided in the embodiments of this application;
[0027] Figure 4 A schematic block diagram of a comic image generation apparatus provided in this application embodiment;
[0028] Figure 5 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0029] With reference to the drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts are within the scope of the present application.
[0030] The flowcharts shown in the drawings are only illustrative, and do not necessarily include all contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be further decomposed, combined or partially merged, so that the actual execution order can be changed according to actual conditions.
[0031] The embodiments of the present application provide a cartoon image generation method, device, equipment and storage medium. The cartoon image generation method can be applied to a terminal device, which can be a mobile phone, a tablet computer, a notebook computer, a desktop computer or the like. It can also be applied to a server, which can be a separate server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0032] Some embodiments of the present application will be described in detail below with reference to the drawings. In the case of no conflict, the embodiments described below and the features in the embodiments can be combined with each other.
[0033] Please refer to Figure 1 , Figure 1 A flowchart of a cartoon image generation method provided by an embodiment of the present application is shown. It should be noted that the cartoon image generation method provided by the embodiment of the present application can be used in a terminal device, and of course can also be used in a server.
[0034] As Figure 2 shown, the cartoon image generation method is applied to a server, which is in communication connection with a terminal device. The server can send the shot image obtained by the cartoon image generation method to the terminal device. Of course, it is not limited to this, and is not limited herein.
[0035] In specific implementation, the terminal device includes but is not limited to any one of a mobile phone, a tablet computer, a notebook computer, and a desktop computer; the server can be a separate server or a server cluster, and can also be a cloud server providing cloud computing services.
[0036] As Figure 1 shown, the cartoon image generation method includes steps S101 to S105.
[0037] In step S101, an initial role image corresponding to each role in the novel text is obtained.
[0038] The novel text generally includes multiple roles, and each role is closely related to the plot of the novel text. In combination with the description of different roles in the novel text, the initial role image corresponding to each role is obtained, so as to shape a more stereoscopic role image.
[0039] For example, the initial role image in the embodiment of the present application can be generated according to the role information of each role. For example, the role information of the role includes basic information (name, age, personality), personality characteristics (positive characteristics, negative characteristics), background story (family background, growing experience, social relationship), ability and skill (magic, superpower, professional skill), role positioning and relationship (main character, supporting character, relationship with other characters), appearance and dressing (facial features, facial contour, clothing), etc.
[0040] In other embodiments, the embodiment of the present application can also obtain the initial role image corresponding to each role imported by the user.
[0041] In step S102, based on a first language sub-model of a preset language model, a shot text corresponding to at least one shot in the novel text is determined.
[0042] It should be noted that the shot is an important link in the creation process of visual media such as movies, animations, advertisements and comics, mainly refers to the continuous story or plot is divided into a series of pictures or shots that can be shot or drawn. The shot text, also known as the storyboard text or shot script, is a very important part of the production of visual media such as movies, animations, advertisements and comics, which describes the content, action, dialogue, sound effects, picture composition and other information of the shot in detail.
[0043] The preset language model in the embodiments of the present application can include a first language sub-model for splitting novel text. Based on the first language sub-model, the novel text can be split into multiple shot texts corresponding to respective shots. Specifically, the semantic structure in the novel text can be analyzed based on the first language sub-model to find semantic breakpoints and split. For example, the first language sub-model can include a BERT pre-training language model, a GPT model, etc. It should be noted that the BERT pre-training language model is based on a Transformer architecture and is trained based on a large-scale text corpus to capture deep language features. The GPT model is based on a Transformer architecture and learns statistical features and internal structures of language through unsupervised learning on a large-scale text corpus. This pre-training method enables the GPT model to generate coherent and grammatically correct text, suitable for various text generation tasks. The Transformer architecture is a model composed of multiple self-attention mechanisms and positional feedforward neural networks, which can capture long-range dependencies in sequences and has the advantage of parallel computing.
[0044] In step S103, based on a second language sub-model of the language model, the shot corresponding shot prompt, the role prompt, and the role distribution position of each role are determined according to the shot text.
[0045] For example, the shot corresponding shot prompt refers to a keyword for guiding the specific content and performance of each shot. The role prompt can include the content of the role speaking, keywords for indicating actions, etc. These prompts can help the working members (such as directors, photographers, animators, etc.) quickly understand the visual style, emotional atmosphere, action requirements, etc. of each shot. The shot prompt can include picture composition, light color; the role prompt can include the action expression of each role, role dialogue, voiceover, etc.
[0046] The role distribution position in the embodiments of the present application refers to the position of one or more roles in a shot. For example, a shot can include a role A, which is in the middle, upper left, or lower right of the shot; for another example, a shot can include three roles, namely role A, role B, and role C, role A and role B are on the right side of the shot, and role C is on the left side of the shot.
[0047] Specifically, the embodiment of the present application can extract key information in the screenplay text based on the second language sub-model of the language model to obtain screenplay corresponding screenplay prompt words, role prompt words and role distribution positions of each role. For example, the screenplay text is "on the evening of October 20, 2020, at 20:05, in XX convenience store, role A and role B stand face to face, role A holds a bottle of beverage and points to role B and says, "Why sell expired products", role B has his hands up and extends to role A and says, "Our store will not sell expired products, please provide me with the receipt of the bottle of beverage you purchased". The embodiment of the present application can input the screenplay text into the BERT pre-training language model to perform semantic analysis on the screenplay text through the BERT pre-training language model and determine the screenplay prompt words (for example, on the evening of October 20, 2020, at 20:05, in XX convenience store) of the convenience store scene, the lines and actions of role A, the lines and actions of role B and the distribution positions of role A and role B. The second language sub-model of the language model in the embodiment of the present application can extract prompt words corresponding to different roles in the screenplay scene respectively to ensure that the roles and the description text are consistent.
[0048] Exemplarily, the second language sub-model in the embodiment of the present application can also include a BERT pre-training language model and a GPT model.
[0049] In step S104, based on the first image sub-model of the preset text-to-image model, the posture of the role in the screenplay is determined according to the initial role image, the role prompt words corresponding to the screenplay and the role distribution positions.
[0050] The posture of the role generally refers to the body posture, action and expression of the role in the screenplay. These postures can convey the emotions, personality and plot of the role. For example, a brave role may present a posture of standing up straight, and a self-abased role may present a posture of lowering his head and shoulders.
[0051] It can be understood that the text-to-image model in the embodiment of the present application can convert abstract text into specific visual performance by learning the relationship between a large number of images and text descriptions. The text-to-image model is generally composed of several parts, including a text understanding component, an image information creation and an image decoder. The text understanding component refers to converting input text information into digital representation to capture ideas in the text; the image information creation generates image information based on these digital representations and noise data; finally, the image decoder converts the image information into visual images.
[0052] The first image submodel of the text-to-image model in the embodiments of this application can include a diffusion model, an autoregressive model, a generative adversarial network model, etc. It should be noted that the diffusion model generates a new image by gradually adding noise to the image and then learning how to recover the original image from this noise distribution. Among them, the typical diffusion model includes the Stable Diffusion model.
[0053] The embodiments of this application can first process the initial role image and the shot corresponding role prompt word based on the first image submodel of the preset text-to-image model to obtain the role corresponding image, and then determine the pose of the role in the shot according to the role corresponding image obtained by processing and the role distribution position.
[0054] Step S105, based on the second image submodel of the text-to-image model, determining the shot corresponding shot image according to the initial role image corresponding to each role in the shot, the pose of the role and the shot prompt word.
[0055] The shot image refers to a shot design draft, which is mainly used to show the author's arrangement of the plot and the picture, including the proportion of separation, shot editing, pose of role performance and placement of lines, etc. In the process of making a comic, the shot image of a comic is generally a continuous picture displayed in a grid, and the story is told through these pictures. It should be noted that the grid is a form unique to comics, which generally refers to cutting a plane within a specified size, and reasonably arranging the number and size ratio of the grid according to the content. Common grid methods include broken format, regular format, diagonal format, full-page format, etc.
[0056] For example, the second image submodel of the text-to-image model in the embodiments of this application can include the Stable Diffusion model, which is a model based on denoising diffusion probability and can generate images according to text prompts. Based on the second image submodel of the text-to-image model, the shot image including at least one role, the pose of the role and the prompt word can be generated according to the initial role image corresponding to each role in the shot, the pose of the role and the shot prompt word.
[0057] The method for generating a comic image provided by the above embodiment comprises: obtaining initial role images corresponding to respective roles in a novel text; determining, based on a first language submodel of a preset language model, a shot text corresponding to at least one shot in the novel text; determining, based on a second language submodel of the language model, a shot prompt word, a role prompt word and a role distribution position of each role corresponding to the shot according to the shot text; determining, based on a first image submodel of a preset text-to-image model, a posture of each role in the shot according to the initial role images, the role prompt word corresponding to the shot and the role distribution position; and determining, based on a second image submodel of the text-to-image model, a shot image corresponding to the shot according to the initial role images corresponding to respective roles in the shot, the posture of each role and the shot prompt word. The embodiment of the application first obtains initial role images corresponding to respective roles in a novel text; then determines a shot prompt word, a role prompt word and a role distribution position of each role corresponding to each shot based on a preset language model; and then determines a posture of each role in the shot based on a preset text-to-image model, and generates a shot image including multiple roles and corresponding prompt words according to the initial role images corresponding to respective roles, the posture of each role and the shot prompt word, thereby ensuring that the roles in each shot image are consistent with the text description.
[0058] In an exemplary embodiment, step S101 can include step S1011 and step S1012.
[0059] Step S1011 comprises obtaining role information corresponding to respective roles in the novel text based on a third language submodel of the language model.
[0060] Step S1012 comprises obtaining initial role images corresponding to respective roles based on a third image submodel of the text-to-image model according to the role information corresponding to respective roles.
[0061] The preset initial role image in the embodiment of the application refers to a role setting image, that is, an image designed for a fictional character in the production process of an animation, a game, a movie or the like. The role setting image can represent the appearance characteristics of the hairstyle, clothing, skin color, facial features and the like of the character.
[0062] The third language submodel of the language model in the embodiment of the application can also include a BERT pre-training language model and a GPT model, and the third image submodel of the text-to-image model can also include a Stable Diffusion model.
[0063] For example, the embodiment of the present application can first extract the role information (i.e., role information of hair style, clothing, skin color, etc.) related to describing different roles in the novel text based on the third language sub-model of the language model, and input the role information related to describing the role into the Stable Diffusion model, fully understand the extracted role information through the Stable Diffusion model, and generate the initial role image corresponding to each role, i.e., the role setting image.
[0064] In an exemplary embodiment, step S104 can include step S401 and step S402.
[0065] Step S401, based on the first image sub-model of the preset text-to-image model, obtains the target role image corresponding to the role according to the preset initial role image and the role prompt word corresponding to the shot.
[0066] Step S402, determines the posture of the role in the shot according to the target role image and the role distribution position.
[0067] After obtaining the initial role image, the embodiment of the present application can also generate the target role image (i.e., the role image of each shot) consistent with the text description of the role based on the first image sub-model of the text-to-image model according to the initial role image and the extracted role prompt word corresponding to the shot, so as to accurately present the image characteristics of the role in the corresponding shot.
[0068] After obtaining the corresponding target role image (role image), the embodiment of the present application can determine the posture of the role in the shot according to the target role image and the aforementioned obtained role distribution position, which not only can more accurately present the postures of different roles in the shot, but also can make the roles displayed in the shot more vivid and lifelike.
[0069] In an exemplary embodiment, step S402 can include step S4021 and step S4022.
[0070] Step S4021, performs biological body detection and key point position detection on the target role image to obtain the key point coordinates corresponding to the target role image.
[0071] Step S4022, maps the key point coordinates to the role distribution position to obtain the posture of the role in the shot.
[0072] For example, embodiments of this application can use deep learning models such as deep neural networks, convolutional neural networks, and recurrent neural networks to perform biological detection on target character images, that is, to perform human body detection on character images to accurately identify human body regions in the image and obtain the bounding box of the human body; then, features within the bounding box of the human body are extracted using pose estimation methods to obtain the key point coordinates of the human body; then, the key point coordinates are mapped to the distribution positions of the characters. Specifically, the key point coordinates are matched with the distribution positions of the characters to obtain matching results, and the key point coordinates are adjusted and optimized based on the matching results, thereby accurately obtaining the poses of each character in the storyboard.
[0073] In one exemplary embodiment, step S105 may include step S501.
[0074] Step S501: Input the initial character images, character poses, and storyboard prompts corresponding to each character in the storyboard into the second image sub-model of the Wensheng image model, so as to fuse the initial character images, character poses, and storyboard prompts in the storyboard through the second image sub-model to obtain the storyboard image.
[0075] In this embodiment, the initial character image, the character's posture, and the storyboard prompts in the storyboard can be fused using the second image sub-model of the textual image model to create a storyboard image that matches the character and text description.
[0076] In an exemplary embodiment, the second image sub-model may include a first adapter network and a second adapter network; step S501 may include steps S5011 to S5013.
[0077] Step S5011: Extract the pose information corresponding to the pose of the character through the first adapter network of the second image sub-model.
[0078] Step S5012: Extract the appearance features corresponding to the initial character image in the storyboard through the second adapter network of the second image sub-model.
[0079] Step S5013: Fuse the storyboard prompts, pose information, and appearance features to obtain the storyboard image.
[0080] like Figure 3 As shown, the second image sub-model may include a first adapter network and a second adapter network. The first adapter network can be used to control the pose of the generated image to be consistent with the pose of the character, and the second adapter network can be used to control the appearance features of the character in the generated image to be consistent with the appearance of the character in the initial character image.
[0081] Specifically, the embodiment of the present application can extract pose information corresponding to the pose of the role through the first adapter network of the second image submodel, so as to control the pose of the generated image to be consistent with the pose of the role. The second adapter network of the second image submodel can extract appearance features of the role from the initial role image, so as to control the appearance features of the person in the generated image to be consistent with the appearance of the role in the initial role image. Finally, the pose information, the appearance features and the shot script words are fused together to obtain a shot image that meets the requirements of shot context and role interaction, so that multiple roles can be presented in the same shot image, and the consistency of the appearance, pose and text description of the person is maintained.
[0082] In an exemplary embodiment, step S105 is followed by step S106.
[0083] Step S106, synthesizing the shot images corresponding to the multiple shots into a video.
[0084] After obtaining the shot images corresponding to the multiple shots, the embodiment of the present application can splice and synthesize the multiple shot images into a video in order.
[0085] Please refer to Figure 4 , Figure 4 A schematic block diagram of a comic image generation device provided by the embodiment of the present application. The comic image generation device can be configured in a server or a terminal device, and is used to execute the comic image generation method described above.
[0086] As Figure 4 shown, the comic image generation device includes an acquisition module 110, a text determination module 120, a position determination module 130, a pose determination module 140 and an image determination module 150.
[0087] The acquisition module 110 is configured to acquire initial role images corresponding to respective roles in a novel text.
[0088] The text determination module 120 is configured to determine a shot text corresponding to at least one shot in the novel text based on a first language submodel of a preset language model.
[0089] The position determination module 130 is configured to determine a shot prompt word, a role prompt word and a role distribution position of each role corresponding to the shot based on a second language submodel of the language model according to the shot text.
[0090] The pose determination module 140 is configured to determine a pose of a role in the shot based on a first image submodel of a preset text-to-image model according to the initial role image, the role prompt word and the role distribution position corresponding to the shot.
[0091] The image determination module 150 is configured to determine the screenplay image corresponding to the screenplay based on the second image sub-model of the text-to-image model, according to the initial role image corresponding to each role in the screenplay, the posture of the role, and the screenplay prompt word.
[0092] In an exemplary embodiment, the acquisition module 110 can include a first acquisition sub-module and a second acquisition sub-module.
[0093] The first acquisition sub-module is configured to acquire the role information corresponding to each role in the novel text based on the third language sub-model of the language model.
[0094] The second acquisition sub-module is configured to obtain the initial role image corresponding to each role based on the third image sub-model of the text-to-image model according to the role information corresponding to each role.
[0095] In an exemplary embodiment, the posture determination module 140 can include a target determination sub-module and a posture determination sub-module.
[0096] The target determination sub-module is configured to obtain the target role image corresponding to the role based on the first image sub-model of the preset text-to-image model according to the preset initial role image and the role prompt word corresponding to the screenplay.
[0097] The posture determination sub-module is configured to determine the posture of the role in the screenplay according to the target role image and the role distribution position.
[0098] In an exemplary embodiment, the posture determination sub-module can include a coordinate acquisition sub-module and a mapping sub-module.
[0099] The coordinate acquisition sub-module is configured to perform biological body detection and key point position detection on the target role image to obtain the key point coordinates corresponding to the target role image.
[0100] The mapping sub-module is configured to map the key point coordinates to the role distribution position to obtain the posture of the role in the screenplay.
[0101] In an exemplary embodiment, the image determination module 150 can be specifically configured to input the initial role image corresponding to each role in the screenplay, the posture of the role, and the screenplay prompt word into the second image sub-model of the text-to-image model, so as to perform fusion processing on the initial role image, the posture of the role, and the screenplay prompt word in the screenplay by the second image sub-model to obtain the screenplay image.
[0102] In an exemplary embodiment, the second image sub-model includes a first adapter network and a second adapter network. The image determination module 150 can include a first extraction sub-module, a second extraction sub-module, and a fusion sub-module.
[0103] The first extraction submodule is configured to extract pose information corresponding to the pose of the role through a first adapter network of the second image submodel.
[0104] The second extraction submodule is configured to extract appearance features corresponding to the initial role image in the shot through a second adapter network of the second image submodel.
[0105] The fusion submodule is configured to fuse the shot prompt word, the pose information, and the appearance features to obtain a shot image.
[0106] In an exemplary embodiment, the apparatus can further include a synthesis module.
[0107] The synthesis module is configured to synthesize the shot images corresponding to the multiple shots into a video.
[0108] It should be noted that, for the convenience and brevity of description, the specific working processes of the apparatus and the modules and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein.
[0109] The method of the present application can be used in a variety of general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0110] For example, the method and apparatus described above can be implemented in the form of a computer program that can run on a computer device.
[0111] Please refer to Figure 5 , Figure 5 A structural schematic block diagram of a computer device provided by an embodiment of the present application. The computer device can be a server or a terminal device.
[0112] As Figure 5 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus, wherein the memory can include a storage medium and an internal memory.
[0113] The storage medium may store an operating system and a computer program. The computer program includes program instructions that, when executed, cause the processor to perform the steps of any comic image generation method.
[0114] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0115] Internal memory provides an environment for the execution of computer programs stored in storage media. When these computer programs are executed by a processor, the processor can perform the steps of any comic image generation method.
[0116] This network interface is used for network communication, such as sending assigned tasks.
[0117] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0118] It should be understood that a processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other convertible logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0119] In one embodiment, the processor is configured to execute a computer program and, when executing the computer program, perform the following steps:
[0120] Obtain the initial character images corresponding to each character in the novel text;
[0121] Based on the first language sub-model of the preset language model, determine the storyboard text corresponding to at least one storyboard in the novel text;
[0122] The second language sub-model based on the language model determines the corresponding scene prompts, character prompts, and the character distribution positions of each character based on the scene text;
[0123] determine the pose of the character in the screenplay based on the initial character image, the character prompt word corresponding to the screenplay, and the character distribution position;
[0124] determine the screenplay image corresponding to the screenplay based on the initial character image corresponding to each character in the screenplay, the pose of the character, and the screenplay prompt word.
[0125] In an exemplary embodiment, obtaining the initial character image corresponding to each character in the novel text comprises:
[0126] obtaining the character information corresponding to each character in the novel text based on the third language sub-model of the language model;
[0127] obtaining the initial character image corresponding to each character based on the third image sub-model of the text-to-image model according to the character information corresponding to each character.
[0128] In an exemplary embodiment, determining the pose of the character in the screenplay based on the initial character image, the character prompt word corresponding to the screenplay, and the character distribution position based on the first image sub-model of the preset text-to-image model comprises:
[0129] obtaining the target character image corresponding to the character based on the initial character image and the character prompt word corresponding to the screenplay based on the first image sub-model of the preset text-to-image model;
[0130] determining the pose of the character in the screenplay based on the target character image and the character distribution position.
[0131] In an exemplary embodiment, determining the pose of the character in the screenplay based on the target character image and the character distribution position comprises:
[0132] performing biological body detection and key point position detection on the target character image to obtain key point coordinates corresponding to the target character image;
[0133] mapping the key point coordinates to the character distribution position to obtain the pose of the character in the screenplay.
[0134] In an exemplary embodiment, determining the screenplay image corresponding to the screenplay based on the initial character image corresponding to each character in the screenplay, the pose of the character, and the screenplay prompt word based on the second image sub-model of the text-to-image model comprises:
[0135] inputting the initial character image corresponding to each character in the screenplay, the pose of the character, and the screenplay prompt word into the second image sub-model of the text-to-image model to perform fusion processing on the initial character image, the pose of the character, and the screenplay prompt word in the screenplay through the second image sub-model to obtain the screenplay image.
[0136] In an example embodiment, the second image sub-model comprises a first adapter network and a second adapter network; the initial character image in the storyboard, the posture of the character, and the storyboard prompt word are fused by the second image sub-model to obtain a storyboard image, comprising:
[0137] The posture information corresponding to the posture of the character is extracted by the first adapter network of the second image sub-model.
[0138] The appearance feature corresponding to the initial character image in the storyboard is extracted by the second adapter network of the second image sub-model.
[0139] The storyboard prompt word, the posture information, and the appearance feature are fused to obtain the storyboard image.
[0140] In an example embodiment, after the second image sub-model based on the text-to-image model determines the storyboard image corresponding to the storyboard according to the initial character image corresponding to each character in the storyboard, the posture of the character, and the storyboard prompt word, the method further comprises:
[0141] The storyboard images corresponding to the multiple storyboards are synthesized into a video.
[0142] It should be noted that, for the convenience and brevity of description, the specific working process of generating a comic image can be referred to the corresponding process in the embodiments of the comic image generation method described above, and will not be described here.
[0143] The embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium stores a computer program. The method implemented by the computer program executed by the processor can refer to each embodiment of the comic image generation method of the present application.
[0144] The computer readable storage medium can be an internal storage unit of the computer device, such as a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0145] It should be understood that the terms used herein in the specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and the appended claims of the present application, unless otherwise clearly indicated by the context, the singular forms "a", "an" and "the" are intended to include the plural forms as well.
[0146] It should also be understood that, in the specification and the appended claims, the terms "and / or" is used to mean one or more of the associated listed items, as well as any combination of any of the associated listed items. It will be apparent to those skilled in the art that various modifications and variations can be made to the present application without departing from the spirit or scope of the application. Thus, it is intended that the present application cover modifications and variations of this application provided they come within the scope of the appended claims and their equivalents.
[0147] The above-mentioned embodiment serial numbers of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments. The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of generating a caricature image, characterized by, The method comprises the following steps: acquiring role information corresponding to each role in a novel text based on a third language submodel of a preset language model; obtaining initial role images corresponding to each role based on a third image submodel of a preset text-to-image model according to the role information corresponding to each role; determining a shot text corresponding to at least one shot in the novel text based on a first language submodel of the language model; determining a shot prompt word, a role prompt word and a role distribution position of each role corresponding to the shot based on a second language submodel of the language model according to the shot text; obtaining target role images corresponding to the roles based on a first image submodel of the text-to-image model according to the initial role images and the role prompt word corresponding to the shot; performing biological body detection and key point position detection on the target role images to obtain key point coordinates corresponding to the target role images; mapping the key point coordinates to the role distribution position to obtain the poses of the roles in the shot; determining a shot image corresponding to the shot based on a second image submodel of the text-to-image model according to the initial role images corresponding to each role in the shot, the poses of the roles and the shot prompt word.
2. The caricature image generation method according to claim 1, characterized in that, The method of determining a shot image corresponding to the shot based on a second image submodel of the text-to-image model according to the initial role images corresponding to each role in the shot, the poses of the roles and the shot prompt word comprises the following steps: inputting the initial role images corresponding to each role in the shot, the poses of the roles and the shot prompt word into the second image submodel of the text-to-image model to perform fusion processing on the initial role images, the poses of the roles and the shot prompt word in the shot through the second image submodel to obtain a shot image.
3. The caricature image generation method according to claim 2, characterized in that, The second image submodel comprises a first adapter network and a second adapter network; the fusion processing on the initial role images, the poses of the roles and the shot prompt word in the shot through the second image submodel to obtain a shot image comprises the following steps: extracting pose information corresponding to the poses of the roles through the first adapter network of the second image submodel; extracting appearance features corresponding to the initial role images in the shot through the second adapter network of the second image submodel; fusing the shot prompt word, the pose information and the appearance features to obtain a shot image.
4. The method of claim 1-3, wherein, After the method of determining a shot image corresponding to the shot based on a second image submodel of the text-to-image model according to the initial role images corresponding to each role in the shot, the poses of the roles and the shot prompt word, the method further comprises the following steps: synthesizing a video from the shot images corresponding to the shots.
5. A caricature image generating apparatus characterized by comprising: The method comprises the following steps: acquiring role information corresponding to each role in a novel text based on a third language submodel of a preset language model; obtaining an initial role image corresponding to each of the roles according to role information corresponding to each of the roles based on a third image submodel of a preset text-to-image model; a text determination module configured to determine a shot text corresponding to at least one shot in the novel text based on a first language submodel of a preset language model; a position determination module configured to determine a shot prompt, a role prompt, and a role distribution position of each of the roles corresponding to the shot based on a second language submodel of the language model and the shot text; a pose determination module configured to obtain a target role image corresponding to each of the roles based on a first image submodel of a preset text-to-image model and the initial role image and the role prompt corresponding to the shot; performing biological body detection and key point position detection on the target role image to obtain key point coordinates corresponding to the target role image, and mapping the key point coordinates to the role distribution position to obtain a pose of the role in the shot; an image determination module configured to determine a shot image corresponding to the shot based on a second image submodel of the text-to-image model, the initial role image corresponding to each of the roles in the shot, the pose of the role, and the shot prompt.
6. A computer device, comprising: The computer device comprises a memory and a processor; The memory is configured to store a computer program; The processor is configured to execute the computer program and implement the method for generating a comic image according to any one of claims 1 to 4 when executing the computer program.
7. A computer-readable storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the steps of the method for generating a comic image according to any one of claims 1 to 4.
Citation Information
Patent Citations
Story video generation method and device, storage medium and equipment
CN117332118A
Cartoon image generation method and device, computer equipment and storage medium
CN117523064A