AI cartoon generation method and related equipment
Through the AI comic generation method, voice analysis and in-depth analysis technology are used to automatically generate comics and explain them, which solves the problem of difficult Chinese characters input in traditional methods, improves the efficiency and quality of comic generation, and improves the user experience.
Patent Information
- Application Number
- CN202510573897.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional comic generation methods are difficult to meet the needs of children, the elderly and people with disabilities, and text input is difficult.
Using the AI comic generation method, the step-by-step processing of the speech analytical sub-model, the comic generation sub-model and the explanation generation sub-model are automatically generated, including speech recognition, in-depth analysis, story generation, storyboard production and explanation text generation.
It improves the efficiency and quality of comic generation, ensures the consistency and logic of content, and improves the user's language understanding and cognitive abilities, especially the interactive experience between visually impaired and language learners.
Smart Images

Figure CN120495487A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer and communication technology, and more specifically, to an AI comic generation method and related equipment. Background Art
[0002] With the development of AI technology, more and more image generation applications have emerged. Traditional comic generation methods usually rely on manually input text or image files. However, for children, the elderly, and people with disabilities, text input is more difficult, and traditional comic generation methods cannot meet their needs. Summary of the Invention
[0003] The embodiments of the present application provide an AI comic generation method and related equipment, which can at least to a certain extent meet the usage needs of children, the elderly and the disabled.
[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0005] According to one aspect of an embodiment of the present application, an AI comic generation method is provided, comprising: receiving a target voice, the target voice containing story content; inputting the target voice into a comic generation model to obtain a target comic and a target explanation; wherein the comic generation model comprises a voice parsing sub-model and a comic generation sub-model, and inputting the target voice into the comic generation model to obtain a target comic and a target explanation specifically comprises: inputting the target voice into the voice parsing sub-model to extract a target story content; inputting the target story content into the comic generation sub-model to obtain a target comic; and inputting the target voice, the target story content, and the target comic into an explanation generation sub-model to obtain a target explanation.
[0006] According to one aspect of an embodiment of the present application, an AI comic generation device is provided, which includes: a target voice receiving module for receiving a target voice, wherein the target voice contains story content; a target comic generation module for inputting the target voice into a comic generation model to obtain a target comic and a target explanation; wherein the comic generation model includes a voice parsing sub-model and a comic generation sub-model, and the target comic generation module specifically includes: a voice parsing sub-module for inputting the target voice into the voice parsing sub-model to extract the target story content; a comic generation sub-module for inputting the target story content into the comic generation sub-model to obtain a target comic; and an explanation generation sub-module for inputting the target voice, the target story content, and the target comic into the explanation generation sub-model to obtain a target explanation.
[0007] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the AI comic generation method as described in the above embodiment is implemented.
[0008] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the AI comic generation method as described in the above embodiments.
[0009] According to one aspect of an embodiment of the present application, a computer program product is provided, comprising one or more computer programs, which, when executed by one or more processors, implement the steps of the AI comic generation method as described in the above embodiment.
[0010] In the technical solutions provided in some embodiments of this application, the three major tasks of speech analysis, comic generation, and explanation generation are handled step by step through sub-models with clear division of labor. The modules work together to achieve automated comic generation, improving processing efficiency and ensuring quality. The speech analysis sub-model can accurately extract key information from the user's speech, ensuring the quality and coherence of the subsequent comic and explanation content. The combination of the comic generation sub-model and the explanation generation sub-model ensures that the output content is not only a comic with both pictures and text, but also helps users understand the plot and improve their language comprehension and cognitive abilities.
[0011] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, explaining the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0013] Figure 1 A schematic diagram shows an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied.
[0014] Figure 2 A flow chart of an AI comic generation method provided in an embodiment of the present application is shown.
[0015] Figure 3 Shown according to Figure 2A specific implementation flowchart of step S210 in the AI comic generation method shown in the corresponding embodiment.
[0016] Figure 4 Shown according to Figure 2 A specific implementation flowchart of step S220 in the AI comic generation method shown in the corresponding embodiment.
[0017] Figure 5 Shown according to Figure 2 A specific implementation flowchart of step S230 in the AI comic generation method shown in the corresponding embodiment.
[0018] Figure 6 A specific implementation flow chart of a comic generation model training provided by an embodiment of the present application is shown.
[0019] Figure 7 A schematic diagram of the structure of an AI comic generation device provided in an embodiment of the present application is shown.
[0020] Figure 8 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0021] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.
[0022] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0023] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0024] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0025] Figure 1 A schematic diagram shows an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied.
[0026] like Figure 1 As shown, the system architecture may include terminal devices (such as Figure 1 101, tablet computer 102, and portable computer 103, which may also be a desktop computer, etc.), network 104, and server 105. Network 104 is a medium for providing a communication link between the terminal device and server 105. Network 104 can include various connection types, such as wired communication links, wireless communication links, etc.
[0027] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as needed. For example, the server 105 may be a server cluster consisting of multiple servers.
[0028] The user can use the terminal device to interact with the server 105 through the network 104 to receive or send messages, etc. The server 105 can be a server that provides various services. For example, the user uses the terminal device 103 (or the terminal device 101 or 102) to upload the target voice to the server 105. The target voice contains story content. The server 105 can input the target voice into the comic generation model to obtain the target comic and target explanation. Specifically, the comic generation model includes a speech parsing sub-model and a comic generation sub-model. The server 105 can first input the target voice into the speech parsing sub-model to extract the target story content; then input the target story content into the comic generation sub-model to obtain the target comic; finally, input the target voice, the target story content and the target comic into the explanation generation sub-model to obtain the target explanation.
[0029] It should be noted that the AI comic generation method provided in the embodiments of the present application is generally executed by the server 105, and accordingly, the AI comic generation device is generally provided in the server 105. However, in other embodiments of the present application, the terminal device may also have similar functions as the server, thereby executing the AI comic generation solution provided in the embodiments of the present application.
[0030] The following is a detailed description of the implementation details of the technical solution of the embodiment of the present application:
[0031] Figure 2 A flowchart of an AI comic generation method according to an embodiment of the present application is shown. The AI comic generation method can be executed by a server, which can be Figure 1 Refer to the server shown in . Figure 2 As shown, the AI comic generation method at least includes:
[0032] S100: Receive a target speech, where the target speech contains story content.
[0033] S200, inputting the target speech into a comic generation model to obtain a target comic and a target explanation.
[0034] Specifically, the comic generation model includes a speech analysis sub-model, a comic generation sub-model, and an explanation generation sub-model. The specific execution steps of S200 may include:
[0035] S210: Input the target speech into the speech analysis sub-model to extract the target story content.
[0036] S220: Input the target story content into the comic generation sub-model to obtain a target comic.
[0037] S230: Input the target speech, the target story content, and the target cartoon into an explanation generation sub-model to obtain a target explanation.
[0038] In the embodiments of this application, the three major tasks of speech analysis, comic generation, and explanation generation are handled step by step through clearly defined sub-models. The modules collaborate to achieve automated comic generation, improving processing efficiency and ensuring quality. The speech analysis sub-model accurately extracts key information from the user's speech, ensuring the quality and coherence of the subsequent comic and explanation content. The combination of the comic generation sub-model and the explanation generation sub-model ensures that the output content is not only a richly illustrated comic, but also helps users understand the plot and enhance their language comprehension and cognitive abilities.
[0039] At S100, user voice input is received. This voice input is typically a story narrated by the user, containing basic elements such as characters, scenes, and plot. The product of this application must be able to recognize and receive this voice signal, including filtering out background noise. This process relies on voice recognition technology to ensure that the voice input is correctly parsed and processed by the system.
[0040] In S200, the target speech is input into a comprehensive comic generation model. This model includes multiple sub-models, each responsible for different tasks, ensuring that the story can be accurately parsed, converted into comics, and explained.
[0041] In S210 , the target speech is converted into text and the main story content is extracted from it. The speech parsing sub-model uses speech recognition technology to recognize and convert the speech signal into text, while performing syntactic and semantic analysis to extract the main plot, characters, and scenes of the story.
[0042] For example, if the user says: "There is a puppy and a cat going on an adventure together", it will be parsed into "roles: puppy, cat, action: going on an adventure together".
[0043] Specifically, in some embodiments, the specific implementation of step S210 can be found in Figure 3 . Figure 3 is based on Figure 2 The detailed description of step S210 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S210 may include the following steps:
[0044] S212: Input the target speech into a speech engine to extract the core elements of the story.
[0045] S214: Generate target story content based on the core elements of the story.
[0046] In this embodiment, by introducing a speech engine for in-depth analysis, the core elements of the story can be more accurately extracted from the target speech, rather than simply converting speech to text. Accurate element extraction provides high-quality input for subsequent story generation, avoiding unnecessary misunderstandings and omissions. At the same time, through story generation based on core elements, the fragmented information in the speech is transformed into logical, complete, and engaging story content, ensuring that the content in the comic generation process is more compact and coherent, facilitating subsequent comic generation and explanation generation, and enhancing the user's interactive experience.
[0047] The above-mentioned extraction and generation methods are more adaptable. They can not only process the basic information in the user's voice, but also cope with more complex voice input, process story content of different types and styles, and provide targeted services for various user needs.
[0048] In S212, the target speech is input into the speech engine for processing. The speech engine automatically converts the content of the target speech using speech recognition and natural language processing technologies. The speech engine not only needs to recognize the words and sentences in the speech, but also needs to analyze the key information based on the grammar, context, and intonation of the speech. For example, when a child tells an adventure story, the speech engine needs to extract elements such as "characters" (such as a puppy or a cat), "actions" (such as going on an adventure together), and "locations" (such as a forest).
[0049] In this step, the system accurately extracts the core elements of the story through the speech engine's analysis of the voice input. These elements typically include characters, actions, scenes and locations, plot, and emotions. Characters are people, animals, or other important entities in the story. Actions describe the main actions or events performed by the characters. Scenes and locations are the locations or backgrounds of the story. Plot and emotions describe the main conflicts, developments, or emotions within the story.
[0050] The speech engine, through in-depth analysis of speech, can identify these elements and extract them as the basis for generating the target story content. The extracted core elements may only be isolated pieces of information, but this information will help to generate more precise content in subsequent steps.
[0051] Specifically, in some embodiments, the specific implementation of step S212 can refer to the following embodiments. Figure 3 The detailed description of step S212 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S212 may include the following steps:
[0052] The target speech is preprocessed to obtain effective speech information.
[0053] Analyze the effective voice information and extract the core elements of the story.
[0054] In this example, the preprocessing step removes noise, interference, and unnecessary parts of the speech, ensuring that the subsequent model processes only clear and valid speech information. Furthermore, by further analyzing the valid speech information, the core elements of the story (such as characters, actions, scenes, plot, etc.) can be accurately extracted, avoiding interference from irrelevant information, thereby improving the quality of subsequent story generation.
[0055] This embodiment ensures the quality of input information, and can generate more coherent, reasonable and attractive story content based on accurate core elements, thereby improving the effect of comic generation.
[0056] Specifically, after the target speech is input into the speech engine, it is pre-processed using noise cancellation technology to remove background noise and irrelevant sounds, eliminate repetitive parts of the speech, and make the speech logical. Common noise cancellation techniques include spectral filtering and time domain analysis.
[0057] If the target speech is long, preprocessing also involves segmenting it into smaller segments or slices. This allows for more precise analysis of each segment. A speech may consist of multiple parts, each representing a different storyline. Proper segmentation allows for separate analysis of each segment, ensuring that each contains useful information.
[0058] For low-quality speech data (e.g., those with a strong accent or rapid speech), pre-processing also includes speech enhancement. This algorithm optimizes low-quality speech data by adjusting the frequency range and clarity of the speech to improve its intelligibility.
[0059] In some embodiments, the pre-processed voice information can be converted into text through voice recognition technology. At this time, the system will recognize the vocabulary, grammatical structure, etc. in the voice and generate preliminary text information.
[0060] The pre-processed effective speech information (converted into text) will be analyzed, and natural language processing technology will be used to identify key information in the text. This information can be entities (such as characters, places, events), actions, emotions, etc., in order to accurately extract the core elements that help build the story. In the text, entities such as "characters", "actions" and "scenes" can be identified and classified. For example, if the text contains the words "cat" and "dog", these two entities can be identified as "characters"; if the text contains words such as "exploration" and "travel", it can be identified as an "action" element; if the text mentions nouns such as "forest" and "city", it can be classified as a "place" or "scene".
[0061] In some embodiments, the sentence structure in the text can also be analyzed to extract the plot and the relationship between the elements. For example, based on the sentence structure (such as "The cat met the dog in the forest"), the order of the plot can be extracted and the interactive relationship between "cat" and "dog" can be determined.
[0062] Sentiment analysis can also analyze the emotional tone revealed in speech. For example, if the speech contains words such as "horror" or "excitement," it can infer that these emotions may affect the development of the storyline and incorporate them into the generation of story content.
[0063] When analyzing effective speech information, we can also comprehensively consider grammatical structure, context, and contextual cues. For example, in a story for children, the characters of "cat" and "dog" may require more anthropomorphic language. Deep learning models can identify and understand the details of the context, ensuring that the generated core elements are more consistent with the intended storyline.
[0064] In S214, after extracting the core story elements, these elements need to be reorganized and supplemented to generate the complete target story content. Specifically, the extracted core elements are first integrated in a logical order to form a coherent story structure. For example, by combining the characters "puppy and cat," "adventure" as the action, and "forest" as the setting, a basic plot such as "In a mysterious forest, the puppy and cat go on an adventure together" can be formed. Building on this simple core element, the plot can be further expanded and enriched to make the story more complete and coherent. For example, when generating the target story, the action "adventure" can be used to further deduce the story's development, such as "They encountered a mysterious monster" or "They helped a lost animal." The generation of the target story content also needs to conform to the expression rules of natural language. Therefore, language optimization is performed during the generation process to ensure grammatical correctness and clear expression. Emotional color, description, or dialogue can be added to make the story content more engaging, especially for specific users.
[0065] Specifically, in some embodiments, the specific implementation of step S214 can refer to the following embodiments. Figure 3 The detailed description of step S214 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S214 may include the following steps:
[0066] Capture similar stories based on the core elements of the story.
[0067] The similar stories are templated to obtain a target story template.
[0068] Generate target story content based on the target story template and the story core elements.
[0069] In this embodiment, stories similar to the core elements of the target story are captured to ensure that the generated story content is consistent and logical, and the templating process allows different core elements to be adapted to the same structure, avoiding the monotony of creation and increasing the diversity of stories. At the same time, by pre-capturing similar stories and templating them, there is no need to create each story from scratch. Instead, the story can be quickly generated by relying on the templated structure, which greatly improves the generation speed and saves manual intervention and creation time. The templated approach can also avoid confusion or deviation from the theme of the story content while ensuring that the plot is reasonable. At the same time, the core elements can be flexibly filled in the template to ensure that the generated story is both innovative and meets user expectations.
[0070] First, the core elements of the story are extracted from the target voice information. These elements can be characters, actions, scenes, plots, etc. Each core element has a certain degree of semantic uniqueness and relevance. Once the core elements are extracted, a similarity calculation model (such as a natural language processing model based on deep learning) is used to search for stories with similar plots in the database. The plot structure, characters, scenes, etc. of these stories need to be highly matched with the core elements of the target story. The matching process may involve technologies such as text analysis, plot mapping, and topic classification. For example, if the core elements of the target story include "magic" and "adventure", stories with similar themes, such as "magic school" or "adventure travel" can be found in the database.
[0071] After finding similar stories, analyze their structure. The structure of a story includes its introduction, development, climax, and conclusion. Once you've identified these structural sections and ensure they're universal and flexible, you can directly create a template.
[0072] In some embodiments, the template-based process involves abstracting the story's structure and plot. For example, a "magician versus dragon" scenario might be transformed into a "hero versus enemy" scenario, or an "adventure at a magic school" scenario might be abstracted into a "school life + adventure elements" story framework. This abstraction allows the template to adapt to different plot inputs.
[0073] In other embodiments, during the templating process, some details in the story (such as character names, specific actions, etc.) can also be used as "replaceable" modules to ensure that these modules can be customized according to the specific core elements of the target story.
[0074] Finally, the captured similar story templates are combined with the core elements of the target story to automatically generate the complete story content. During this process, core elements (such as characters, scenes, and actions) can be added to the corresponding positions in the template to ensure that the generated story content meets the requirements of the target elements. The automatically generated story will include an introduction, development, climax, and conclusion, and each part will be closely connected and logically coherent, while fully reflecting the specific needs of the target story.
[0075] In some embodiments, the story content can be optimized or fine-tuned according to specific rules during the generation process. For example, the story rhythm, character behavior, or plot development may need to be adjusted to ensure that the final generated story content is more attractive and logical.
[0076] In S220, after obtaining the story content in text form, the comic generation sub-model generates a comic based on the extracted story content. The comic generation sub-model includes sub-modules such as scene generation, character design, and storyboard production. The system generates appropriate comic scenes based on the story plot. For example, if the story mentions a puppy and a cat going on an adventure together, comic scenes containing these two characters and their adventure scenes will be generated.
[0077] This step abstracts and decomposes the content, transforming it into elements suitable for comic expression.
[0078] Specifically, in some embodiments, the specific implementation of step S220 can be found in Figure 4 . Figure 4 is based on Figure 2 The detailed description of step S220 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S220 may include the following steps:
[0079] S222: Split the target story content to obtain the content of each storyboard.
[0080] S224: Generate a corresponding target comic according to the content of each storyboard.
[0081] In this embodiment, by splitting the target story content into multiple storyboards and generating corresponding comic images based on each storyboard, the story's presentation can be more precisely controlled, ensuring that the details of each storyboard are fully displayed and that the overall comic's structure and plot are clear. Splitting the story into storyboards can better showcase key elements in the comic, such as action, character expressions, and scene changes, making the storyline more readable and visually impactful. The comic generation sub-model can use a unified style and model parameters when generating each storyboard, ensuring that the entire comic remains consistent in style and avoiding style mismatches. In this way, the resulting comic is more coherent and aesthetically pleasing.
[0082] In S222, the overall content of the target story is first analyzed to identify key plot points. These points can include character action changes, scene transitions, emotional transitions, and so on. Based on these plot changes, the story content is then automatically split into multiple parts, each corresponding to a storyboard in the comic. Once the target story content is split, the corresponding content is extracted for each storyboard, including character settings, background settings, actions, and plot points.
[0083] Character design refers to the information about the characters included in each shot, such as their positions, clothing, and expressions. Background design refers to the background content of each scene, such as indoor scenes, street scenes, and natural environments. Action and plot design refers to the specific actions or key plot points that each shot should display, such as the character's movement trajectory and specific expressions.
[0084] Specifically, in some cases, these actions can be split into multiple storyboards according to the action paragraphs in the story (for example, "the protagonist starts running", "the enemy swings his weapon", etc.). Each storyboard corresponds to a different action step or stage. In other cases, whenever the scene of the story changes (such as from indoors to outdoors, from day to night), it can be split into new storyboards according to the scene changes. In some cases, it can be split into different storyboards according to the emotional changes of the characters in the story to show the changes in the characters' expressions or psychological activities.
[0085] Specifically, in some embodiments, the specific implementation of step S222 can refer to the following embodiments. Figure 4 The detailed description of step S222 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S222 may include the following steps:
[0086] The target story content is processed through an attention recurrent network to obtain each sentence vector.
[0087] The target story content is split according to the sentence vectors and the core elements of the story to obtain story segments.
[0088] According to the preset storyboard window value and each story segment, the content of each storyboard is obtained.
[0089] In this embodiment, the use of an attention recurrent network can better model each sentence in the target story, process and understand the relationship between sentences, accurately identify different fragments of the story, and then split them into reasonable storyboards. At the same time, by setting the storyboard window, the story fragments contained in each storyboard can be flexibly controlled, making the plot switching of the comic more natural, ensuring that the content of each storyboard is not only independent and complete, but also the connections between each other can also transition smoothly. Splitting based on the core elements of the story can ensure that the important plot in each storyboard is fully displayed, making the comic not only a presentation of the plot, but also a better expression of the emotion and theme of the story.
[0090] Specifically, the target story content is first input into a recurrent neural network (ARNN) based on the attention mechanism. Through network processing, each sentence in the target story is mapped into a sentence vector. This sentence vector contains the semantic information of the sentence and its relationship with the context. The role of the attention mechanism is to identify which sentences or words are most critical to the advancement and understanding of the plot in different contexts, thereby strengthening the semantic representation of these parts. Through the attention mechanism, it can capture the underlying structure and rhythm information of long stories, establish an accurate plot foundation at the sentence level, and lay the foundation for subsequent storyboarding.
[0091] After obtaining a vector representation for each sentence, further analysis and segmentation are performed based on the core story elements. These core elements may include character relationships, time changes, and plot climaxes. By analyzing these elements, it is determined which sentences or paragraphs belong to the same plot segment. Based on the sentence vectors and core elements, the target story content is automatically divided into multiple segments, each representing a scene or stage of story development in the comic. The segmentation of segments depends not only on the structure of the text itself but can also be adjusted based on external factors such as plot twists, emotional fluctuations, and character actions. This combined semantic analysis and core element segmentation method ensures that each segment contains sufficient information, avoids overly detailed or overly coarse segmentation, and improves the coherence of the story.
[0092] The frame window value is a crucial parameter that determines the size of the story segment contained in each frame. This value can be set based on factors such as the frequency of plot changes and the speed of scene transitions. For example, a dynamic, rapidly evolving plot may require a shorter frame window, while a static, emotionally rich scene may require a longer frame window. Once the story is broken down into segments, these segments can be segmented based on the frame window value to generate specific frame content. Each frame will contain one or more story segments and clearly indicate the key elements of each frame, such as characters, setting, action, and background. The frame window allows you to control the rhythm and emotional expression of each frame, ensuring that each frame in the comic carries sufficient story information. This also creates a more natural transition and connection between frames, reducing the sense of fragmentation in the plot.
[0093] In S224, the content of each storyboard is input into a dedicated comic generation sub-model. The model generates a corresponding comic image by understanding and analyzing the content of each storyboard (characters, actions, background, etc.) and combining it with the comic style requirements.
[0094] Specifically, the model needs to draw the character's posture, expression, and interaction with other elements based on the character settings and action requirements in each storyboard. Comic style details (such as line thickness and shading) are also determined at this stage. Then, based on the background settings in the storyboard, the corresponding scene image is generated. The system needs to ensure that the scene layout conforms to the visual standards of comics, such as perspective relationships, lighting and shadow effects, etc. Finally, the style of the entire comic is unified. For example, the character drawing style (such as cartoon style, realistic style, etc.) needs to be consistent across all storyboards. In addition, color, line treatment, composition methods, etc. also need to be unified to ensure the overall coordination of the comic work. During the generation process, each storyboard image can be optimized and details adjusted to make it more consistent with the comic's artistic style and plot expression. For example, subtle adjustments can be made to the character's expression, or more details can be added to the background to enhance the layering of the picture.
[0095] After all the storyboards are generated, these storyboard images are integrated into a complete comic work to ensure that the plot of the entire comic is smooth and visually coherent.
[0096] Specifically, in some embodiments, the specific implementation of step S224 can refer to the following embodiments. Figure 4 The detailed description of step S224 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S224 may include the following steps:
[0097] According to the content of each storyboard, a corresponding storyboard key frame is generated.
[0098] A storyboard frame sequence is formed according to the storyboard key frames.
[0099] A target comic is generated according to the storyboard frame sequence.
[0100] In this embodiment, by generating storyboard keyframes and forming a frame sequence based on them, the content of each storyboard can be precisely controlled to ensure that the composition of each picture conforms to the original storyline and rhythm. At the same time, by generating a frame sequence, the transition between pictures can be guaranteed to be natural, avoiding abruptness or discontinuity. The entire comic generation process, from storyboard content to target comic, is fully automated. This greatly improves the efficiency of comic creation, reduces the time and cost of manual drawing, and can speed up the comic production process while maintaining creativity. At the same time, through precise keyframe generation and sequence arrangement, it can ensure that each frame has high-quality performance in both visual and narrative aspects, enhancing the overall effect of the comic. At the same time, it is also possible to reasonably adjust the number of frames and picture content of each storyboard according to the changes in the plot and the needs of the character's actions, making the entire comic more vivid and expressive.
[0101] Specifically, the system takes the storyboard content as input and automatically identifies key plot elements, such as character actions, emotions, and background changes, based on each storyboard's description. Artificial intelligence algorithms (such as convolutional neural networks or generative adversarial networks) are used to generate corresponding keyframes—the core image in each storyboard. Keyframes typically include important elements of the scene, such as characters, actions, and environment, forming a static image that captures the core of the storyboard's content.
[0102] It should be noted that each storyboard contains a certain paragraph or scene of the storyline, which may include information such as characters, background, and actions.
[0103] Each generated keyframe is then analyzed to determine the temporal or plot relationship between each frame, ensuring the order and logic of the storyboard content. Based on the timing and plot development of each keyframe, these keyframes are arranged in an appropriate order to form a storyboard frame sequence. The storyboard frame sequence determines the arrangement order and composition layout of each page and each frame of the comic. The generation of the frame sequence requires the system to be able to flexibly arrange the display of each storyboard according to the rhythm of the plot, including how to transition between pictures and how to express emotional changes. By properly arranging the frame sequence, the story's plot coherence can be ensured, and the visual transition can be smooth, avoiding incoherent or difficult-to-understand parts when reading.
[0104] Finally, the generated storyboard frame sequence is used as input, and the actual comic book production process begins based on the information in this frame sequence. Specifically, each comic page is automatically generated based on the frame sequence. Each comic page consists of multiple consecutive storyboards, and the complete comic page is gradually generated based on the previously generated storyboard keyframes and frame sequences. During this process, necessary detailed optimizations can also be performed, such as adjusting character expressions, modifying background elements, and enhancing visual effects.
[0105] In S230, the target speech, story content, and generated comic are fed into the explanation generation sub-model. This model generates accompanying explanation text, typically a brief explanation or description of the comic. For example, if the generated comic depicts a puppy and a cat on an adventure, the explanation generation sub-model might generate explanation text for this scene, such as "This is a brave puppy exploring a mysterious forest with the cat." This text is automatically generated based on the storyline, using language that is as concise as possible and tailored to the specific user's cognitive level.
[0106] Specifically, in some embodiments, the specific implementation of step S230 can be found in Figure 5 . Figure 5 is based on Figure 2 The detailed description of step S230 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S230 may include the following steps:
[0107] S232: Generate explanation text corresponding to each target comic according to the target voice, the target comic and the target story content.
[0108] S234, converting the explanation text into explanation voice, and associating the explanation voice with the target comic to form a target explanation.
[0109] In this embodiment, the automatically generated explanation text and voice can be accurately matched to each page and each frame of the comic, ensuring a high degree of consistency between the explanation content and the visual content, enhancing the user's sense of immersion. The addition of the explanation voice also makes the comic no longer just a static picture, but a dynamic expression of sound and explanation content, increasing the interactivity with the reader. The plot expression of the comic is no longer solely dependent on the picture. The ups and downs of sound and the expression of voice can enhance the transmission of emotion, improve the narrative effect and appeal of the comic. The explanation voice helps users better understand the content of the comic, and is particularly suitable for the visually impaired, language learners, and readers who want more background information.
[0110] This embodiment deeply integrates the target voice, story content and comics to generate personalized and multi-dimensional explanation content, which not only enhances the artistry and fun of the comics, but also improves the overall user experience.
[0111] In S232, the target story content is first analyzed to extract key information, such as characters, scenes, actions, and dialogue. Next, based on the comic's visual content and combined with the target audio prompts, a corresponding explanatory text is generated. For example, if the comic depicts a conversation, the generated explanatory text might include a description of the emotional tone and background of the conversation.
[0112] Based on the combination of the target voice and the target story content, one or more explanatory texts are automatically generated. These explanatory texts will explain the screen content in detail and may include information such as the characters' emotions, the development of the story, and changes in the environment.
[0113] Specifically, in some embodiments, the specific implementation of step S232 can refer to the following embodiments. Figure 5 The detailed description of step S232 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S232 may include the following steps:
[0114] An explanation element is obtained according to the target voice and the target story content.
[0115] The explanation elements are matched with the target comics to obtain matching results.
[0116] Based on the matching results, an explanation text corresponding to each target comic is generated.
[0117] In this embodiment, by extracting explanation elements from the target voice and target story content and matching them with the target comic content, it is ensured that the generated explanation text can accurately reflect the plot and details of the comic screen, avoiding inconsistent or mismatched content. At the same time, through the feedback of the matching results, a more targeted explanation text can be generated according to the specific comic plot, improving the accuracy and quality of the text. The explanation content is not only synchronized with the comic screen, but also more in line with the user's reading experience. Personalized explanation text is generated based on the content of different comics (such as characters, plots, etc.). The explanation text of each comic will be different according to its specific content, enhancing the expressiveness and interactivity of the comic.
[0118] Specifically, the target story content is analyzed through natural language processing (NLP) technology, and key explanation elements are extracted based on a comprehensive analysis of the plot, voice and story content.
[0119] The above-mentioned explanation elements may include character information, emotional information, event information, and background information.
[0120] The target comic's content is then analyzed to identify key elements, including characters, plot, and scenes. Image recognition technology can be used to parse each page or frame of the comic, identifying objects, actions, and expressions. The extracted explanatory elements are then matched with individual frames in the target comic. For example, if a frame in the comic shows Character A smiling happily, the system can match the corresponding "Character A" smile with the "happy" emotion information.
[0121] In some embodiments, a matching result can be generated for each frame, marking the required explanation content for each frame, including how to describe the character's emotions, actions, and relationship with the scene and dialogue.
[0122] Based on the matching results, a detailed explanation text is generated according to the matched explanation elements (such as characters, emotions, background, actions, etc.). For example, if a frame in a comic shows character A looking worriedly into the distance, the generated explanation text may be: "Character A looks worried, staring into the distance, as if thinking about something important." In some embodiments, the tone of the explanation text can also be adjusted according to the emotional atmosphere and context of the comic. For example, in a tense or suspenseful scene, the tone may be more tense or low; while in a humorous scene, the tone may be more relaxed. After the explanation text is generated, grammar checking and sentence optimization can be performed to ensure that the text is smooth and natural and fits the context.
[0123] In step S234, after the explanation text is generated, speech synthesis technology (such as TTS, Text-to-Speech) can be used to convert the text into speech. Speech synthesis technology analyzes the sentence structure, punctuation, modal particles, and other elements in the explanation text to generate speech output that matches the text content.
[0124] In some embodiments, models trained with deep learning techniques (such as neural network-based speech synthesis models) can also be used to ensure the naturalness and intelligibility of the speech. The converted speech must not only have clear pronunciation but also convey the correct emotion and context.
[0125] The generated narration voice can be further optimized, such as adjusting the speaking speed, intonation, and timbre, to make it more suitable for specific scenes or plots. For example, the voice may need to be more passionate during certain plot climaxes, while it may need to be softer during sad or calm scenes.
[0126] After the explanation speech is generated, it is associated with the target comic.
[0127] This association can be achieved through timestamps, meaning each comic page or frame is marked with a start time, and the corresponding commentary is played based on this time. In some cases, dynamic analysis of the comic content can also be used to determine when to play which audio segment, ensuring a perfect match between the comic and the commentary. This association not only provides explanations for each comic page, but also dynamically adjusts the playback of the commentary based on user actions (such as turning pages, pausing, fast-forwarding, etc.), further enhancing the user experience. This allows users to gain a richer emotional experience and background knowledge while reading comics.
[0128] Specifically, in some embodiments, the specific implementation of step S234 can refer to the following embodiments. Figure 5 The detailed description of step S234 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S234 may include the following steps:
[0129] The explanation text is compared with the target speech to determine the similarity of the explanation texts.
[0130] If the similarity of the explanation text is greater than a predetermined similarity threshold, the target speech is segmented according to the explanation text to obtain a corresponding explanation speech.
[0131] If the similarity of the explanation text is less than a predetermined similarity threshold, the explanation text is converted into explanation speech.
[0132] The explanation voice is associated with the target comic to form a target explanation.
[0133] In this embodiment, by determining the similarity between the explanation text and the target voice, optimization can be made when generating the explanation voice. If the explanation text is similar to the target voice, existing voice resources can be directly utilized, saving generation time; if the similarity is low, a new voice is directly generated to avoid voice mismatches. After the explanation voice is generated, it is associated with the target comic, so that each explanation text can be matched with the most appropriate voice, ensuring a high degree of consistency between the voice and the comic content, and enhancing the user's immersive experience. Dynamically adjusting the generation method based on the similarity between the explanation text and the target voice enhances the flexibility and adaptability of voice generation, allowing the generated explanation voice to more accurately reflect the comic content, avoid unnecessary voice generation, and save computing resources. When the similarity is high, existing voice resources are directly used to avoid regenerating the voice, thereby improving generation efficiency and reducing the consumption of computing resources.
[0134] Specifically, speech analysis technology is used to compare the explanation text and the target speech. Natural language processing (NLP) technology is usually used to analyze the text content and match it with the audio features of the target speech (such as pitch, tone, and speaking speed) to determine their similarity. Similarity can be quantified by calculating the differences between the two in terms of semantics, emotion, tone, etc. Common methods include text similarity algorithms (such as cosine similarity and Jaccard similarity) and speech feature similarity analysis. For example, if the content, emotion, and tone of the target speech are highly consistent with those of the explanation text, the similarity is high.
[0135] If the similarity exceeds a predetermined threshold, the explanation text closely matches the target speech, and the target speech can be directly segmented. The target speech can be segmented based on the segmentation of the explanation text to generate the corresponding explanation speech. This saves speech generation time and ensures consistency between speech and text. Specifically, segmentation can be performed based on the segmentation of the explanation text to ensure that the generated explanation speech matches the scenes or plot segments of the comic content. For example, the switching position of the speech can be adjusted according to the content of each comic page.
[0136] If the similarity falls below a predetermined threshold, the explanation text and the target speech are insufficiently matched, and the explanation speech needs to be regenerated. Specifically, Text-to-Speech (TTS) technology can be used to convert the explanation text directly into speech, adjusting the speech's tone, speed, pitch, and other characteristics based on the text's emotion and context.
[0137] The generated audio commentary is associated with the target comic, ensuring that the audio corresponds to each frame or page of the comic. Synchronization between the audio commentary and the comic's visuals can be achieved based on the comic's content, scene changes, and plot progression. For example, when the comic's plot changes, the corresponding audio commentary automatically switches. Through the close coordination of audio and visuals, a target commentary is ultimately generated. The target commentary is a combined audio and visual output. While viewing the comic, users can simultaneously hear the audio commentary that matches the visuals, enhancing their immersion.
[0138] Please refer to Figure 6 , the training method of the above comic generation model specifically includes:
[0139] S102: Obtain a target speech sample set, where the target speech sample set includes a plurality of target speech samples, each of which is pre-marked with a corresponding comic label and explanation label.
[0140] S104: Input all target speech samples in the target speech sample set into a cartoon generation model one by one to obtain output target cartoons and target explanations.
[0141] S106, updating parameters of the comic generation model according to the output target comic and target explanation and the marked comic label and explanation label until a predetermined end condition is reached, and the training is terminated to obtain a trained comic generation model.
[0142] Specifically, in some embodiments, the specific implementation of step S106 can refer to the following embodiments. Figure 6 The detailed description of step S106 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S106 may include the following steps:
[0143] If, in the target speech sample set, less than a predetermined number of target speech samples are input into the comic generation model, and the target comic and target explanation outputted therefrom are consistent with the marked comic label and explanation label, then updating the parameters of the comic generation model;
[0144] If, in the target speech sample set, more than a predetermined number of target speech samples are input into the comic generation model and the target comic and target explanation output are consistent with the marked comic label and explanation label, then the predetermined end condition is met, the training is ended, and a trained comic generation model is obtained.
[0145] Specifically, in other embodiments, the specific implementation of step S106 can refer to the following embodiments. Figure 6The detailed description of step S106 in the AI comic generation method shown in the corresponding embodiment, in the AI comic generation method, step S106 may include the following steps:
[0146] Determine a loss function based on the output target cartoon and target explanation and the marked cartoon label and explanation label;
[0147] The parameters of the comic generation model are updated according to the loss function until a predetermined end condition is reached, and the training is terminated to obtain a trained comic generation model.
[0148] In this embodiment, the loss function reaching the predetermined end condition may be that the loss function converges or the loss function is less than a predetermined loss (eg, 0.001).
[0149] In the above embodiment, the website risk is determined by using a neural network model obtained through multiple trainings to obtain corresponding cartoon labels and explanation labels. The neural network model is obtained through multiple trainings. The more samples it trains, the more accurate the results obtained. It basically does not require maintenance during operation, which also reduces maintenance costs and improves the efficiency and accuracy of website risk determination.
[0150] The following describes an embodiment of the device of the present application, which can be used to execute the AI comic generation method described in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the AI comic generation method described above in the present application.
[0151] Figure 7 A block diagram of an AI comic generation device according to an embodiment of the present application is shown.
[0152] Reference Figure 7 As shown, the AI comic generation device 700 according to one embodiment of the present application includes a target speech receiving module 710 and a target comic generation module 720.
[0153] The target speech receiving module 710 is used to receive the target speech, which contains story content; the target comic generation module 720 is used to input the target speech into the comic generation model to obtain the target comic and target explanation.
[0154] Among them, the above-mentioned comic generation model includes a speech analysis sub-model and a comic generation sub-model, and the target comic generation module 720 specifically includes a speech analysis sub-module 722, a comic generation sub-module 724 and an explanation generation sub-module 726.
[0155] The speech analysis submodule 722 is used to input the target speech into the speech analysis submodel to extract the target story content; the comic generation submodule 724 is used to input the target story content into the comic generation submodel to obtain the target comic; the explanation generation submodule 726 is used to input the target speech, the target story content and the target comic into the explanation generation submodel to obtain the target explanation.
[0156] In some feasible embodiments of the present application, the speech analysis submodule 722 specifically includes: a core element extraction unit, which is used to input the target speech into the speech engine to extract the core elements of the story; and a story content generation unit, which is used to generate the target story content based on the core elements of the story.
[0157] In some feasible embodiments of the present application, the core element extraction unit specifically includes: an effective voice information subunit, which is used to preprocess the target voice to obtain effective voice information; and a story core element subunit, which is used to parse the effective voice information and extract the core elements of the story.
[0158] In some feasible embodiments of the present application, the story content generation unit specifically includes: a similar story capture sub-unit, which is used to capture similar stories based on the core elements of the story; a target story template sub-unit, which is used to template the similar stories to obtain a target story template; and a story content generation sub-unit, which is used to generate target story content based on the target story template and the core elements of the story.
[0159] In some feasible embodiments of the present application, the comic generation submodule 724 specifically includes: a story content splitting unit, used to split the target story content to obtain each storyboard content; a target comic generation unit, used to generate a corresponding target comic according to each storyboard content.
[0160] In some feasible embodiments of the present application, the story content splitting unit specifically includes: a sentence vector sub-unit, which is used to process the target story content through an attention recurrent network to obtain each sentence vector; a story fragment sub-unit, which is used to split the target story content according to each sentence vector and the core elements of the story to obtain each story fragment; and a storyboard content sub-unit, which is used to obtain each storyboard content according to a preset storyboard window value and each story fragment.
[0161] In some feasible embodiments of the present application, the target comic generation unit specifically includes: a storyboard key frame sub-unit, used to generate corresponding storyboard key frames according to the content of each storyboard; a storyboard frame sequence sub-unit, used to form a storyboard frame sequence according to the storyboard key frames; and a target comic sub-unit, used to generate a target comic according to the storyboard frame sequence.
[0162] In some feasible embodiments of the present application, the explanation generation submodule 726 specifically includes: an explanation text generation unit, which is used to generate an explanation text corresponding to each target comic based on the target voice, the target comic and the target story content; a target explanation formation unit, which is used to convert the explanation text into an explanation voice, and associate the explanation voice with the target comic to form a target explanation.
[0163] In some feasible embodiments of the present application, the explanation text generation unit specifically includes: an explanation element sub-unit, which is used to obtain explanation elements based on the target voice and the target story content; a matching result sub-unit, which is used to match the explanation elements with each target comic to obtain a matching result; and an explanation text sub-unit, which is used to generate an explanation text corresponding to each target comic based on the matching result.
[0164] In some feasible embodiments of the present application, the target explanation forming unit specifically includes: a similarity comparison subunit, used to compare the explanation text with the target voice to determine the similarity of the explanation text; a first voice subunit, used to segment the target voice according to the explanation text to obtain the corresponding explanation voice if the similarity of the explanation text is greater than a predetermined similarity threshold; a second voice subunit, used to convert the explanation text into an explanation voice if the similarity of the explanation text is less than a predetermined similarity threshold; a comic association subunit, used to associate the explanation voice with the target comic to form a target explanation.
[0165] In some feasible embodiments of the present application, the AI comic generation device also includes: a sample acquisition module, used to obtain a target voice sample set, the target voice sample set contains multiple target voice samples, and each target voice sample is pre-marked with a corresponding comic label and explanation label; a sample input module, used to input all target voice samples in the target voice sample set into the comic generation model one by one to obtain output target comics and target explanations; a model parameter adjustment module, used to update the parameters of the comic generation model according to the output target comics and target explanations and the marked comic labels and explanation labels until a predetermined end condition is reached, the training is ended, and a trained comic generation model is obtained.
[0166] In the embodiments of this application, the three major tasks of speech analysis, comic generation, and explanation generation are handled step by step through clearly defined sub-models. The modules collaborate to achieve automated comic generation, improving processing efficiency and ensuring quality. The speech analysis sub-model accurately extracts key information from the user's speech, ensuring the quality and coherence of the subsequent comic and explanation content. The combination of the comic generation sub-model and the explanation generation sub-model ensures that the output content is not only a richly illustrated comic, but also helps users understand the plot and enhance their language comprehension and cognitive abilities.
[0167] Figure 8 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown.
[0168] It should be noted that Figure 8 The computer system of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0169] like Figure 8 As shown, the computer system includes a central processing unit (CPU) 1801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1802 or the program loaded from the storage part 1808 into the random access memory (RAM) 1803, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 1803. The CPU 1801, ROM 1802 and RAM 1803 are connected to each other via a bus 1804. An input / output (I / O) interface 1805 is also connected to the bus 1804.
[0170] The following components are connected to the I / O interface 1805: an input section 1806 including a keyboard, a mouse, and the like; an output section 1807 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 1808 including a hard disk; and a communication section 1809 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to the I / O interface 1805 as needed. Removable media 1811, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1810 as needed, so that computer programs read from the removable media can be installed in the storage section 1808 as needed.
[0171] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1809, and / or installed from a removable medium 1811. When the computer program is executed by the central processing unit (CPU) 1801, the various functions defined in the system of the present application are executed.
[0172] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0173] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0174] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0175] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device implements the method described in the above embodiments.
[0176] This specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figures 1 to 6 The method of the embodiment shown, the specific execution process can be found in Figures 1 to 6 The detailed description of the illustrated embodiment will not be repeated here.
[0177] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0178] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0179] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.
[0180] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. An AI comic generation method, characterized in that: The AI comic generation method includes: receiving a target speech, wherein the target speech includes story content; Inputting the target speech into a comic generation model to obtain a target comic and a target explanation; The comic generation model includes a speech analysis sub-model and a comic generation sub-model. Inputting the target speech into the comic generation model to obtain the target comic and target explanation specifically includes: Inputting the target speech into the speech parsing sub-model to extract the target story content; Inputting the target story content into the comic generation sub-model to obtain a target comic; The target speech, the target story content and the target comic are input into the explanation generation sub-model to obtain the target explanation.
2. The AI comic generation method according to claim 1, wherein: Inputting the target speech into the speech parsing sub-model to extract the target story content specifically includes: Inputting the target speech into a speech engine to extract the core elements of the story; Generate target story content based on the core elements of the story.
3. The AI comic generation method according to claim 2, wherein: The step of inputting the target speech into a speech engine and extracting the core elements of the story specifically includes: Preprocessing the target speech to obtain effective speech information; Analyze the effective voice information and extract the core elements of the story.
4. The AI comic generation method according to claim 2, wherein: Generating target story content according to the core elements of the story specifically includes: Capture similar stories based on the core elements of the story; Template the similar stories to obtain a target story template; Generate target story content based on the target story template and the story core elements.
5. The AI comic generation method according to claim 1, wherein: Inputting the target story content into the comic generation sub-model to obtain the target comic specifically includes: Splitting the target story content to obtain the content of each storyboard; Generate corresponding target comics based on the content of each storyboard.
6. The AI comic generation method according to claim 5, wherein: The target story content is split to obtain the contents of each storyboard, specifically including: Process the target story content through the attention recurrent network to obtain each sentence vector; Splitting the target story content according to the sentence vectors and the core elements of the story to obtain story segments; According to the preset storyboard window value and each story segment, the content of each storyboard is obtained.
7. An AI comics generating device, characterized in that: The AI comic generation device includes: A target speech receiving module is used to receive a target speech, wherein the target speech contains story content; A target comic generation module is used to input the target speech into a comic generation model to obtain a target comic and a target explanation; The comic generation model includes a speech analysis sub-model and a comic generation sub-model, and the target comic generation module specifically includes: A speech analysis submodule, configured to input the target speech into the speech analysis submodel to extract target story content; A comic generation submodule, configured to input the target story content into the comic generation submodel to obtain a target comic; The explanation generation submodule is used to input the target speech, the target story content and the target cartoon into the explanation generation submodel to obtain the target explanation.
8. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the AI comic generation method according to any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the AI comic generation method according to any one of claims 1 to 6.
10. A computer program product comprising one or more computer programs, characterized in that When the one or more computer programs are executed by one or more processors, the steps of the AI comic generation method according to any one of claims 1 to 6 are implemented.