Content generation method, electronic device, storage medium, program product, and program
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-07-31
AI Technical Summary
The current e-books suffer from low efficiency in generating illustrations and videos, and it is difficult to ensure the consistency of key elements in images and videos, resulting in a poor user reading experience.
By performing semantic understanding on the story text, the target plot is automatically determined and the invariant features of key elements are identified. Then, a machine learning model is used to generate videos, ensuring the consistency of key elements.
It improves the efficiency and accuracy of video generation, provides a richer and more vivid visual experience, and enhances the user's reading experience.
Smart Images

Figure CN122498153A_ABST
Abstract
Description
Content generation methods, electronic devices, storage media, program products and programs Technical Field
[0001] This disclosure relates to the fields of artificial intelligence and computer technology, and in particular to a content generation method, electronic device, storage medium, program product, and program. Background Technology
[0002] With the development of internet technology, more and more users are reading e-books through e-reading platforms, applications, and other means.
[0003] Currently, most e-books are displayed primarily in text format. Some e-books also include images created by the author alongside the text to help users better understand the words. Summary of the Invention
[0004] According to some embodiments of this disclosure, a content generation method is provided, including: performing semantic understanding on the text of a story to determine a target plot in the story for generating a video; performing semantic understanding on the target plot and its context to determine key elements and invariant features corresponding to the key elements; generating a video corresponding to the target plot based on information about the invariant features and the target plot; and displaying the video corresponding to the target plot.
[0005] According to other embodiments of this disclosure, an electronic device is provided, including: a processor; and a memory coupled to the processor for storing instructions that, when executed by the processor, cause the processor to perform a content generation method according to any embodiment of this disclosure.
[0006] According to further embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, wherein when executed by a processor, the program causes the processor to implement the content generation method of any embodiment of the present disclosure.
[0007] According to further embodiments of the present disclosure, a computer program product is provided, including instructions that, when executed by a processor, cause the processor to perform a content generation method as described in any embodiment of the present disclosure.
[0008] According to further embodiments of the present disclosure, a computer program is provided, including instructions that, when executed by a processor, cause the processor to perform a content generation method as described in any embodiment of the present disclosure.
[0009] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0010] Embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the drawings described below are merely illustrative of some embodiments of this disclosure and are not intended to limit the scope of this disclosure. In the drawings:
[0011] Figure 1 shows a flowchart illustrating a content generation method according to some embodiments of this disclosure;
[0012] Figures 2 to 5 are schematic diagrams of the display interfaces of some embodiments of this disclosure;
[0013] Figure 6 shows a schematic diagram of the system architecture of some embodiments of this disclosure;
[0014] Figure 7 shows a schematic diagram of the structure of a content generation apparatus according to some embodiments of the present disclosure;
[0015] Figure 8 shows a schematic diagram of the structure of an electronic device according to some embodiments of the present disclosure;
[0016] Figure 9 shows a schematic diagram of the structure of an electronic device according to some embodiments of the present disclosure. Detailed Implementation
[0017] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0018] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of components and steps set forth in these embodiments should be interpreted as merely exemplary and does not limit the scope of this disclosure.
[0019] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". The term "based on" means "at least partially based on".
[0020] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.
[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.
[0023] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0024] In the context of this disclosure, "image" can refer to any of a variety of images, such as color images, grayscale images, etc. It should be noted that the type of image is not specifically limited in the context of this specification. Furthermore, an image can be any suitable image, such as a raw image obtained by a camera device, or an image from which specific processing has been performed, such as preliminary filtering, dealiasing, color adjustment, contrast adjustment, normalization, etc. It should be noted that preprocessing operations may also include other types of preprocessing operations known in the art, which will not be described in detail here.
[0025] In e-books, especially story-based e-books, displaying images alongside text enhances the reading experience. Currently, most e-book illustrations are still hand-drawn by creators, resulting in high costs and low efficiency. With the development of artificial intelligence and machine learning models, some solutions exist that use large models to generate images from text. Generating illustrations for e-books requires manually selecting paragraphs and inputting them into a large model to generate corresponding images. This process also requires human intervention. Furthermore, inputting multiple paragraphs into a large model often results in inconsistencies in the images; for example, the same character may appear different in different images. Directly using these images as illustrations provides a poor user experience. Maintaining image consistency requires manual adjustments. Moreover, the inventors realized that images are static, and for dynamic scenes in stories, such as battle scenes or climaxes, illustrations alone are insufficient to provide adequate reading assistance. Adding videos to the story text could provide a richer, more vivid, and concrete visual experience, enabling users to better understand the content. However, generating videos from story text also suffers from the aforementioned problems: it requires manual selection of suitable segments to input into the model and manual assurance of consistency among characters and scenes in the video. Because videos are far more complex than images, inconsistencies in characters and scenes are even more pronounced in videos.
[0026] To address the aforementioned issues, this disclosure proposes a content generation method that performs semantic understanding on the story text, automatically identifies the target plot for video generation, and then, through semantic understanding of the target plot and its context, identifies key elements and their corresponding invariant features. Based on the invariant feature information and the target plot, the method generates and displays the video corresponding to the target plot. This content generation method automatically determines the target plot for video generation and automatically identifies the invariant features corresponding to key elements, improving the efficiency of video generation. Furthermore, generating the video based on invariant feature information ensures the consistency of some features of key elements in the video, making the generated video more accurate. In addition, the video presentation provides users with a richer, more vivid, and concrete visual experience, better assisting users in reading and enhancing the user experience.
[0027] Figures 1 to 6 below illustrate some embodiments of the content generation method of this disclosure.
[0028] Figure 1 is a flowchart of some embodiments of the content generation method of this disclosure. As shown in Figure 1, the method of this embodiment includes steps S102 to S108.
[0029] In step S102, semantic understanding is performed on the text of the story to determine the target plot in the story used to generate the video.
[0030] The target plot is a plot suitable for generating a video. For example, target plots include at least one of the following: plots involving key events for one or more characters, plots that drive the plot forward, flashbacks, plots describing scenes, and plots describing atmosphere, not limited to the examples given. These plots, used to generate videos, can help users better understand the story's content and are suitable for presentation in video format, providing a better visual experience.
[0031] A primary machine learning model can be used to perform semantic understanding of the entire story's text, automatically selecting target plots for video generation. Examples of primary machine learning models include LLM (Large Language Model), but are not limited to the examples given.
[0032] In step S104, semantic understanding is performed on the target plot and its context to determine key elements and the invariant features corresponding to the key elements.
[0033] The context of a target plot can include one or more paragraphs preceding and / or following the target plot, as well as paragraphs describing the characteristics of key elements in the story. For example, paragraphs introducing a character in the story typically describe the character's appearance and other characteristics. If a character appears in the target plot, paragraphs describing that character's characteristics can serve as the context of the target plot.
[0034] For example, key elements include at least one type of element from the character, object, scene, or environment. By semantically understanding the target plot and its context, one or more key elements in the target plot can be identified, and the invariant features of each key element can be determined. For example, if the protagonist A in the target plot is a key element, the protagonist A's physical appearance or facial features remain unchanged.
[0035] The first machine learning model can be used to perform semantic understanding of the target plot and its context, and to determine the key elements and the invariant features corresponding to the key elements.
[0036] In step S106, a video corresponding to the target plot is generated based on the information of the unchanged features and the target plot.
[0037] Information about invariant features can take the form of descriptive information about the invariant features, keywords, related images, etc. Referring to the information about invariant features in the generated video corresponding to the target plot can ensure that these features of key elements in the generated video remain consistent, thereby improving the accuracy of the generated video.
[0038] For example, a second machine learning model can be used to generate a video corresponding to the target plot based on information about invariant features and the target plot. This second machine learning model can generate video directly from text, or it can first generate multiple images from the text and then generate a video from those images.
[0039] In step S108, the video corresponding to the target plot is displayed.
[0040] Videos corresponding to target plots can be displayed in one or more distribution scenarios. For example, videos corresponding to target plots can be displayed in the reading interface or the audio playback interface. Identifying one or more key target plots allows for the display of corresponding videos in the story's recommendation or introduction interface. The display method and location of videos corresponding to target plots can be configured according to actual needs and are not limited to the examples given above.
[0041] In the method described in the above embodiments, semantic understanding of the story text is performed to automatically determine the target plot in the story used to generate the video. Then, through semantic understanding of the target plot and its context, key elements and their corresponding invariant features are determined. Based on the information of the invariant features and the target plot, a video corresponding to the target plot is generated and displayed. The method described in the above embodiments can automatically determine the target plot used to generate the video and can automatically identify key elements and their corresponding invariant features without human intervention, thus improving the efficiency of video generation. Furthermore, generating the video based on invariant feature information ensures the consistency of some features of key elements in the video, making the generated video more accurate. In addition, by presenting the video corresponding to the target plot, the text description is visualized and made more concrete, providing users with a richer, more vivid, and specific visual experience, better assisting users in reading, enabling users to understand the content of the target plot more quickly and accurately, and improving the user experience.
[0042] The following describes, with some examples, how to perform semantic understanding on the text of a story to determine the target plot in the story used to generate a video.
[0043] In some embodiments, a first machine learning model is used to perform semantic understanding on the text of the story to identify one or more characters in the story and the plot related to one or more characters; semantic understanding is performed on the plot related to one or more characters to identify key events corresponding to each of the one or more characters in the development of the plot; and the plots corresponding to the key events are identified as target plots for generating videos.
[0044] The story text and a first prompt can be input into a first machine learning model. The first machine learning model performs semantic understanding of the story text based on the first prompt, identifying one or more characters and their related plot points, and identifying key events corresponding to each character as target plot points. For example, the first prompt includes task description information, which instructs the first machine learning model to perform a deep understanding of the story, such as understanding the plot, character settings, themes, story structure, core plot points, and key information points, understanding their roles in the story's development, logical causal relationships, and chronological order. The task description information also instructs the first machine learning model to identify key events corresponding to each character. For example, a key event corresponding to a character may include at least one of the following: a breakthrough event in the character's growth, a crisis event related to the character, a key event in the character's emotional change, or a conflict event related to the character, and is not limited to the examples given. For example, a character's breakthrough event may include an event where the character learns a certain skill; a crisis event related to the character may include an event where the character faces a life-or-death crisis; a key event in the character's emotional change may include an event where the character makes a romantic confession; a conflict event related to the character may include an event where the character fights an enemy, etc., and is not limited to the examples given.
[0045] Key events for a character can influence plot development, showcase a character's growth, or reflect emotional changes. By automatically identifying these key events and using their corresponding plot points as target scenes, videos can be generated to better assist users in reading and to help them understand the story's content and development more quickly and accurately.
[0046] In some embodiments, a first machine learning model is used to perform semantic understanding of the text of the story to identify at least one plot point, either a plot that drives the plot or a flashback plot, as a target plot point for generating a video.
[0047] The first machine learning model can perform semantic understanding of the story's text based on the first cue information, and identify at least one plot point—either a plot that drives the story forward or a flashback plot—as the target plot point. For example, the first cue information includes task description information, which instructs the first machine learning model to perform a deep understanding of the story and identify plot points that drive the story forward and flashback plots.
[0048] There may be overlap between plot points that drive the storyline and key events for specific characters, but the process of generating videos will not be performed multiple times for the same plot point. For example, plot points that drive the storyline include at least one of the following: plot twists, revelations of important information, and conflict or contradictions, and are not limited to the examples given. For instance, plot twists include situations where the plot suddenly changes direction just when the user expects it to unfold in a certain way; revelations of important information include the revelation of the protagonist's background; and conflict or contradictions include large-scale battles, and are not limited to the examples mentioned above.
[0049] For a flashback scene, related scenes can also be identified as target scenes. For example, if the flashback scene recalls the process of protagonists A and B meeting and falling in love, then the process of protagonists A and B meeting and falling in love can also be considered a target scene.
[0050] By automatically identifying plot points and flashback sequences that drive the story forward, and using these as target plot points, the generated videos can more vividly and concretely showcase the plot's development, highlighting key moments and helping users recall relevant content. This better assists users in reading and allows for a faster and more accurate understanding of the story's content and development.
[0051] In some embodiments, a first machine learning model is used to perform semantic understanding of the text of the story to determine plots describing at least one of the scenes and atmosphere as target plots for generating a video.
[0052] The first machine learning model can perform semantic understanding of the story text based on the first cue information to determine the plot describing at least one of the scene and atmosphere as the target plot. For example, the first cue information includes task description information, which instructs the first machine learning model to perform deep understanding of the story and identify plots describing at least one of the scene and atmosphere. For example, plots describing the scene include plots describing beautiful scenery; plots describing the atmosphere include plots describing tense, cheerful, or other atmospheres, and so on, and are not limited to the examples given.
[0053] By automatically identifying plot points that describe scenes and atmosphere, and using these plot points as target plots to generate videos, users can feel immersed in the story and have a better reading experience when watching the videos.
[0054] The various embodiments for determining the target plot described above can be executed individually or in any combination. The first prompt message may also include one or more examples of the target plot.
[0055] By selecting a target scenario suitable for video generation from the story, it is possible to further determine the key elements in the target plot and the invariant features corresponding to the key elements. The following describes, with some examples, how to perform semantic understanding on the target plot and its context to determine the key elements and the invariant features corresponding to the key elements.
[0056] In some embodiments, a first machine learning model is used to perform semantic understanding on the target plot and the context of the target plot, and to determine multiple elements corresponding to the target plot, wherein the multiple elements belong to at least one type of character, item, scene, environment; to determine whether the features of each of the multiple elements change in the target plot; elements whose features remain partially or completely unchanged are identified as key elements, and the unchanged features corresponding to the key elements are determined.
[0057] The first machine learning model uses semantic information from the target plot and its context to determine multiple elements corresponding to the target plot. For example, the target plot includes elements corresponding to the two character types A and B, the scene type element corresponding to the room, and elements corresponding to multiple item types such as tables and chairs within the room. Based on the understanding of the target plot and its context, the features of each element can be determined, as well as whether those features change within the target plot. For example, it can be determined that the facial features and clothing of characters A and B remain unchanged, and the shape and color of the chairs in the room remain unchanged, but the position of the chairs has changed.
[0058] For example, after the first machine learning model determines the target plot, it can perform semantic understanding of the target plot and its context based on the second prompt information to identify key elements and their corresponding invariant features. For instance, the second prompt information may include task description information, which instructs the first machine learning model to identify key elements and their corresponding invariant features. The second prompt information may also include examples of one or more key elements and their invariant features.
[0059] The method in the above embodiments uses a first machine learning model to perform semantic understanding of the target plot and its context, automatically determines key elements and the invariant features corresponding to the key elements, and improves the efficiency and accuracy of subsequently generating content corresponding to the target plot.
[0060] In some embodiments, the key element includes key elements of at least one of the character type or item type. For each key element of a character type or item type, in response to the absence of a description of the appearance of the key element in the text of the target plot, the appearance features of the key element are determined based on the context of the target plot and using an associative function, and the appearance features are used as invariant features corresponding to the key element.
[0061] The target scenario may not necessarily describe the physical characteristics of every key element in detail. However, generating the video requires the physical characteristics of each element. Therefore, you can use the context of the target scenario to make associations and supplement the physical characteristics of key elements. For example, the target scenario may only describe that there is a table in the room, but it does not specifically describe the shape of the table. You need to combine the target scenario and its context to associate the shape of the table, such as its color, shape, and material. If it is a wealthy family, the table may be more luxurious and larger; conversely, the table may be older and smaller.
[0062] For example, if the text of the target plot does not include the location description text corresponding to the key element and the location of the key element remains unchanged, the location feature of the key element is determined based on the context of the target plot and using the association function, and the location feature is used as the unchanging feature corresponding to the key element.
[0063] A first machine learning model can be used to determine the shape and location features of key elements based on the context of the target scenario and second prompts, employing associative functions. The second prompts include task description information, which instructs the first machine learning model to associate the shape and location features of key elements.
[0064] The method described in the above embodiments uses the associative function to determine the shape features, position features, and other features of key elements as the unchanging features corresponding to the key elements. This can more accurately determine the unchanging features corresponding to the key elements and make the content of the images in the subsequently generated video more accurate.
[0065] In some embodiments, key elements include key elements of at least one type in a scene and environment. For each key element of a scene type or environment type, features of fixed items corresponding to the key elements are generated based on the target plot and the context of the target plot, and using an associative function. The features of the fixed items are then identified as the invariant features corresponding to the key elements.
[0066] The target scenario may only provide a simple description of the scene or environment, requiring the use of associative thinking to supplement it with the characteristics of fixed objects within the scene and environment. For example, if the target scenario only describes night, one could associate it with the moon in the night sky.
[0067] A first machine learning model can be used to determine the features of fixed items corresponding to key elements based on the context of the target scenario and second prompt information, and by employing an association function. The second prompt information includes task description information, which instructs the first machine learning model to associate the features of the fixed items.
[0068] The method described in the above embodiment uses the association function to determine the features of the fixed items corresponding to the key elements. As the unchanging features corresponding to the key elements, the unchanging features corresponding to the key elements can be determined more accurately, and the content of the images in the subsequently generated video can be more accurate.
[0069] The following describes, with reference to some embodiments, how to generate a video corresponding to a target plot based on information about unchanging features and the target plot.
[0070] In some embodiments, storyboard text is generated based on the text of the target plot, wherein the storyboard text includes a picture description text and a shooting method description text for each of the multiple shots; multiple images are generated based on information of invariant features and the storyboard text; and a video corresponding to the target plot is generated based on the multiple images.
[0071] To generate smoother, more vivid videos with richer content, actual filming techniques can be simulated. Based on a semantic understanding of the target plot, the images in the video are divided into images corresponding to different shots, and descriptive text for each shot and descriptive text for the filming method are determined. Each shot may include one or more frames, and the descriptive text for each frame may include feature descriptions of various elements within the frame, such as keywords. For example, the descriptive text for the filming method includes a description of the shot type, such as close-up, medium shot, and long shot.
[0072] For example, a first machine learning model can be used to generate storyboard text based on the text of the target scene and third-party cueing information. The third-party cueing information includes task description information, which instructs the first machine learning model to generate storyboard text based on the text of the target scene. The third-party cueing information may also include requirements for the format and content of the output storyboard text, as well as relevant examples of generated storyboard text.
[0073] The above embodiment generates storyboard text based on the text of the target plot, then generates multiple images, and finally generates a video based on the multiple images. On the one hand, this can improve the quality of the video, making it smoother and more vivid. On the other hand, each shot can maintain the consistency of some important features, thus improving the accuracy of the generated video.
[0074] In some embodiments, a first machine learning model is used to divide the text of the target plot into text corresponding to each shot; based on the text corresponding to each shot, multiple elements in each shot and feature description information of each of the multiple elements are determined; based on the feature description information of the multiple elements in each shot, a picture description text for each shot is generated; the main elements in each shot are determined; and based on the main elements in each shot and the content of the target plot, a shooting method description text for each shot is generated.
[0075] When a shot corresponds to multiple frames, the feature description information of each element in each frame can be determined, and a frame description text for each frame can be generated. The frame description texts of the multiple frames corresponding to the shot are then used as the frame description text for that shot. For example, the frame description text for each frame includes feature description information of at least one element from the characters, objects, scenes, and environment.
[0076] Based on the semantic information of the text corresponding to each shot, the main elements in each shot can be identified. For example, the main elements might be people or objects. The target plot will contain descriptive text corresponding to the shooting method. For example, "He rode a horse galloping from afar," which indicates that the shot can be fixed in a certain position, filming the process from far to near. For example, for a particular shot, it can be identified whether the text for that shot is a close-up of the main element; if it is, a close-up shooting method is used. For example, for a particular shot, if it is identified as a scene with multiple people, a long shot shooting method is used. The third cue information can also include some examples of identifying the shooting method, which can assist the first machine learning model in understanding how to determine the shooting method.
[0077] The method described in the above embodiments identifies multiple elements in each shot or frame, specifically determines the feature description information of each element, and combines it with the text corresponding to each shot to determine the shooting method description text for each shot, which can improve the accuracy and richness of the content of the subsequently generated video.
[0078] In some embodiments, based on the text corresponding to each shot, multiple elements in each shot are identified, wherein the multiple elements belong to at least one type: person, object, scene, environment; it is determined whether the text corresponding to the shot containing the person type element includes descriptive information of the person's specific characteristics, wherein the person's specific characteristics include at least one of action, posture, and expression; in response to the absence of descriptive information of the person's specific characteristics in the text corresponding to the shot containing the person type element, descriptive information of the person's specific characteristics corresponding to the person type element is generated based on the context using an associative function.
[0079] For elements of the character type, if the text corresponding to the shot containing the element includes descriptive information about the character's specific characteristics, then that descriptive information is retrieved. If the text corresponding to the shot containing the element does not include such descriptive information, the context can be searched to see if such descriptive information exists. If the context also does not contain such descriptive information, the association function can be used to generate descriptive information about the character's specific characteristics based on the semantic information of the text corresponding to the shot containing the element and its context. For example, if character A is an element, and the text corresponding to the shot containing character A does not include descriptive information about at least one of character A's actions, posture, or expression, the context can be combined with the association function to generate descriptive text indicating that character A's posture is standing, facing the camera, and has a calm expression.
[0080] The task description information in the third prompt can also instruct the first machine learning model to parse or generate description information for at least one of the character's actions, postures, and expressions.
[0081] The method described in the above embodiments identifies elements of a person type and generates descriptive information of at least one of the person's actions, postures, and expressions based on the association function and context. This can make the subsequently generated video more vivid, expressive, and accurate.
[0082] In some embodiments, it is determined whether the text corresponding to each shot includes descriptive information of an item type element; for text that does not include descriptive information of an item type element, based on the text and its context, an association function is used to determine the item type element corresponding to the text and the feature description information of the item type element.
[0083] Some shots may only include descriptions of characters, scenes, and environments in their corresponding text, excluding descriptions of specific objects. It's possible to determine if the context includes descriptive information about object types. Further, the association function can be used to supplement the object type elements in the scene and determine their characteristic descriptions. Alternatively, in shots with few object type elements (e.g., fewer than a threshold), the association function can be used to supplement the object type elements and their characteristic descriptions based on the shot's text and context; or, if the characteristic descriptions of object type elements are insufficiently detailed, the association function can be used to supplement their characteristic descriptions based on the shot's text and context.
[0084] The task description information in the third prompt can also instruct the first machine learning model to parse or generate feature description information for the item.
[0085] The method described in the above embodiments can enrich the descriptive information of the shot, making the content of the generated video richer and enhancing the visual effect.
[0086] In some embodiments, it is determined whether the text corresponding to each shot includes descriptive information of scene and / or environment type elements; for text that does not include descriptive information of scene and / or environment type elements, based on the text and its context, an association function is used to determine the scene and / or environment type elements corresponding to the text and the feature descriptive information of the scene and / or environment type elements.
[0087] For each shot, combining the corresponding text and its context can generally determine the elements and their characteristic descriptions of the scene and environment type. If this is not possible, the association function can be used. Alternatively, when the characteristic descriptions of the elements of the scene and / or environment type are insufficiently detailed, the association function can be used to expand upon these elements. For example, if the text describes the protagonist walking on a street, the street is a scene element; the association function can be used to expand upon scene elements such as buildings and traffic lights on the street, along with their characteristic descriptions.
[0088] The task description information in the third prompt can also instruct the first machine learning model to parse or generate feature description information of the scene and environment.
[0089] The method described in the above embodiments can enrich the descriptive information of the shot, making the content of the generated video richer and enhancing the visual effect.
[0090] Based on the descriptions of the foregoing embodiments, a first machine learning model can be used to generate storyboard text based on third prompt information. For example, the third prompt information may include task description information to describe the task performed by the first machine learning model. For instance, the task description information may instruct the first machine learning model to generate storyboard text, or to identify elements such as people, objects, scenes, and environments in a shot or frame and determine their corresponding feature description information. For elements such as people, objects, scenes, and environments, the corresponding feature description information can be determined through parsing or association. For people, feature description information such as facial expressions, postures / actions needs to be determined. The third prompt information may also include explanations and relevant examples of the feature description information of people, objects, scenes, and environments.
[0091] For example, third-party prompts may include: identifying the main character from the context and writing down their name; determining the character's current facial expression based on the contextual description or dialogue; describing the character's actions and postures based on the context, and if not mentioned, associating them with the scene, such as whether the posture is standing / sitting, facing the camera, facing away from the camera, or turned to the side. For example, third-party prompts may include: analyzing important objects appearing in the same scene and describing them and their state or location, such as a character using a phone or receiving a text message. For example, third-party prompts may include: analyzing the scene; for a scene that is clearly happy and joyful, describing the atmosphere with good weather or bright light. For example, third-party prompts may include: deciding on the shot type based on the elements and content of the shot. For example, using a long shot for scenes with many people; using a medium shot for general scenes and most of the time. For example, third-party prompts may include: identifying or associating the main object from the text description, such as the moon for 'night'. Third-party prompts may include format requirements for the output storyboard text, examples of input and output, etc., and are not limited to the examples given.
[0092] After generating the storyboard text, multiple images can be generated based on the storyboard text. In some embodiments, based on the scene description text and shooting method description text of each shot, as well as information on invariably invariant features, a prompt message corresponding to each shot is generated; the prompt message corresponding to each shot is input into a second machine learning model to obtain the generated image of each shot.
[0093] Information on invariant features includes descriptive text for the invariant features and / or images corresponding to the invariant features. For example, the protagonist's facial features and clothing features are invariant features. The descriptive text of these features can be used as part of the cue information corresponding to the protagonist's shots, and input into the second machine learning model. Alternatively, the protagonist's image can be used as part of the cue information corresponding to the protagonist's shots, and input into the second machine learning model. For each character (role) or main character (character) in the story, an image of that character can be generated based on the descriptive information. When generating images for each shot subsequently, using this image ensures consistency in the appearance of the same character across different images, improving accuracy.
[0094] The prompts for each shot can include a description of the shot's image, a description of the shooting method, and information about invariant features. The second machine learning model is an image generation model; by inputting the prompts for each shot into the second machine learning model, it can generate an image of the shot for each shot.
[0095] The method described in the above embodiment can perform semantic understanding of the story text using a first machine learning model to determine the target plot, and then extract information on key elements and their invariant features. Based on the target plot, it generates image description text and shooting method description text for each shot. Then, based on the image description text, shooting method description text, and invariant feature information for each shot, it generates prompt information corresponding to each shot. The prompt information corresponding to each shot is input into a second machine learning model to obtain the generated image of each shot, thereby improving the accuracy of the generated image and the accuracy of the subsequently generated video.
[0096] For example, a third machine learning model can be used to generate a video corresponding to a target plot based on multiple images.
[0097] In some embodiments, a third machine learning model is used to perform semantic understanding on each of the multiple images to identify one or more important elements in each image. Based on the change information of each important element across the multiple images and the semantic information of the target plot, a dynamic effect is determined for each important element. Based on the multiple images and the dynamic effects of each important element, a video corresponding to the target plot is generated. For example, for the environment, dynamic effects can be determined based on changes in the environment (rain, lightning) within the target plot; for the characters, dynamic effects can be determined based on the target plot (punching a fist, walking, etc.).
[0098] For example, multiple images and a fourth cue can be input into a third machine learning model. This fourth cue might include task description information, which instructs the third machine learning model to understand each image and determine dynamic effects for each important element. The multiple images could be keyframes from a video, rather than all frames, which are then further expanded to generate the video.
[0099] The method described above combines the understanding of multiple images and the target plot to determine the important elements and dynamic effects in the multiple images, and then generates a video corresponding to the target plot based on the multiple images and the dynamic effects of each important element, making the generated video more accurate, vivid and smooth.
[0100] In some embodiments, a target rhythm is determined based on the semantic information of the target plot, and a video corresponding to the target plot is generated based on the target rhythm, multiple images, and the dynamic effects of each important element.
[0101] Some target plots have a slow and relaxed pace, such as a scene describing a beautiful landscape, while others have a fast and tense pace, such as a battle scene. The pacing of the generated video needs to be controlled according to the pacing of the target plot. The faster the target pacing, the faster the generated video can be, and the shorter its duration; the slower the target pacing, the slower the generated video can be, and the longer its duration can be.
[0102] Based on the method described in the above embodiments, the generated video can better match the content of the target plot, conform to the user's reading habits, and improve the user experience.
[0103] In some embodiments, a target duration is determined based on the duration information of the target plot read by the user group, and a video corresponding to the target plot is generated based on the target duration, multiple images, and the dynamic effects of each important element.
[0104] The user group can be multiple users who have read the target content, and the duration information can be the average, maximum, or minimum duration of reading the target content by multiple users. Based on the duration information and a preset deviation value, the target duration can be determined, thus ensuring that the generated video has the target duration.
[0105] The method described above controls the duration of the generated video based on the user group's behavioral habits, making the generated video more compatible with the user's habits and improving the user experience.
[0106] It can also generate videos corresponding to target plots based on target duration, target rhythm, multiple images, and dynamic effects of each important element.
[0107] The above embodiments describe methods for generating images and videos corresponding to target scenes. Videos may also include audio components; the following embodiments describe how to generate the audio components in a video.
[0108] In some embodiments, narration text is generated based on the text of the target plot; narration audio is generated based on the narration text; and a video corresponding to the target plot is generated based on multiple images and narration audio.
[0109] The video corresponding to the target plot can be an explanatory video of the target plot, so that users can better understand the target plot and the content of the video.
[0110] In some embodiments, semantic understanding is performed on the text of the target plot, and a colloquial text describing the target plot from a preset perspective is generated based on the text of the target plot; key parts in the target plot are determined based on the semantic information of the text of the target plot; and the text corresponding to the key parts in the colloquial text describing the target plot from a preset perspective is expanded to generate explanatory text.
[0111] The first machine learning model can be used to generate narration text based on the fifth prompt and the target plot. The fifth prompt includes task description information, which can instruct the description of the target plot from a predetermined perspective. For example, it could describe it from the first-person perspective of the protagonist, changing the titles of various characters. The task description can also instruct the use of colloquial language. Furthermore, it can instruct the expansion of text corresponding to key parts, such as focusing on character traits, identities, actions, behaviors, psychological activities, emotions, and dialogues, highlighting the interactions between different characters to enhance vividness and appeal. In addition, the fifth prompt can also include constraints, such as avoiding repetitive descriptions, not altering content, and ensuring natural transitions.
[0112] The method described in the above embodiments can generate more fluent and natural narration text. While presenting the content of the target plot through video, the narration audio can help users better understand the plot. Furthermore, the narration audio has a different perspective and expression style from the target plot, which can make users feel more fresh and improve the user experience.
[0113] In some embodiments, the target rhythm is determined based on the semantic information of the target plot, and narration audio is generated based on the target rhythm and the narration text.
[0114] In some embodiments, narration audio is generated based on the duration of time a user group reads the target plot, the target duration, and the narration text.
[0115] Matching the rhythm of the target plot and / or the behavioral habits of the user group when generating narration audio can improve the accuracy of the generated audio and enhance the user experience.
[0116] The playback of the generated audio narration must match the playback of the video visuals; therefore, multiple images and videos can be generated based on the narration text. For example, based on the narration text, a storyboard text is generated; based on information about invariably retained features and the storyboard text, multiple images are generated; and based on the multiple images, a video corresponding to the target scene is generated. Furthermore, generating a video corresponding to the target scene based on multiple images can include: generating a video corresponding to the target scene based on the audio duration of a segment in the narration text corresponding to each of the multiple images.
[0117] Alternatively, a reference video can be generated first based on multiple images, and then narration text can be generated based on the playback time information of the reference video and the text of the target plot, so that the narration audio and video visuals are more matched.
[0118] The audio portion of the video may also include sound materials such as ambient sound effects, motion sound effects, and background audio. These sound materials can be automatically generated to improve the quality and effect of the video. In some embodiments, semantic understanding is performed on the text of the target plot to determine the descriptive text of the sound materials, wherein the sound materials include at least one of ambient sound effects, emotional sound effects, motion sound effects, event sound effects, and background music; sound materials are generated based on the descriptive text of the sound materials; and a video corresponding to the target plot is generated based on multiple images and sound materials.
[0119] For example, the system can identify the scene corresponding to the target plot and generate appropriate environmental sound effects based on that scene. For instance, it can generate bird sounds in a forest. It can also use changes in environmental sound effects to represent the passage of time or changes in the tension of the plot. For example, as a fierce battle is about to begin, the environmental sound effects can gradually change from calm, natural sounds to tense sounds like howling wind and beating war drums, hinting at the impending conflict.
[0120] For example, identifying the emotional state of characters in a target plot and generating corresponding emotional sound effects based on that state. For example, playing somber, slow music when a character is sad. For example, using rhythmic and tonal variations in emotional sound effects to highlight emotional fluctuations. For example, using sound effects with unstable rhythms and fluctuating pitches when a character is experiencing inner struggles to express their complex psychological state.
[0121] For example, identifying action descriptions in a target scene and generating corresponding sound effects. Examples include the sounds of weapons clashing in battle and the footsteps of characters running. For example, utilizing the continuity and rhythm of sound effects to simulate the fluidity and speed of action. For example, in a chase scene, rapid footsteps and hurried breathing sounds are played continuously.
[0122] For example, it can identify special events or key plot twists in the target storyline and generate corresponding event sound effects. For instance, when the protagonist solves an important puzzle, a crisp hint sound effect can be played, like the sound of unlocking a door.
[0123] In addition, sound materials can include personalized sound effects for specific characters or scenes. For example, the sound effects when the main character appears.
[0124] Background music can be generated based on at least one of the following: story style, plot type and rhythm, emotional type, and character personality. For example, background music can match the story's style and type, and conform to the historical context. For example, background music can be synchronized with the plot development and rhythm, adjusting its rhythm according to the pace and speed of the target plot; for example, tense plots might use music with a strong, dynamic rhythm. For example, background music can express a character's emotional state. For example, background music can be used to portray a character's personality; as a character displays different aspects of their personality, the background music can change accordingly.
[0125] In some embodiments, the target rhythm is determined based on the semantic information of the target plot, and sound material is generated based on the target rhythm and the narration text.
[0126] In some embodiments, audio material is generated based on the duration of time a user group reads the target plot, the target duration, and the narration text.
[0127] The timing and duration of audio playback can be reasonably controlled according to reading speed and the pace of plot development, thereby improving the quality of the generated audio materials and bringing users a better audiovisual experience.
[0128] In some embodiments, the dialogue of the characters in the target plot is determined, and the dialogue audio is generated based on the character settings and dialogue of each character.
[0129] Character profiles are determined based on a semantic understanding of the story. For example, character profiles include personality traits, age, and gender. This results in audio dialogue that is more closely aligned with the characters, enhancing the user experience.
[0130] In addition to generating videos, this disclosure can also generate individual images and / or audio.
[0131] In some embodiments, semantic understanding is performed on the text of the story to determine the target plot in the story and the type of generated content corresponding to the target plot. The type of generated content includes at least one of audio, image, and video.
[0132] If the type of generated content corresponding to the target plot is audio, generate the corresponding audio based on the target plot. If the type of generated content corresponding to the target plot is image, generate the corresponding image based on the target plot. If the type of generated content corresponding to the target plot is video, generate the corresponding video based on the target plot.
[0133] For a target plot used to generate an image, the image description text can be generated based on the text of the target plot, and the image can be generated based on the image description text.
[0134] For a target plot used to generate audio, a description text for the audio can be generated based on the text of the target plot, and then the audio can be generated based on the description text.
[0135] For example, target plots for video generation can be selected from the story first. Then, suitable plots for generating still images can be selected from the story. For instance, based on semantic understanding of the story's text, scenes, characters, actions, environments, and objects suitable for image representation can be identified. The generated images match the story's type, style, and author's style, maintaining consistency with the text style.
[0136] For example, identifying target plot points from a story that are suitable for audio representation. If a target plot point contains descriptions of specific sounds or scenes, then corresponding audio is generated. For instance, the target plot point might describe the roar of machines, the chirping of cicadas, or specific scenes like a bustling market.
[0137] Machine learning can be used to generate images and audio, and the methods described in the preceding embodiments will not be repeated here. The timing and duration of audio playback can be reasonably controlled according to reading speed and the pace of the plot, improving the quality of the generated sound materials and providing users with a better audiovisual experience.
[0138] The method described in the above embodiments can identify target plots in a story used to generate different types of content, generate different types of content to complement the text, and provide users with a better audiovisual experience.
[0139] In some embodiments, a video playback control is displayed in the display interface corresponding to the target plot, wherein the display interface corresponding to the target plot includes a reading interface for the target plot and / or an audio playback interface corresponding to the target plot; in response to the triggering of the playback control, the video corresponding to the target plot is played.
[0140] As mentioned in the previous embodiments, the video corresponding to the target plot can be displayed in various distribution scenarios. For example, as shown in Figure 2, the reading interface can display the target plot 201, and below the target plot, a video preview area 202 and a playback control 203 can be displayed. In response to the user's triggering of the playback control 203, the video can be played in the video preview area, or it can be played in a new page, window, panel, or other interface. The user can select an audiobook mode, in which the audio playback interface can also display the video playback control.
[0141] In some embodiments, an image corresponding to the target plot is displayed in the display interface corresponding to the target plot. The display interface corresponding to the target plot includes a reading interface for the target plot and / or an audio playback interface for the target plot.
[0142] As shown in Figure 3, for the target plot that generates the corresponding image, the target plot 301 can be displayed in the reading interface, and the image 302 is below the target plot.
[0143] In some embodiments, a playback control for the generated audio corresponding to the target plot is displayed in the display interface corresponding to the target plot. The display interface corresponding to the target plot includes a reading interface for the target plot and / or an audio playback interface for the target plot.
[0144] As shown in Figure 4, the target plot 201 is "beep beep beep", which is suitable for generating the corresponding sound effect audio. An audio playback control 402 can be displayed in the appropriate position. In response to the user's triggering of the audio playback control 402, the sound effect audio can be played.
[0145] As shown in Figure 5, the audio playback controls for background music and sound effects can be displayed in different forms and / or in different locations. An audio playback control 501 for background music can be displayed in the reading interface, playing background music in response to the user's input to the audio playback control 501.
[0146] The story text can generate various forms of content, including audio, images, and videos. As shown in Figure 6, the story text can be processed by the plot understanding module 610, the content generation module 620, and the content presentation module 630 to obtain multimodal interactive content. The plot understanding module 610 includes an image understanding module 611, an audio understanding module 612, and a video understanding module 613, which are used to determine the target plot for generating images, audio, and video, and to generate corresponding prompts for the images, audio, and video, respectively. Refer to the relevant content of the first machine learning model in the aforementioned embodiment. The content generation module 620 includes an image generation module 621, an audio generation module 622, and a video generation module, which are used to generate images, audio, and video, respectively. The content presentation module 630 includes an image presentation module 631, an audio presentation module 632, and a video presentation module 633, which are used to display information related to images, audio, and video, and to play audio and video, etc.
[0147] This disclosed content generation method can create an atmosphere and mood that matches the storyline, enhancing the story's appeal and emotional resonance. It allows users to more intuitively understand the textual descriptions, attracting their attention and deepening their impression of the story's development. It can also enrich and enhance the characters' personalities, giving users a more direct understanding of their traits and temperaments. Furthermore, it can alleviate visual fatigue, making the reading process more relaxed and enjoyable.
[0148] This disclosure also provides a content generation apparatus, which will be described below with reference to FIG7.
[0149] Figure 7 is a structural diagram of some embodiments of the content generation apparatus of this disclosure. As shown in Figure 7, the content generation apparatus 70 of this embodiment includes: a first determining module 710, a second determining module 720, a generation module 730, and a display module 740.
[0150] The first determining module 710 is configured to perform semantic understanding on the text of the story to determine the target plot in the story used to generate the video.
[0151] The second determining module 720 is configured to perform semantic understanding of the target plot and the context of the target plot, and to determine key elements and the invariant features corresponding to the key elements.
[0152] The generation module 730 is configured to generate a video corresponding to the target plot based on information about the unchanged features and the target plot.
[0153] Display module 740 is configured to display the video corresponding to the target plot.
[0154] In some embodiments, the first determining module 710 is configured to use a first machine learning model to perform semantic understanding on the text of the story, determine one or more characters in the story and the plot related to the one or more characters; perform semantic understanding on the plot related to the one or more characters, determine the key events corresponding to each of the one or more characters in the development of the plot; and determine the plot corresponding to the key events as the target plot for generating the video.
[0155] In some embodiments, the first determining module 710 is configured to use a first machine learning model to perform semantic understanding on the text of the story and determine at least one plot point, either a plot that drives the plot or a flashback plot, as a target plot point for generating a video.
[0156] In some embodiments, the first determining module 710 is configured to use a first machine learning model to perform semantic understanding on the text of the story and determine plots describing at least one of the scenes and atmosphere as target plots for generating a video.
[0157] In some embodiments, the second determining module 720 is configured to use a first machine learning model to perform semantic understanding on the target plot and the context of the target plot, determine multiple elements corresponding to the target plot, wherein the multiple elements belong to at least one type of character, item, scene, environment; determine whether the features of each of the multiple elements change in the target plot; determine elements whose features remain partially or completely unchanged as key elements, and determine the unchanging features corresponding to the key elements.
[0158] In some embodiments, the key element includes key elements of at least one type of character type or item type. The second determining module 720 is configured to determine the appearance features of the key element based on the context of the target plot and by using an association function, in response to the absence of appearance description text corresponding to the key element in the text of the target plot, and to treat the appearance features as invariant features corresponding to the key element.
[0159] In some embodiments, the key elements include key elements of at least one type in the scene and environment, and the second determining module 720 is configured to generate features of fixed items corresponding to the key elements based on the target plot and the context of the target plot, and using an association function, and determine the features of the fixed items as the unchanging features corresponding to the key elements.
[0160] In some embodiments, the generation module 730 is configured to generate storyboard text based on the text of the target plot, wherein the storyboard text includes a picture description text and a shooting method description text for each of the multiple shots; generate multiple images based on information of invariant features and the storyboard text; and generate a video corresponding to the target plot based on the multiple images.
[0161] In some embodiments, the generation module 730 is configured to use a first machine learning model to divide the text of the target plot into text corresponding to each shot; determine multiple elements in each shot and feature description information of each element based on the text corresponding to each shot; generate a picture description text for each shot based on the feature description information of the multiple elements in each shot; determine the main elements in each shot; and generate a shooting method description text for each shot based on the main elements in each shot and the content of the target plot.
[0162] In some embodiments, the generation module 730 is configured to identify multiple elements in each shot based on the text corresponding to each shot, wherein the multiple elements belong to at least one type: person, object, scene, environment; determine whether the text corresponding to the shot containing the person type element includes descriptive information of specific characteristics of the person, wherein the specific characteristics of the person include at least one of action, posture, and expression; and in response to the absence of descriptive information of specific characteristics of the person in the text corresponding to the shot containing the person type element, generate descriptive information of specific characteristics of the person corresponding to the person type element using an associative function based on the context.
[0163] In some embodiments, the generation module 730 is configured to determine whether the text corresponding to each shot includes descriptive information of an item type element; for text that does not include descriptive information of an item type element, the generation module 730 uses an association function based on the text and its context to determine the item type element corresponding to the text and the feature description information of the item type element.
[0164] In some embodiments, the information of the invariant features includes descriptive text of the invariant features and / or images corresponding to the invariant features. The generation module 730 is configured to generate prompt information corresponding to each shot based on the image description text and shooting method description text of each shot and the information of the invariant features; and input the prompt information corresponding to each shot into the second machine learning model to obtain the generated image of each shot.
[0165] In some embodiments, the generation module 730 is configured to generate narration text based on the text of the target plot; generate narration audio based on the narration text; and generate a video corresponding to the target plot based on multiple images and narration audio.
[0166] In some embodiments, the generation module 730 is configured to perform semantic understanding on the text of the target plot, generate colloquial text describing the target plot from a preset perspective, determine the key parts of the target plot based on the semantic information of the text of the target plot, and expand the text corresponding to the key parts of the colloquial text describing the target plot from a preset perspective to generate explanatory text.
[0167] In some embodiments, the generation module 730 is configured to perform semantic understanding on the text of the target plot, determine the descriptive text of the sound material, wherein the sound material includes at least one of ambient sound effects, emotional sound effects, action sound effects, event sound effects and background music; generate sound material based on the descriptive text of the sound material; and generate a video corresponding to the target plot based on multiple images and sound material.
[0168] In some embodiments, the display module 740 is configured to display a video playback control in a display interface corresponding to the target plot, wherein the display interface corresponding to the target plot includes a reading interface for the target plot and / or an audio playback interface corresponding to the target plot; and in response to the triggering of the playback control, the video corresponding to the target plot is played.
[0169] Figure 8 shows a block diagram of an electronic device according to some embodiments of the present disclosure.
[0170] Memory 81 is used to store one or more computer-readable instructions. Memory 81 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 81 may, for example, store operating systems, application programs, boot loaders, databases, and other programs, as well as various application programs and various data.
[0171] The processor 82 is configured to execute computer-readable instructions to implement the content generation method or the method described in any of the foregoing embodiments. Specific implementations of each step of the method can be found in the above embodiments; repeated details will not be elaborated here.
[0172] Processor 82 can be configured to perform the steps shown in Figure 1. Processor 82 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be based on x86 or ARM architectures, etc.
[0173] The processor 82 and the memory 81 can communicate with each other directly or indirectly. For example, the processor 82 and the memory 81 can communicate via a network. The network can include wireless networks, wired networks, and / or any combination of wireless and wired networks. The processor 82 and the memory 81 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0174] It should be noted that the components of the electronic device 8 shown in Figure 8 are exemplary and not limiting. The electronic device 8 may have other components depending on the specific application requirements. The processor 82 can control other components in the electronic device 8 to perform the desired functions.
[0175] Electronic device 8 can be implemented by software, firmware and / or hardware, and can be integrated into a device with the relevant application installed.
[0176] Figure 9 shows a block diagram of an electronic device according to some other embodiments of the present disclosure.
[0177] The electronic device 9 shown in Figure 9 can be a computer system with a dedicated hardware structure, capable of performing corresponding functions when relevant applications are installed.
[0178] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet computers (PCs), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.
[0179] As shown in Figure 9, the Central Processing Unit (CPU) 91 executes various processes based on programs stored in Read-Only Memory (ROM) 92 or programs loaded from Storage Section 98 into Random Access Memory (RAM) 93. RAM 93 stores data required as needed when the CPU 91 executes various processes. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. ROM 92, RAM 93, and Storage Section 98 can be various forms of computer-readable storage media. It should be noted that although ROM 92, RAM 93, and Storage Section 98 are shown separately in Figure 9, one or more of them can be combined or located in the same or different memories or storage modules.
[0180] CPU 91, ROM 92 and RAM 93 are interconnected via bus 94. Input / output interface 95 is also connected to bus 94.
[0181] The following components are connected to the input / output interface 95: input section 96, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 97, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 98, including hard disk, magnetic tape, etc.; and communication section 99, including network interface cards such as LAN cards, modems, etc. The communication section 99 allows communication processing to be performed via a network such as the Internet. It is readily understood that although some parts of the electronic device 9 shown in Figure 9 communicate via bus 94, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.
[0182] As needed, drive 910 is also connected to input / output interface 95. Removable media 911, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 910 as needed, so that computer programs read from them can be installed into storage section 98 as needed.
[0183] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 911.
[0184] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to perform the methods described in any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer instructions can be downloaded and installed from a network via communication section 99, or installed from storage section 98, or installed from ROM 92. When the computer program is executed by CPU 91, the methods of embodiments of this disclosure are performed.
[0185] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0186] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.
[0187] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. Computer instructions are stored on the computer-readable storage medium that, when executed by a processor, implement the methods described in any of the foregoing embodiments.
[0188] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0189] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0190] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the methods described in any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.
[0191] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0192] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0193] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0194] According to some embodiments of this disclosure, a content generation method is provided, comprising: performing semantic understanding on the text of a story to determine a target plot in the story for generating a video; performing semantic understanding on the target plot and its context to determine key elements and invariant features corresponding to the key elements; generating a video corresponding to the target plot based on information about the invariant features and the target plot; and displaying the video corresponding to the target plot.
[0195] In some embodiments, semantic understanding of the story text to determine the target plot in the story for generating a video includes: using a first machine learning model to perform semantic understanding of the story text to determine one or more characters in the story and plots related to one or more characters; performing semantic understanding of the plots related to one or more characters to determine key events corresponding to each of the one or more characters during the development of the plot; and determining the plots corresponding to the key events as the target plots for generating the video.
[0196] In some embodiments, semantic understanding of the text of a story to determine target plots in the story for generating a video includes: using a first machine learning model to perform semantic understanding of the text of the story to determine at least one plot, either a plot that drives the plot forward or a flashback plot, as a target plot for generating a video.
[0197] In some embodiments, semantic understanding of the text of a story to determine target plots in the story for generating a video includes: using a first machine learning model to perform semantic understanding of the text of the story to determine plots describing at least one of a scene and atmosphere as target plots for generating a video.
[0198] In some embodiments, semantic understanding of the target plot and its context, and determination of key elements and their corresponding invariant features, includes: using a first machine learning model to perform semantic understanding of the target plot and its context, determining multiple elements corresponding to the target plot, wherein the multiple elements belong to at least one type: character, item, scene, environment; determining whether the features of each of the multiple elements change in the target plot; identifying elements whose features remain partially or entirely unchanged as key elements, and determining the invariant features corresponding to the key elements.
[0199] In some embodiments, the key element includes key elements of at least one type of character type or item type, and determining the invariant features corresponding to the key element includes: in response to the absence of appearance description text corresponding to the key element in the text of the target plot, determining the appearance features of the key element based on the context of the target plot and using an association function, and using the appearance features as the invariant features corresponding to the key element.
[0200] In some embodiments, key elements include key elements of at least one type in the scene and environment, and determining the invariant features corresponding to the key elements includes: generating features of fixed items corresponding to the key elements based on the target plot and the context of the target plot, and using an association function, and determining the features of the fixed items as the invariant features corresponding to the key elements.
[0201] In some embodiments, generating a video corresponding to a target plot based on information of invariant features and the target plot includes: generating shot description text based on the text of the target plot, wherein the shot description text includes a scene description text and a shooting method description text for each of the multiple shots; generating multiple images based on information of invariant features and the shot description text; and generating a video corresponding to the target plot based on the multiple images.
[0202] In some embodiments, generating shot description text based on the text of the target plot includes: using a first machine learning model to divide the text of the target plot into text corresponding to each shot; determining multiple elements in each shot and feature description information of each element based on the text corresponding to each shot; generating a shot description text based on the feature description information of the multiple elements in each shot; determining the main elements in each shot; and generating a shooting method description text for each shot based on the main elements in each shot and the content of the target plot.
[0203] In some embodiments, determining multiple elements in each shot and feature description information of each element based on the text corresponding to each shot includes: identifying multiple elements in each shot based on the text corresponding to each shot, wherein the multiple elements belong to at least one type: person, object, scene, environment; determining whether the text corresponding to a shot that includes a person type element includes description information of specific characteristics of a person, wherein the specific characteristics of a person include at least one of action, posture, and expression; and generating description information of specific characteristics of a person corresponding to an element of a person type based on context using an associative function, in response to the absence of description information of specific characteristics of a person in the text corresponding to a shot that includes a person type element.
[0204] In some embodiments, determining multiple elements in each shot and the feature description information of each element based on the text corresponding to each shot includes: determining whether the text corresponding to each shot includes description information of an item type element; for text that does not include description information of an item type element, determining the item type element corresponding to the text and the feature description information of the item type element based on the text and its context using an association function.
[0205] In some embodiments, the information of the invariant features includes descriptive text of the invariant features and / or images corresponding to the invariant features. Generating multiple images based on the information of the invariant features and the shot description text includes: generating prompt information corresponding to each shot based on the image description text and shooting method description text of each shot and the information of the invariant features; inputting the prompt information corresponding to each shot into a second machine learning model to obtain the generated image of each shot.
[0206] In some embodiments, generating a video corresponding to a target plot based on multiple images includes: generating narration text based on the text of the target plot; generating narration audio based on the narration text; and generating a video corresponding to the target plot based on multiple images and narration audio.
[0207] In some embodiments, generating explanatory text based on the text of the target plot includes: performing semantic understanding on the text of the target plot; generating colloquial text describing the target plot from a preset perspective based on the text of the target plot; determining key parts of the target plot based on the semantic information of the text of the target plot; and expanding the text corresponding to the key parts of the colloquial text describing the target plot from a preset perspective to generate explanatory text.
[0208] In some embodiments, generating a video corresponding to a target plot based on multiple images includes: performing semantic understanding on the text of the target plot to determine the descriptive text of sound material, wherein the sound material includes at least one of environmental sound effects, emotional sound effects, action sound effects, event sound effects, and background music; generating sound material based on the descriptive text of the sound material; and generating a video corresponding to the target plot based on multiple images and sound material.
[0209] In some embodiments, displaying the video corresponding to the target plot includes: displaying a video playback control in a display interface corresponding to the target plot, wherein the display interface corresponding to the target plot includes a reading interface for the target plot and / or an audio playback interface corresponding to the target plot; and playing the video corresponding to the target plot in response to the triggering of the playback control.
[0210] According to some other embodiments of this disclosure, a content generation apparatus is provided, comprising: a first determining module configured to perform semantic understanding on the text of a story to determine a target plot in the story for generating a video; a second determining module configured to perform semantic understanding on the target plot and the context of the target plot to determine key elements and invariant features corresponding to the key elements; a generation module configured to generate a video corresponding to the target plot based on information of the invariant features and the target plot; and a display module configured to display the video corresponding to the target plot.
[0211] According to further embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute a content generation method as described in any embodiment of the present disclosure based on instructions stored in the memory.
[0212] According to further embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the content generation method of any embodiment of the present disclosure.
[0213] According to some other embodiments of the present disclosure, a computer program is provided, comprising: instructions that, when executed by a processor, cause the processor to perform a content generation method of any embodiment of the present disclosure.
[0214] According to further embodiments of the present disclosure, a computer program product is provided, including instructions that, when executed by a processor, implement the content generation method of any embodiment of the present disclosure.
[0215] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A content generation method, comprising: Perform semantic understanding on the text of the story to determine the target plot in the story used to generate the video; Semantic understanding is performed on the target plot and its context to determine key elements and the invariant features corresponding to those key elements; Based on the information of the invariant features and the target plot, generate a video corresponding to the target plot; Display the video corresponding to the target plot.
2. The content generation method according to claim 1, wherein, The semantic understanding of the story text to determine the target plot in the story used to generate the video includes: Using a first machine learning model, semantic understanding is performed on the text of the story to identify one or more characters in the story and the plot related to the one or more characters; Semantic understanding of the plot related to the one or more characters is performed to determine the key events corresponding to each of the one or more characters in the plot development process; The plot corresponding to the key event is determined as the target plot for generating the video.
3. The content generation method according to claim 1 or 2, wherein, The semantic understanding of the story text to determine the target plot in the story used to generate the video includes: Using a first machine learning model, semantic understanding is performed on the text of the story to identify at least one plot point, either a plot that drives the plot or a flashback plot, as a target plot point for generating a video.
4. The content generation method according to any one of claims 1-3, wherein, The semantic understanding of the story text to determine the target plot in the story used to generate the video includes: Using a first machine learning model, semantic understanding is performed on the text of the story to determine plot elements that describe at least one of the scene and atmosphere, as target plot elements for generating the video.
5. The content generation method according to any one of claims 1-4, wherein, The semantic understanding of the target plot and its context, and the determination of key elements and their corresponding invariant features, includes: Using a first machine learning model, semantic understanding is performed on the target plot and the context of the target plot to determine multiple elements corresponding to the target plot, wherein the multiple elements belong to at least one type: character, item, scene, environment; Determine whether the characteristics of each of the plurality of elements change in the target plot; Elements whose features remain partially or completely unchanged are identified as key elements, and the features that remain unchanged corresponding to the key elements are also identified.
6. The content generation method according to claim 5, wherein, The key elements include key elements of at least one of the character type or item type, and the feature that remains unchanged corresponding to the key element includes: In response to the absence of a description of the appearance of the key element in the text of the target plot, the appearance features of the key element are determined based on the context of the target plot and using an association function, and the appearance features are used as the unchanging features corresponding to the key element.
7. The content generation method according to claim 5 or 6, wherein, The key elements include at least one type of key element in the scene and environment, and the invariant features corresponding to the key elements include: Based on the target plot and its context, and using the association function, the features of the fixed items corresponding to the key elements are generated, and the features of the fixed items are determined as the unchanging features corresponding to the key elements.
8. The content generation method according to any one of claims 1-7, wherein, The step of generating the video corresponding to the target plot based on the information of the invariant features and the target plot includes: Based on the text of the target plot, a storyboard text is generated, wherein the storyboard text includes a visual description text and a shooting method description text for each of the multiple shots; The multiple images are generated based on the information of the invariant features and the storyboard text; Based on the multiple images, a video corresponding to the target plot is generated.
9. The content generation method according to claim 8, wherein, The step of generating storyboard text based on the text of the target plot includes: Using a first machine learning model, the text of the target plot is divided into text corresponding to each shot; Based on the text corresponding to each shot, determine multiple elements in each shot and feature description information of each of the multiple elements; Based on the feature description information of multiple elements in each shot, generate a scene description text for each shot; Identify the main elements in each of the shots; Based on the main elements in each shot and the content of the target plot, generate a description text of the shooting method for each shot.
10. The content generation method according to claim 9, wherein, The step of determining multiple elements in each shot and the feature description information of each element based on the text corresponding to each shot includes: Based on the text corresponding to each shot, identify multiple elements in each shot, wherein the multiple elements belong to at least one type: person, object, scene, environment; Determine whether the text corresponding to the shot containing elements of a person type includes descriptive information about the person's specific characteristics, wherein the specific characteristics of the person include at least one of action, posture, and expression; In response to the fact that the text corresponding to the shot containing the element of the character type does not contain descriptive information about the specific characteristics of the character, the descriptive information about the specific characteristics of the character corresponding to the element of the character type is generated based on the context using an associative function.
11. The content generation method according to claim 9 or 10, wherein, The step of determining multiple elements in each shot and the feature description information of each element based on the text corresponding to each shot includes: Determine whether the text corresponding to each shot includes descriptive information about the item type; For text that does not include descriptive information about item types, based on the text and its context, an association function is used to determine the item type element corresponding to the text and the feature description information of the item type element.
12. The content generation method according to any one of claims 8-11, wherein, The information about the invariant features includes descriptive text of the invariant features and / or images corresponding to the invariant features. Generating the plurality of images based on the information about the invariant features and the storyboard text includes: Based on the image description text and shooting method description text of each shot, as well as the information of the unchanging features, generate the prompt information corresponding to each shot; The prompt information corresponding to each shot is input into the second machine learning model to obtain the generated image of each shot.
13. The content generation method according to any one of claims 8-12, wherein generating the video corresponding to the target plot based on the plurality of images comprises: Based on the text of the target plot, generate explanatory text; Based on the narration text, generate narration audio; Based on the multiple images and the narration audio, a video corresponding to the target plot is generated.
14. The content generation method according to claim 13, wherein, The step of generating explanatory text based on the text of the target plot includes: Semantic understanding is performed on the text of the target plot, and colloquial text describing it from a preset perspective is generated based on the text of the target plot. Based on the semantic information of the target plot's text, determine the key parts of the target plot; The explanatory text is generated by expanding the text corresponding to the key parts of the colloquial text described from a preset perspective.
15. The content generation method according to any one of claims 8-14, wherein generating the video corresponding to the target plot based on the plurality of images comprises: Semantic understanding is performed on the text of the target plot to determine the descriptive text of the sound material, wherein the sound material includes at least one of environmental sound effects, emotional sound effects, action sound effects, event sound effects, and background music; The sound material is generated based on the description text of the sound material; Based on the multiple images and the audio materials, a video corresponding to the target plot is generated.
16. The content generation method according to any one of claims 1-15, wherein, The video displaying the target plot includes: The video playback controls are displayed in the display interface corresponding to the target plot, wherein the display interface corresponding to the target plot includes a reading interface for the target plot and / or an audio playback interface corresponding to the target plot; In response to the triggering of the playback control, the video corresponding to the target plot is played.
17. An electronic device comprising: Memory; as well as A processor coupled to the memory, the processor being configured to execute the content generation method as described in any one of claims 1 to 16 based on instructions stored in the memory.
18. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the content generation method of any one of claims 1 to 16.
19. A computer program product comprising: Instructions, wherein when executed by a processor, the instructions implement the content generation method of any one of claims 1 to 16.
20. A computer program comprising: Instructions, wherein when executed by a processor, the instructions implement the content generation method of any one of claims 1 to 16.