Method, device and related product for generating video based on e-book
By identifying the target content elements in the e-book and using a generative model to generate video data, the problem of low efficiency and low accuracy caused by users manually generating prompts is solved, and efficient and accurate generation of video data is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2026-04-14
AI Technical Summary
When using generative models to generate ebook video data, users need to spend a considerable amount of time coming up with prompts, resulting in low efficiency and inaccuracy in video data generation.
By determining the video content type of the first video data, selecting the corresponding e-book and identifying the target content elements within it, a generative model is used to generate the second video data, ensuring that the video effects of the second video data match those of the first video data, thus avoiding the need for users to manually input prompts.
It improves the accuracy and efficiency of video data generation, ensuring that the generated video data matches the video effects and content expected by the user.
Smart Images

Figure CN120529144B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus and related products for video generation based on e-books. Background Technology
[0002] With the development of generative models, e-book readers can utilize these models to generate video data for characters, plot points, and other content within their e-books, thus meeting their needs for in-depth e-book reading. However, due to the complexity of video data, generating it using generative models requires precise prompts to ensure the output matches the user's expectations. In practice, users often spend considerable time devising suitable prompts, leading to low efficiency in video data generation. Summary of the Invention
[0003] This disclosure provides a video generation method, apparatus, and related products based on e-books. It can generate second video data based on first video data for relevant content of an e-book, and the video effects of the second video data match the video effects of the first video data, thereby improving the accuracy and efficiency of video data generation.
[0004] In a first aspect, embodiments of this disclosure provide a video generation method based on e-books, including:
[0005] Display first video data, respond to a request to generate second video data based on video effects of the first video data, and determine the video content type of the first video data;
[0006] The method involves determining an ebook for generating the second video data, and, based on the video content type of the first video data, determining target content elements within the ebook; these target content elements are used to generate the video content of the second video data; the video content type of the second video data is the same as that of the first video data.
[0007] The second video data is generated using a generative model based on the target content elements and the first video data; the video effects of the second video data are matched with the video effects of the first video data.
[0008] Secondly, embodiments of this disclosure provide a video generation apparatus based on an e-book, comprising:
[0009] The display module is used to display first video data, and in response to a request to generate second video data based on video effects of the first video data, determine the video content type of the first video data.
[0010] A determining module is configured to determine an ebook for generating the second video data, and, based on the video content type of the first video data, determine a target content element in the ebook; the target content element is used to generate the video content of the second video data; the video content type of the second video data is the same as the video content type of the first video data.
[0011] The generation module is used to generate the second video data based on the target content elements and the first video data using a generative model; the video effects of the second video data are matched with the video effects of the first video data.
[0012] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method described in the first aspect above.
[0013] Fourthly, embodiments of this disclosure provide a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the method described in the first aspect.
[0014] Fifthly, embodiments of this disclosure provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the method described in the first aspect above.
[0015] In one or more embodiments of this disclosure, firstly, first video data is displayed; in response to a request to generate second video data based on video effects of the first video data, the video content type of the first video data is determined; then, an ebook for generating the second video data is determined, and, based on the video content type of the first video data, a target content element is determined in the ebook; the target content element is used to generate the video content of the second video data; the video content type of the second video data is the same as the video content type of the first video data; next, the second video data is generated based on the target content element and the first video data using a generative model; the video effects of the second video data are matched with the video effects of the first video data. As can be seen, through this embodiment, in the scenario of generating second video data based on first video data, when a user initiates a request to generate second video data for the video effects of the first video data, the video content type of the first video data can be determined, and the e-book used to generate the second video data can be determined. Based on the video content type of the first video data, the target content elements for generating the video content of the second video data are determined in the e-book. Then, through a generative model, the second video data is generated based on the target content elements and the first video data. Thus, the user does not need to manually input prompts into the generative model. The second video data can be generated based on the relevant content of the e-book from the first video data, and the video effects of the second video data match the video effects of the first video data, improving the accuracy of video data generation and increasing the efficiency of video data generation. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in one or more embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating an embodiment of the video generation method based on e-books provided in this disclosure;
[0018] Figure 2 This is a schematic diagram of a playback page for first video data provided in an embodiment of the present disclosure;
[0019] Figure 3 A schematic diagram illustrating the determination of an eBook according to an embodiment of this disclosure;
[0020] Figure 4 A schematic diagram illustrating the determination of an eBook according to another embodiment of this disclosure;
[0021] Figure 5This is a schematic diagram of one frame of the second video data provided in an embodiment of the present disclosure;
[0022] Figure 6 A schematic diagram of the structure of an e-book-based video generation device provided in an embodiment of this disclosure;
[0023] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0024] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this disclosure, the technical solutions in one or more embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of the embodiments. Based on one or more embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this disclosure.
[0025] It is understood that before using the technical solutions disclosed in the embodiments of this disclosure, relevant parties should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and authorization from the relevant parties should be obtained.
[0026] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0027] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0028] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0029] This disclosure provides a method, apparatus, and related products for generating videos based on e-books. It can generate second video data based on first video data for relevant content in an e-book, and the video effects of the second video data match those of the first video data, improving the accuracy and efficiency of video data generation. The e-book-based video generation method can be applied to and implemented on terminal devices, including but not limited to laptops, tablets, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smartphones, smart speakers, smartwatches, smart TVs, in-vehicle terminals, and various other types of user terminals.
[0030] Figure 1 This is a flowchart illustrating an embodiment of the video generation method based on e-books provided in this disclosure, as shown below. Figure 1 As shown, the process includes:
[0031] Step S102: Display the first video data, and in response to the request to generate the second video data based on the video effects of the first video data, determine the video content type of the first video data;
[0032] Step S104: Determine the ebook used to generate the second video data, and, based on the video content type of the first video data, determine the target content element in the ebook; the target content element is used to generate the video content of the second video data; the video content type of the second video data is the same as the video content type of the first video data;
[0033] Step S106: Generate second video data based on the target content elements and the first video data using a generative model; the video effects of the second video data are matched with the video effects of the first video data.
[0034] In one or more embodiments of this disclosure, firstly, first video data is displayed; in response to a request to generate second video data based on video effects of the first video data, the video content type of the first video data is determined; then, an ebook for generating the second video data is determined, and, based on the video content type of the first video data, a target content element is determined in the ebook; the target content element is used to generate the video content of the second video data; the video content type of the second video data is the same as the video content type of the first video data; next, the second video data is generated based on the target content element and the first video data using a generative model; the video effects of the second video data are matched with the video effects of the first video data. As can be seen, through this embodiment, in the scenario of generating second video data based on first video data, when a user initiates a request to generate second video data for the video effects of the first video data, the video content type of the first video data can be determined, and the e-book used to generate the second video data can be determined. Based on the video content type of the first video data, the target content elements for generating the video content of the second video data are determined in the e-book. Then, through a generative model, the second video data is generated based on the target content elements and the first video data. Thus, the user does not need to manually input prompts into the generative model. The second video data can be generated based on the relevant content of the e-book from the first video data, and the video effects of the second video data match the video effects of the first video data, improving the accuracy of video data generation and increasing the efficiency of video data generation.
[0035] In step S102 above, the first video data is displayed. The first video data can be video data with video effects, or it can be a GIF (Graphics Interchange Format) animation with video effects. The length of the first video data can be from tens of seconds to several seconds, and this embodiment does not limit it. Video effects are ways to modify, process, and transform video images through technical means to enhance visual expressiveness, create a specific atmosphere, or achieve creative expression. Video effects include transition effects, color grading and stylization effects, and visual effects. Transition effects can be exemplified by fade-in / fade-out effects, occlusion transition effects, scaling or rotation transition effects, etc. Color grading and stylization effects can be exemplified by filter effects, artistic painting effects, etc. Visual effects can be exemplified by smoke effects, fire effects, magic effects, character effects, object effects, etc.
[0036] In step S102 above, in response to a request to generate second video data based on video effects from the first video data, the video content type of the first video data is determined. The terminal device can detect the user's request to generate second video data based on video effects from the first video data and determine the video content type of the first video data according to this request.
[0037] When determining the video content type of the first video data, a video content understanding model trained using AI (Artificial Intelligence) technology can be used to understand the content of the first video data, thereby determining its video content type. Video content type refers to the type of content contained in the video data, which includes, but is not limited to, character-based, scene-based, and plot-based content. When the content contained in the first video data is a character, the video content type is character-based; examples of characters include male protagonists and female protagonists. When the content contained in the first video data is a scene, the video content type is scene-based; examples of scenes include family gatherings and office scenes. When the content contained in the first video data is a plot, the video content type is plot-based; examples of plots include a protagonist overcoming difficulties to achieve self-transcendence or a protagonist winning a basketball championship. The difference between scene and plot can be understood as follows: a scene shows exciting moments and is relatively short, such as 1-2 seconds, while a plot shows a more complete story, such as the main process from the male protagonist participating in the basketball game to winning the game, and is usually longer, such as 5-10 seconds.
[0038] In one example, the terminal device can determine the video content type of the first video data. In another example, the terminal device can send a notification message to the server, informing the server to determine the video content type of the first video data. The server can then send the determination result back to the terminal device.
[0039] Figure 2 This is a schematic diagram of a playback page for first video data provided in an embodiment of the present disclosure, as shown below. Figure 2 As shown, the playback page displays one frame from the first video data, which includes a girl, grass, and colorful bubbles. In the first video data, the girl is running on the grass, accompanied by colorful bubbles. The video effects in the first video data are bubble effects and running effects. The visual characteristics of the bubble effect are "colorful bubbles, round and spherical in shape, varying in size, floating or rising, accompanied by flickering light and shadow." It is understandable that... Figure 2 In the image shown, the girl and the bubble effect are presented statically, but when the first video data is actually played, the girl in the first video data is running, and the bubble effect is presented dynamically. Figure 2 The playback page shown displays a component named "Cut Same Style" at the bottom. Users can trigger this component to send a request to the terminal device to generate second video data. After detecting that the user has triggered the "Cut Same Style" component, the terminal device determines the video content type of the first video data. Figure 2In the first video data, the video content type is the role category.
[0040] After determining the video content type of the first video data, in step S104 above, the e-book used to generate the second video data is determined.
[0041] In one embodiment, determining the ebook used to generate the second video data includes:
[0042] The user-specified ebook is selected as the ebook used to generate the second video data.
[0043] In this embodiment, the method for determining the e-book used to generate the second video data can be to first determine the e-book specified by the user, and then determine the e-book specified by the user as the e-book used to generate the second video data. The e-book specified by the user can be the e-book corresponding to the book title entered by the user, or it can be an e-book selected by the user from multiple e-books.
[0044] Figure 3 This is a schematic diagram illustrating the determination of an e-book according to an embodiment of the present disclosure, such as... Figure 3 As shown, when the user clicks the above Figure 2 After using the "Cut the Same" component, the terminal device displays a pop-up window specifying the ebook. This page includes an input box for the ebook title, allowing the user to enter the name of the ebook used to generate the second video data. For example, if the user enters "AAAA," then the ebook "AAAA" will be the user-specified ebook, and thus, "AAAA" can be selected as the ebook for generating the second video data.
[0045] In another example, when the user clicks the above Figure 2 After using the "Cut the Same" component, the terminal device displays an e-book selection page. This page shows multiple e-books, and the user can choose one to use as the source for generating the second video data. Once the terminal device detects this user selection, it uses that e-book as the source for generating the second video data.
[0046] As can be seen, by using this embodiment to determine the e-book specified by the user as the e-book for generating the second video data, the user's video generation needs can be accurately matched, thus improving the accuracy of video generation.
[0047] In another embodiment, determining the ebook used to generate the second video data includes:
[0048] Obtain the element information input by the user, and determine the content element represented by the element information as the target content element;
[0049] E-books containing the target content elements are identified as e-books for generating the second video data.
[0050] In this embodiment, the element information input by the user can be obtained first. This element information is the element information of the content elements in the e-book. When the content element is a character, the element information includes the character name. When the content element is a scene, the element information includes the scene name. When the content element is a plot, the element information includes the plot name. Then, the content element represented by the element information input by the user is determined as the target content element. Next, the e-book including the target content element is determined as the e-book used to generate the second video data.
[0051] Figure 4 A schematic diagram illustrating the determination of an e-book as provided in another embodiment of this disclosure, such as... Figure 4 As shown, when the user clicks the above Figure 2 After using the "Cut the Same Style" component, the terminal device displays a specified page of the e-book via a pop-up window. This page includes input boxes for character names, scene names, and plot names. Users can input information into any one or more of these boxes. The content element represented by the user's input is then identified as the target content element, and the e-book containing this target content element is designated as the e-book used to generate the second video data. For example, if the user enters "Xiao A" in the character name input box, Xiao A is the name of the female protagonist in e-book 1. Therefore, the female protagonist Xiao A is the target content element, and e-book 1 containing the female protagonist Xiao A is the e-book used to generate the second video data.
[0052] As can be seen, through this embodiment, by obtaining the element information input by the user, determining the content element represented by the element information as the target content element, and determining the e-book including the target content element as the e-book for generating the second video data, it is possible to anchor the e-book with the content element specified by the user, ensuring that the generated video data is strongly correlated with the content element specified by the user, accurately matching the user's video generation needs, and improving the accuracy of video generation.
[0053] In step S104 above, a target content element is determined in the e-book based on the video content type of the first video data. This target content element is used to generate the video content of the second video data, and the video content type of the second video data is the same as that of the first video data. In this embodiment, the target content element matches the video content type of the first video data. When the video content type of the first video data is a character, the target content element can be a character in the e-book, such as the male protagonist or female protagonist. After determining the target content element, the video content of the second video data can be generated based on the target content element. For example, when the target content element is the male protagonist in the e-book, the video content of the second video data includes the male protagonist in the e-book. The video content type of the generated second video data is the same as that of the first video data. That is, when the video content type of the first video data is a character, the video content type of the second video data is also a character; when the video content type of the first video data is a scene, the video content type of the second video data is also a scene; and when the video content type of the first video data is a plot, the video content type of the second video data is also a plot.
[0054] In one example, the terminal device can determine the target content element in the e-book based on the video content type of the first video data. In another example, the terminal device can send a notification message to the server, instructing the server to determine the target content element in the e-book based on the video content type of the first video data. The server can then send the element information of the target content element to the terminal device, thereby enabling the terminal device to determine the target content element.
[0055] In one embodiment, determining target content elements in the e-book based on the video content type of the first video data includes:
[0056] Search the ebook for candidate content elements that match the video content type of the first video data, and display the element information of the found candidate content elements;
[0057] In response to the selection instruction for feature information, the candidate content features corresponding to the selected feature information are identified as the target content features.
[0058] In this embodiment, when determining target content elements in the e-book based on the video content type of the first video data, candidate content elements matching the video content type of the first video data can be searched in the e-book first, and the element information of the found candidate content elements can be displayed. In one example, the terminal device can search for candidate content elements matching the video content type of the first video data in the e-book and display the element information of the found candidate content elements. In another example, the terminal device can also send a notification message to the server, notifying the server to search for candidate content elements matching the video content type of the first video data in the e-book. The server can then send the element information of the found candidate content elements to the terminal device, thereby allowing the terminal device to display the element information of the found candidate content elements.
[0059] Then, the terminal device can detect the user's selection operation on the element information, generate a selection instruction on the element information based on this operation, and then, in response to the selection instruction, determine the candidate content element corresponding to the selected element information as the target content element.
[0060] In one example, the video content type of the first video data is, for example, a "role" category. Correspondingly, the candidate content elements are "roles," and the element information for each candidate content element is its name. For instance, the candidate content elements found in eBook 1 that match the video content type of the first video data include: female lead, male lead, female supporting character, and male supporting character. The element information for the female lead is "A," the element information for the male lead is "B," the element information for the female supporting character is "C," and the element information for the male supporting character is "D." Assuming the user selects "A" from the above element information, then the candidate content element corresponding to "A," i.e., the female lead in eBook 1, can be identified as the target content element.
[0061] As can be seen, this embodiment improves content matching efficiency by finding candidate content elements in the ebook that match the video content type of the first video data and displaying the element information of the found candidate content elements. This allows filtering of ebook content based on the video content type of the first video data. By responding to selection instructions for element information, the candidate content elements corresponding to the selected element information are determined as target content elements. This ensures that the generated video data is strongly correlated with the user-specified content elements, accurately matching the user's video generation needs.
[0062] In one embodiment, searching for candidate content elements in the ebook that match the video content type of the first video data includes:
[0063] When the video content type of the first video data is a character type, at least one target character is searched in the e-book as a candidate content element;
[0064] When the video content type of the first video data is scene-based, at least one target scene is searched in the e-book as a candidate content element.
[0065] When the video content type of the first video data is plot-based, at least one target plot is searched in the e-book as a candidate content element.
[0066] In this embodiment, when searching for candidate content elements in the e-book that match the video content type of the first video data, if the video content type of the first video data is a character type, at least one target character matching this video content type is searched in the e-book as a candidate content element. If the video content type of the first video data is a scene type, at least one target scene matching this video content type is searched in the e-book as a candidate content element. If the video content type of the first video data is a plot type, at least one target plot matching this video content type is searched in the e-book as a candidate content element.
[0067] The target character can be any character in the ebook, or a popular character that has received positive reviews from users. The target scene can be any scene in the ebook, or a highlight scene that has received positive reviews from users. The target plot can be any plot in the ebook, or a highlight scene that has received positive reviews from users.
[0068] In one example, the ebook is ebook 1. When the video content type of the first video data is a character type, at least one target character found in ebook 1 that matches the video content type of character type is the female lead A, the male lead B, the female supporting character C, and the male supporting character D. Then, the female lead A, the male lead B, the female supporting character C, and the male supporting character D in ebook 1 can be used as candidate content elements.
[0069] When the video content type of the first video data is scene type, and at least one target scene that matches the scene type in e-book 1 is a family gathering scene or an office scene, then the family gathering scene or office scene in e-book 1 can be used as candidate content elements.
[0070] When the video content type of the first video data is plot-based, if at least one target plot in eBook 1 that matches the plot-based video content type is the plot of "the female protagonist winning the dance competition championship" or the plot of "the male protagonist studying hard and getting into university", then the plots of "the female protagonist winning the dance competition championship" and "the male protagonist studying hard and getting into university" in eBook 1 can be used as candidate content elements.
[0071] As can be seen, through this embodiment, when the video content type of the first video data is a character type, at least one target character is searched in the e-book as a candidate content element; when the video content type of the first video data is a scene type, at least one target scene is searched in the e-book as a candidate content element; and when the video content type of the first video data is a plot type, at least one target plot is searched in the e-book as a candidate content element. This allows filtering of e-book content based on the video content type of the first video data, thereby obtaining candidate content elements that match the video content type of the first video data, thus improving content matching efficiency.
[0072] After identifying the target content elements in the e-book, in step S106 above, a second video data is generated based on the target content elements and the first video data using a generative model. The video effects of the second video data are matched with the video effects of the first video data. The generative model may include a video generation model trained using AI technology and a Large Language Model (LLM) trained using AI technology. The video generation model can generate second video data whose video effects match those of the first video data. Matching the video effects of the second video data with those of the first video data can mean that the video effects of the second video data are identical to those of the first video data.
[0073] In one example, the terminal device can generate second video data based on target content elements and first video data using a generative model. In another example, the terminal device can send a notification message to the server, instructing the server to generate second video data based on the target content elements and first video data using a generative model. The server can then send the second video data to the terminal device, which will then display the second video data.
[0074] In one embodiment, a second video data is generated based on target content elements and first video data using a generative model, including:
[0075] Based on the target content elements, generate content prompt information corresponding to the second video data;
[0076] Based on the video effects of the first video data, generate the corresponding effect prompt information for the second video data;
[0077] A second video data is generated using a generative model based on content prompts and effect prompts; the content prompts and effect prompts are used to guide the generative model in generating the second video data.
[0078] In this embodiment, when generating second video data based on target content elements and first video data using a generative model, content prompts corresponding to the second video data can be generated first, based on the target content elements. These content prompts can be used to guide the generative model in generating the content of the second video data. Then, based on the video effects of the first video data, effect prompts corresponding to the second video data are generated. These effect prompts can be used to guide the generative model in generating the video effects of the second video data. Next, the content prompts and effect prompts are input into the generative model. After the generative model performs content understanding on the content prompts and effect prompts, the second video data is generated.
[0079] In one example, the ebook is ebook 1, and the target content element is the female protagonist, Xiao A, in ebook 1. Based on Xiao A in ebook 1, the content prompt information corresponding to the generated second video data is, for example: Xiao A has a high ponytail, with pearl hair clips adorning her fluffy hair. Her bright almond-shaped eyes shine beneath her smooth forehead, and she shouts in a clear and lively voice, "Go! We can definitely win!" The video effects of the first video data are running and bubble effects. Based on these running and bubble effects, the effect prompt information corresponding to the second video data is: "A person is running on the grass, surrounded by colorful bubbles. The bubbles are round and spherical, varying in size, floating or rising, accompanied by flickering light and shadow." The above content prompt information and effect prompt information are input into a generative model, which generates the second video data. The content element contained in the second video data is the female protagonist, Xiao A, from ebook 1. The video effects of the second video data are the same as those of the first video data, namely running and bubble effects.
[0080] Figure 5 This is a schematic diagram of one frame of the second video data provided in an embodiment of the present disclosure. Figure 5 For based on Figure 2 The second video data is generated from the first video data shown. Figure 5 As shown, this image is a frame from the second video data, displaying the female protagonist, Xiao A, from eBook 1, and showing the video effects of the second video data. The video effects in the second video data are the same as those in the first video data, namely running and bubble effects. In the second video data, the female protagonist, Xiao A, is running on the grass, accompanied by colorful bubbles. It is understandable that... Figure 5 The running and bubble effects in the shown image are static, but when the second video data is actually played, the female protagonist, Xiao A, is running, and the bubble effects are dynamic. Furthermore, when the second video data is actually played, Xiao A shouts in a clear and lively voice, "Go! We can definitely win!" (Comparison) Figure 2 and Figure 5 It can be seen that the second video data has the same video effects as the first video data. The difference between the second video data and the first video data is that the content of the second video data is the content of the e-book determined above, thus realizing the effect of "one-click cutting the same style" of the video data of interest by the user using the content of the e-book.
[0081] In another example, the e-book is e-book 1, and the target content element is the family gathering scene in e-book 1. Based on the description of the family gathering scene in e-book 1, the content prompt information corresponding to the generated second video data is, for example: "In the cozy living room, snacks and fruits are placed on the coffee table, a family photo hangs on the wall, and warm yellow lighting creates a lively atmosphere. The father holds up his phone and shouts in a slightly hoarse but penetrating voice, 'Everyone, look at the camera.'" The video effect of the first video data is, for example, a dynamic depth-of-field switching effect, which simulates the process of a camera "finding focus" during real shooting. The effect prompt information corresponding to the second video data generated based on this dynamic depth-of-field switching effect is: "When the father holds up his phone and shouts 'Everyone, look at the camera,' the camera first focuses on the phone screen, capturing family members as blurred outlines on the screen. When the father utters the last syllable of 'camera,' the camera quickly refocuses on the main body of the crowd, and the background is instantly blurred." The above content prompts and effects prompts are input into the generative model, which generates the second video data. The content elements in the second video data are the family gathering scene in eBook 1. The video effects of the second video data are the same as those of the first video data, which is a dynamic depth-of-field switching effect.
[0082] As can be seen, through this embodiment, by generating content prompt information corresponding to the second video data based on the target content elements, and generating effect prompt information corresponding to the second video data based on the video effects of the first video data, the content prompt information and effect prompt information are used to prompt the generative model to generate the second video data. By generating the second video data through the generative model based on the content prompt information and effect prompt information, the second video data can be more matched with the target content elements and the video effects of the first video data, thereby improving the accuracy of video generation.
[0083] In one embodiment, content prompt information corresponding to the second video data is generated based on the target content elements, including:
[0084] The first text describing the visual content of the target content element is obtained from the e-book, and the second text describing the audio content of the target content element is obtained from the e-book; the audio content of the target content element includes the character's voice and lines associated with the target content element, or includes the background sound associated with the target content element.
[0085] Based on the first text and the second text, generate content prompt information corresponding to the second video data.
[0086] In this embodiment, when generating content prompt information corresponding to the second video data based on the target content element, the first text describing the screen content of the target content element can be obtained from the e-book. The screen content of the target content element may include the character image, object props, and scene arrangement associated with the target content element. For example, when the target content element is the female protagonist in e-book 1, the screen content of the target content element may include the female protagonist's appearance, expression, actions, and the scene in which the female protagonist is located. Similarly, when the target content element is a family gathering scene in e-book 1, the screen content of the target content element may include the appearance, posture, furniture, and home decoration of the characters in that family gathering scene. Furthermore, when the target content element is a highlight scene in e-book 1, the screen content of the target content element may include the appearance, posture, and background environment of the characters in that scene. The first text is the original text of the e-book.
[0087] Furthermore, a second text is retrieved from the ebook to describe the audio content of the target content element. The audio content of the target content element includes the character's voice and dialogue associated with the target content element, or it includes background sounds associated with the target content element. For example, when the target content element is the female protagonist, Xiao A, in ebook 1, the audio content of the target content element includes Xiao A's voice and dialogue, which could be Xiao A's classic lines. As another example, when the target content element is a family gathering scene in ebook 1, the audio content of the target content element includes background sounds associated with the family gathering scene, such as the sound of a television set or distant fireworks. Similarly, when the target content element is a highlight scene in ebook 1, the audio content of the target content element includes background sounds associated with the scene, such as the sound of wind or the sound of tides. The second text is the original text of the ebook.
[0088] In one example, the ebook is ebook 1, and the target content element is the plot of "the female protagonist winning the dance competition." The first text obtained from ebook 1 describing the visual content of "the female protagonist winning the dance competition" is "The red ribbon falls on the podium, and Xiao A tightly grips the gold medal." The second text obtained from ebook 1 describing the audio content of "the female protagonist winning the dance competition" is "Thunderous applause erupts from the audience, and Xiao A excitedly shouts: 'I did it!'"
[0089] After obtaining the first and second texts, content prompts corresponding to the second video data are generated based on the first and second texts. When generating the content prompts for the second video data, the first and second texts can be combined as the content prompts. Alternatively, a generative model can be used to understand the content of the first and second texts and then generate the content prompts for the second video data.
[0090] As can be seen, by obtaining the first text describing the visual content of the target content element from the e-book and the second text describing the audio content of the target content element from the e-book through this embodiment, and generating content prompt information corresponding to the second video data based on the first text and the second text, the content prompt information corresponding to the generated second video data can better match the target content element in the original text of the e-book, thereby improving the accuracy of video generation.
[0091] In one embodiment, content prompt information corresponding to the second video data is generated based on the first text and the second text, including:
[0092] A reference image is generated based on the first text; the reference image is used to represent the content of the image based on the first text.
[0093] A reference voice-over is generated based on the second text; the reference voice-over is used to represent the sound content based on the second text.
[0094] Based on the reference image and reference voiceover, generate content prompts corresponding to the second video data.
[0095] In this embodiment, when generating content prompt information corresponding to the second video data based on the first text and the second text, a reference image can first be generated based on the first text using a generative model. This reference image is used to represent the visual content of the target content element based on the first text. Then, a reference voice-over can be generated based on the second text using a generative model. This reference voice-over is used to represent the audio content of the target content element based on the second text. The reference voice-over can be audio data.
[0096] In one embodiment, one or more reference images can be generated based on the first text, and each reference image can be provided to a user, who can then select one of the images as the target reference image for the second video data. Additionally, one or more reference voice-overs can be generated based on the second text, and each reference voice-over can be provided to the user, who can then select one of the reference voice-overs as the target reference voice-over for the second video data.
[0097] Finally, a generative model is used to generate content prompts corresponding to the second video data based on the reference image and reference voice-over. For example, the user-selected reference image, the user-selected reference voice-over, the first text, and the second text are collectively used as the content prompts corresponding to the second video data. In this embodiment, the content of the video frame in the second video data is consistent with the content of the user-selected reference image; the content of the user-selected reference image is the content of the video frame in the second video data.
[0098] In one example, the first text is "The red ribbon falls from the podium, and Little A clutches the gold medal tightly." The second text is "Thunderous applause erupts from the audience, and Little A shouts excitedly: 'I did it!'" In this example, multiple reference images are generated based on the first text, and multiple reference voice-overs are generated based on the second text. The reference voice-overs are audio data. Next, the user selects a target reference image from among the reference images, and a target reference voice-over from among the reference voice-overs. The first text, the second text, the target reference image, and the target reference voice-over are then used to define the content prompt information corresponding to the second video data.
[0099] As can be seen, through this embodiment, a reference image is generated based on the first text, which is used to represent the screen content based on the first text; a reference voice-over is generated based on the second text, which is used to represent the sound content based on the second text; and content prompt information corresponding to the second video data is generated based on the reference image and the reference voice-over. This can ensure that the screen and sound of the second video data are highly consistent with the original text of the e-book, and improve the accuracy of video generation.
[0100] After generating content prompt information corresponding to the second video data, special effects prompt information corresponding to the second video data is generated based on the video effects of the first video data. In one embodiment, generating special effects prompt information corresponding to the second video data based on the video effects of the first video data includes:
[0101] Generate descriptive information for video effects that describe the first video data;
[0102] Based on the description information and the first video data, generate special effects prompts corresponding to the second video data.
[0103] In this embodiment, when generating effect prompt information corresponding to the second video data based on the video effects of the first video data, descriptive information describing the video effects of the first video data can be generated first. Then, based on the descriptive information and the first video data, the effect prompt information corresponding to the second video data is generated. Specifically, the video effects of the first video data can be summarized using a large language model to obtain descriptive information about the video effects of the first video data. Then, the descriptive information and the first video data are used together as the effect prompt information corresponding to the second video data. The effect prompt information may also carry text such as "Please refer to this description to reproduce the effects in the video data."
[0104] In one example, the first video data features a dynamic starlight effect. This effect is characterized by shimmering, fragmented light spots of varying sizes, predominantly cool-toned with warm gold accents, randomly moving and interacting with ambient light and shadow, with darker areas being more prominent and brighter areas less so. Using a large language model, the generated description of this dynamic starlight effect is, for example, "Fragmented light spots, predominantly cool-toned with warm gold accents, or drifting, flowing, and scaling, flickering with the brightness and darkness of the ambient light and shadow, with a lively rhythm." This description, along with the first video data, is then used as the effect cue information for the second video data.
[0105] As can be seen, this embodiment can generate descriptive information for video effects describing the first video data, and generate effect prompts for the second video data based on the descriptive information and the first video data, ensuring that the video effects of the generated second video data match the video data of the first video data, thereby improving the accuracy of video generation.
[0106] It should be noted that the titles, examples, special effects, prompts, and accompanying drawings in this embodiment are all illustrative examples and are not intended to limit this embodiment. The server mentioned in this embodiment can be a server or server cluster used to generate the second video data.
[0107] In summary, through this embodiment, in the scenario of generating second video data based on first video data, when a user initiates a request to generate second video data for the video effects of the first video data, the video content type of the first video data can be determined, and the e-book used to generate the second video data can be identified. Based on the video content type of the first video data, the target content elements for generating the second video data are determined in the e-book. Then, through a generative model, the second video data is generated based on the target content elements and the first video data. Thus, the user does not need to manually input prompts into the generative model; the second video data can be generated based on the relevant content of the e-book from the first video data. Furthermore, the video effects of the second video data match the video effects of the first video data, improving the accuracy and efficiency of video data generation.
[0108] Figure 6 This is a schematic diagram of the structure of a video generation device based on an e-book provided in an embodiment of the present disclosure, as shown below. Figure 6 As shown, the device includes:
[0109] Display module 601 is used to display first video data, and in response to a request to generate second video data based on video effects of the first video data, determine the video content type of the first video data.
[0110] The determining module 602 is configured to determine an e-book for generating the second video data, and, based on the video content type of the first video data, determine a target content element in the e-book; the target content element is used to generate the video content of the second video data; the video content type of the second video data is the same as the video content type of the first video data.
[0111] The generation module 603 is used to generate the second video data based on the target content elements and the first video data using a generative model; the video effects of the second video data are matched with the video effects of the first video data.
[0112] Optionally, the determining module 602 is specifically used to: determine the e-book specified by the user, and determine the e-book specified by the user as the e-book used to generate the second video data.
[0113] Optionally, the determining module 602 is specifically used to: obtain element information input by the user, determine the content element represented by the element information as the target content element; and determine the e-book including the target content element as the e-book for generating the second video data.
[0114] Optionally, the determining module 602 is specifically configured to: search for candidate content elements in the e-book that match the video content type of the first video data, and display the element information of the found candidate content elements; in response to a selection instruction for the element information, determine the candidate content element corresponding to the selected element information as the target content element.
[0115] Optionally, the determining module 602 is further configured to: when the video content type of the first video data is a character type, search for at least one target character in the e-book as the candidate content element; when the video content type of the first video data is a scene type, search for at least one target scene in the e-book as the candidate content element; when the video content type of the first video data is a plot type, search for at least one target plot in the e-book as the candidate content element.
[0116] Optionally, the generation module 603 is specifically used to: generate content prompt information corresponding to the second video data based on the target content elements; generate special effects prompt information corresponding to the second video data based on the video effects of the first video data; and generate the second video data through the generative model based on the content prompt information and the special effects prompt information; the content prompt information and the special effects prompt information are used to prompt the generative model to generate the second video data.
[0117] Optionally, the generation module 603 is further configured to: obtain a first text from the e-book describing the visual content of the target content element, and obtain a second text from the e-book describing the audio content of the target content element; the audio content of the target content element includes the character's timbre and lines associated with the target content element, or includes the background sound associated with the target content element; and generate the content prompt information corresponding to the second video data based on the first text and the second text.
[0118] Optionally, the generation module 603 is further configured to: generate a reference image based on the first text; the reference image is used to represent the content of the screen based on the first text; generate a reference voice-over based on the second text; the reference voice-over is used to represent the sound content based on the second text; and generate the content prompt information corresponding to the second video data based on the reference image and the reference voice-over.
[0119] Optionally, the generation module 603 is further configured to: generate description information for describing the video effects of the first video data; and generate the effect prompt information corresponding to the second video data based on the description information and the first video data.
[0120] In this embodiment, in a scenario where second video data is generated based on first video data, when a user requests the generation of second video data for video effects of the first video data, the video content type of the first video data can be determined, and the e-book used to generate the second video data can be identified. Based on the video content type of the first video data, the target content elements for generating the second video data are determined in the e-book. Then, through a generative model, the second video data is generated based on the target content elements and the first video data. This eliminates the need for the user to manually input prompts into the generative model. The second video data can be generated based on the relevant content of the e-book from the first video data, and the video effects of the second video data match those of the first video data, improving the accuracy and efficiency of video data generation.
[0121] The e-book-based video generation device in this embodiment can implement the various processes of the above-described e-book-based video generation method embodiment and achieve the same effects and functions, which will not be repeated here.
[0122] One embodiment of this disclosure also provides an electronic device. Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, as shown below. Figure 7 As shown, electronic devices can vary considerably due to differences in configuration or performance. They may include one or more processors 701 and memories 702, with the memory 702 storing one or more application programs or data. The memory 702 can be temporary or persistent storage. The application programs stored in the memory 702 may include one or more modules (not shown), each module including a series of computer-executable instructions from the electronic device. Furthermore, the processor 701 may be configured to communicate with the memory 702, executing the series of computer-executable instructions stored in the memory 702 on the electronic device. The electronic device may also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input or output interfaces 705, one or more keyboards 706, etc.
[0123] In one specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the following process:
[0124] Display first video data, respond to a request to generate second video data based on video effects of the first video data, and determine the video content type of the first video data;
[0125] The method involves determining an ebook for generating the second video data, and, based on the video content type of the first video data, determining target content elements within the ebook; these target content elements are used to generate the video content of the second video data; the video content type of the second video data is the same as that of the first video data.
[0126] The second video data is generated using a generative model based on the target content elements and the first video data; the video effects of the second video data are matched with the video effects of the first video data.
[0127] In this embodiment, in a scenario where second video data is generated based on first video data, when a user requests the generation of second video data for video effects of the first video data, the video content type of the first video data can be determined, and the e-book used to generate the second video data can be identified. Based on the video content type of the first video data, the target content elements for generating the second video data are determined in the e-book. Then, through a generative model, the second video data is generated based on the target content elements and the first video data. This eliminates the need for the user to manually input prompts into the generative model. The second video data can be generated based on the relevant content of the e-book from the first video data, and the video effects of the second video data match those of the first video data, improving the accuracy and efficiency of video data generation.
[0128] The electronic device in this embodiment can implement the various processes of the above-described video generation method embodiment based on e-books and achieve the same effects and functions, which will not be repeated here.
[0129] Another embodiment of this disclosure also provides a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the following process:
[0130] Display first video data, respond to a request to generate second video data based on video effects of the first video data, and determine the video content type of the first video data;
[0131] The method involves determining an ebook for generating the second video data, and, based on the video content type of the first video data, determining target content elements within the ebook; these target content elements are used to generate the video content of the second video data; the video content type of the second video data is the same as that of the first video data.
[0132] The second video data is generated using a generative model based on the target content elements and the first video data; the video effects of the second video data are matched with the video effects of the first video data.
[0133] In this embodiment, in a scenario where second video data is generated based on first video data, when a user requests the generation of second video data for video effects of the first video data, the video content type of the first video data can be determined, and the e-book used to generate the second video data can be identified. Based on the video content type of the first video data, the target content elements for generating the second video data are determined in the e-book. Then, through a generative model, the second video data is generated based on the target content elements and the first video data. This eliminates the need for the user to manually input prompts into the generative model. The second video data can be generated based on the relevant content of the e-book from the first video data, and the video effects of the second video data match those of the first video data, improving the accuracy and efficiency of video data generation.
[0134] The computer-readable storage medium in this embodiment can implement the various processes of the above-described e-book-based video generation method embodiment and achieve the same effects and functions, which will not be repeated here.
[0135] Another embodiment of this disclosure also provides a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the following process:
[0136] Display first video data, respond to a request to generate second video data based on video effects of the first video data, and determine the video content type of the first video data;
[0137] The method involves determining an ebook for generating the second video data, and, based on the video content type of the first video data, determining target content elements within the ebook; these target content elements are used to generate the video content of the second video data; the video content type of the second video data is the same as that of the first video data.
[0138] The second video data is generated using a generative model based on the target content elements and the first video data; the video effects of the second video data are matched with the video effects of the first video data.
[0139] In this embodiment, in a scenario where second video data is generated based on first video data, when a user requests the generation of second video data for video effects of the first video data, the video content type of the first video data can be determined, and the e-book used to generate the second video data can be identified. Based on the video content type of the first video data, the target content elements for generating the second video data are determined in the e-book. Then, through a generative model, the second video data is generated based on the target content elements and the first video data. This eliminates the need for the user to manually input prompts into the generative model. The second video data can be generated based on the relevant content of the e-book from the first video data, and the video effects of the second video data match those of the first video data, improving the accuracy and efficiency of video data generation.
[0140] The computer program product in this disclosure embodiment can implement the various processes of the above-described e-book-based video generation method embodiment and achieve the same effects and functions, which will not be repeated here.
[0141] In various embodiments of this disclosure, the computer-readable storage medium includes read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc.
[0142] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0143] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0144] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0145] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing the embodiments of this disclosure, the functions of each unit can be implemented in one or more software and / or hardware.
[0146] Those skilled in the art will understand that one or more embodiments of this disclosure can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0147] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0148] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0149] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0150] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0151] One or more embodiments of this disclosure can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.
[0152] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0153] The above description is merely an embodiment of this disclosure and is not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.
Claims
1. A video generation method based on e-books, characterized in that, include: Display first video data, respond to a request to generate second video data based on video effects of the first video data, and determine the video content type of the first video data; The video content types include character-based, scene-based, and plot-based. Determine the ebook used to generate the second video data, and, based on the video content type of the first video data, determine the target content elements in the ebook that match the video content type; The target content elements are used to generate the video content of the second video data; The video content type of the second video data is the same as that of the first video data; The second video data is generated using a generative model based on the target content elements and the first video data. The video effects of the second video data match the video effects of the first video data; The step of generating the second video data using a generative model based on the target content elements and the first video data includes: Based on the target content elements, generate content prompt information corresponding to the second video data; based on the video effects of the first video data, generate effect prompt information corresponding to the second video data; and generate the second video data using the generative model based on the content prompt information and the effect prompt information.
2. The method according to claim 1, characterized in that, The step of determining the e-book used to generate the second video data includes: The user-specified ebook is selected as the ebook used to generate the second video data.
3. The method according to claim 1, characterized in that, The step of determining the e-book used to generate the second video data includes: Obtain the element information input by the user, and determine the content element represented by the element information as the target content element; The ebook that includes the target content elements is identified as the ebook used to generate the second video data.
4. The method according to claim 1, characterized in that, The step of determining the target content elements in the e-book based on the video content type of the first video data includes: The system searches for candidate content elements in the e-book that match the video content type of the first video data, and displays the element information of the found candidate content elements. In response to the selection instruction for the element information, the candidate content element corresponding to the selected element information is determined as the target content element.
5. The method according to claim 4, characterized in that, The step of searching for candidate content elements in the e-book that match the video content type of the first video data includes: When the video content type of the first video data is a character type, at least one target character is searched in the e-book as the candidate content element; When the video content type of the first video data is a scene type, at least one target scene is searched in the e-book as the candidate content element; When the video content type of the first video data is plot-based, at least one target plot is searched in the e-book as the candidate content element.
6. The method according to claim 1, characterized in that, The step of generating content prompt information corresponding to the second video data based on the target content elements includes: The first text describing the visual content of the target content element is obtained from the e-book, and the second text describing the audio content of the target content element is obtained from the e-book; the audio content of the target content element includes the character's voice and lines associated with the target content element, or includes the background sound associated with the target content element. Based on the first text and the second text, generate the content prompt information corresponding to the second video data.
7. The method according to claim 6, characterized in that, The step of generating the content prompt information corresponding to the second video data based on the first text and the second text includes: A reference image is generated based on the first text; the reference image is used to represent the content of the image based on the first text. A reference voice-over is generated based on the second text; the reference voice-over is used to represent the sound content based on the second text. Based on the reference image and the reference voiceover, the content prompt information corresponding to the second video data is generated.
8. The method according to claim 1, characterized in that, The step of generating special effects prompt information corresponding to the second video data based on the video effects of the first video data includes: Generate descriptive information for video effects that describe the first video data; Based on the description information and the first video data, the special effects prompt information corresponding to the second video data is generated.
9. A video generation device based on e-books, characterized in that, include: The display module is used to display first video data, and in response to a request to generate second video data based on video effects of the first video data, to determine the video content type of the first video data; the video content type includes character type, scene type, and plot type. The determining module is configured to determine the ebook used to generate the second video data, and, based on the video content type of the first video data, determine target content elements in the ebook that match the video content type; The target content elements are used to generate the video content of the second video data; The video content type of the second video data is the same as that of the first video data; The generation module is used to generate the second video data based on the target content elements and the first video data using a generative model; the video effects of the second video data are matched with the video effects of the first video data. Specifically, the generation module is used to: generate content prompt information corresponding to the second video data based on the target content elements; and generate special effects prompt information corresponding to the second video data based on the video effects of the first video data. The second video data is generated using the generative model based on the content prompts and the special effects prompts.
10. An electronic device, characterized in that, include: processor; as well as, A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions that, when executed by a processor, implement the method described in any one of claims 1-8.
12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-8.
Citation Information
Patent Citations
Automatically generating visual content
WO2021173305A1
Video generation method and apparatus, and device, storage medium and program product
WO2024240188A1