Picture book video generation method and device based on large model, electronic equipment and medium

Through the method of generating picture book videos through a large model, the themes entered by users are obtained, multiple dynamic images are generated and stitched into picture book videos, which solves the problems of complex operation and poor user experience in the existing technology, and realizes convenient and efficient picture book video generation and vivid video effects.

CN120264098APending Publication Date: 2025-07-04BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510473797.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The operation of generating picture book videos in the prior art is complicated, and users need to switch between multiple applications. The generated videos are single in form and have poor user experience.

Method used

Through the method of generating picture book videos through a large model, the topics entered by the user are obtained, multiple dynamic images are generated, and they are stitched into picture book videos. The objects in the dynamic images change between multiple frame images, simplifying user operations and improving video vividness.

Benefits of technology

It improves the operation convenience and user experience of picture book video generation, and the generated videos are more vivid and vivid, meeting the diverse needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264098A_ABST
    Figure CN120264098A_ABST
Patent Text Reader

Abstract

The invention provides a picture book video generation method and device based on a large model, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the fields of generation models, large models and the like. According to the specific implementation scheme, input information is obtained, wherein the input information comprises a theme; generating a plurality of dynamic images according to the theme based on the large model; wherein the dynamic image comprises multiple frames of images, and objects in the dynamic image change among the multiple frames of images; and generating a picture book video according to the plurality of dynamic images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to fields such as generative models and large models. More specifically, the present disclosure provides a method, apparatus, electronic device, storage medium, and computer program product for generating picture book videos based on large models. Background Art

[0002] In the context of family education, parents sometimes play picture book videos composed of multiple static pictures for children. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, electronic device, storage medium, and computer program product for generating picture book videos based on large models.

[0004] According to one aspect of the present disclosure, there is provided a method for generating a picture book video based on a large model, including: obtaining input information, where the input information includes a theme; based on the large model, generating a plurality of dynamic images according to the theme; wherein, the dynamic images include multiple frames of images, and the objects in the dynamic images change among the multiple frames of images; generating a picture book video according to the plurality of dynamic images..

[0005] According to another aspect of the present disclosure, there is provided a device for generating a picture book video based on a large model, including: an obtaining module, a dynamic image generating module, and a picture book video generating module. The obtaining module is configured to obtain input information, where the input information includes a theme. The dynamic image generating module is configured to generate a plurality of dynamic images based on the large model according to the theme; wherein, the dynamic images include multiple frames of images, and the objects in the dynamic images change among the multiple frames of images. The picture book video generating module is configured to generate a picture book video according to the plurality of dynamic images.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method provided by the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method provided by the present disclosure.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Description of the Drawings

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a schematic diagram of an application scenario of a picture book video generation method and device according to an embodiment of the present disclosure;

[0012] Figure 2 is a schematic flowchart of a picture book video generation method according to an embodiment of the present disclosure;

[0013] Figures 3A - 3B is a schematic diagram of a front-end page of a picture book video generation method according to an embodiment of the present disclosure;

[0014] Figure 4 is a schematic structural block diagram of a picture book video generation device according to an embodiment of the present disclosure; and

[0015] Figure 5 is a structural block diagram of an electronic device for implementing the picture book video generation method according to an embodiment of the present disclosure. Detailed Embodiments

[0016] The following describes exemplary embodiments of the present disclosure with reference to the drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0017] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information and other processing all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0018] In the technical solution of the present disclosure, the authorization or consent of the user is obtained before obtaining or collecting user personal information.

[0019] In some technical solutions, a single image can be generated based on text, or an image can be processed into a video, but it does not support generating a video based on a user input theme (such as a sentence). Therefore, users need to use multiple products to generate a video. For example, a user generates a story through one application, then uses another application to generate a prompt for generating pictures, then uses another application to generate pictures based on the prompt, then uses the pictures to generate a sequence of consecutive image frames, and finally uses video editing software to splice the image frame sequence, music, and subtitles into a video. It can be seen that with this method of generating a video, the operation cost for users is relatively high and the convenience is poor. In addition, the form of the generated video is relatively single. For example, the video is actually multiple segments, and each segment plays a static image, resulting in a poor user experience.

[0020] Embodiments of the present disclosure aim to provide a method for generating a picture book video. This method can generate dynamic images based on a theme input by a user, and then generate a picture book video through the dynamic images, without the user having to operate multiple applications, which is more convenient and improves the user experience. In addition, since the pictures in the picture book video are dynamic images rather than static images, they are more vivid and further improve the user experience.

[0021] The technical solution provided in this embodiment is applicable to generating picture book videos and can also be used to generate other videos. This embodiment does not limit the type of generated videos.

[0022] The technical solution provided by the present disclosure will be elaborated in detail below in conjunction with the accompanying drawings and specific embodiments.

[0023] Figure 1 is a schematic diagram of an application scenario of a method and apparatus for generating a picture book video according to an embodiment of the present disclosure.

[0024] It should be noted that Figure 1 The example shown is only an example of a system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.

[0025] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0026] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on.

[0027] Server 105 can be a server that provides various services. For example, it can be a background management server (only an example) that supports the websites browsed by users using terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as picture book videos generated according to the theme input by the user or the options selected) to the terminal device.

[0028] It should be noted that the picture book video generation method provided by the embodiments of the present disclosure can generally be executed by server 105. Correspondingly, the picture book video generation device provided by the embodiments of the present disclosure can generally be set in server 105. The picture book video generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103, and / or server 105. Correspondingly, the picture book video generation device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103, and / or server 105.

[0029] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0030] Figure 2 are only illustrative. According to the implementation needs, there can be any number of terminal devices, networks, and servers.

[0031] As Figure 2 shown, the picture book video generation method 200 based on a large model can include operation S210 to operation S230.

[0032] In operation S210, input information is obtained, and the input information includes a theme.

[0033] For example, the user can perform operations such as input and selection through the front-end page. For example, the user inputs a text-based theme through the front-end page, and the theme is, for example, "Xiaoming plays basketball". In addition, the user can also select options such as style and video ratio. The client encapsulates the information obtained through these operations into input information and sends it to the electronic device that executes the picture book video generation method, so that the electronic device can obtain the input information.

[0034] In operation S220, based on a large model, multiple dynamic images are generated according to the theme; wherein, the dynamic images include multiple frames of images, and the objects in the dynamic images change among the multiple frames of images.

[0035] For example, the theme can be input into the large model, and the large model is guided by pre-configured prompt information to generate a moving picture based on the theme. This embodiment does not limit the prompt information.

[0036] For example, the object can include some parts of a character, and the character can include a person or an animal. For example, the object includes the facial features, limbs, etc. of a person. The object can also include an object, such as a fallen leaf, the sun, a tree branch, a vehicle, etc. Other areas in the dynamic image except the object can remain unchanged among the multiple frames of images.

[0037] In operation S230, a picture book video is generated according to the multiple dynamic images.

[0038] For example, the dynamic images can be spliced according to a predetermined transition effect to obtain a picture book video.

[0039] The picture book video generation method provided in this embodiment generates dynamic images according to the theme input by the user and then generates a picture book video. And a picture book video is generated through the images. During the process of generating the picture book video, the user does not need to operate multiple application programs, and the operation is more convenient, improving the user experience. In addition, since the pictures in the picture book video are dynamic images rather than static images, they are more vivid and further improve the user experience.

[0040] Next, the process of generating dynamic images based on the theme will be described.

[0041] In this embodiment, the user can also select a video effect through the front-end page. The video effect includes, for example, a still image effect and a moving image effect. The still image effect means that the images included in the picture book video to be generated are static images, and the moving image effect means that the images included in the picture book video to be generated are dynamic images.

[0042] During the process of generating the dynamic images, multiple static images can be first generated based on the large model according to the theme, and then, when the video effect is the moving image effect, the multiple static images are respectively converted into multiple dynamic images based on the large model.

[0043] For example, the theme can be input into the large model, and the large model generates several images related to the theme. During the generation process, the generation steps of the large model can be guided by pre-configured prompt information. This embodiment does not limit the prompt information.

[0044] For example, the ways to convert a static image into a dynamic image may include at least one of the following: adjusting the expression of a character (such as blinking, opening the mouth), adjusting the limb movements of a character (such as raising a hand), and adjusting the movement trajectory of an object.

[0045] In some embodiments, an object that can generate an action in the static image may be determined first. For example, several objects included in the static image are determined by means such as object detection and classification. Object categories that can generate actions may be preconfigured, so as to determine the objects that can generate actions from several objects based on the object categories. In addition, the maximum number of objects that can act in the dynamic image may be configured. If the number of objects that can generate actions determined based on the object categories is greater than the maximum number, screening may be performed based on priorities to avoid problems such as too many action objects in the dynamic image, resulting in an overly complex picture or poor overall coordination of multiple objects acting simultaneously. The action methods of the objects may be preconfigured. For example, when the object is an eye, the action method is blinking. Then, dynamic effects are added to the object based on the action method, so as to convert the static image into a dynamic image.

[0046] In this embodiment, a static image is first generated based on a theme, and then the static image is converted to generate a dynamic image. Compared with the method of directly generating a dynamic image based on a theme, converting through a static image can improve the quality of the dynamic image.

[0047] Next, the process of generating a static image based on a theme will be described.

[0048] In this embodiment, a story text may be first generated based on a large model according to the theme; the story text includes multiple story sub-texts. For example, the theme is input into the large model, and the large model expands and writes the theme to generate the story text. The story text includes multiple story sub-texts. The story sub-texts are, for example, a sentence or a paragraph in the story text. The number of story sub-texts may be within a predetermined range. For example, the predetermined range is 10 to 20.

[0049] Next, multiple image description texts respectively for multiple story sub-texts may be generated based on the large model according to the multiple story sub-texts. For example, the story sub-texts may be input into the large model, and the large model is guided by preconfigured prompt information to appropriately expand and write the plot of the story sub-texts and describe the content included in the plot, so as to obtain the image description texts. The image description texts can describe the semantic content included in the images. For example, the image description text is "The girl climbs up the building, the snow under her feet creaks, and she tightly holds the iron cable with both hands and climbs up".

[0050] Next, multiple static images may be generated according to the multiple image description texts. For example, the image description texts are input into the large model, and the large model generates static images that conform to the descriptions.

[0051] The picture book video generation method provided in this embodiment generates a story text according to the theme input by the user, then generates multiple image description texts for multiple sub-story texts, generates static images based on the image description texts, and generates a picture book video through the static images. In this way, the user does not need to operate multiple applications, and can generate a picture book video by inputting the theme, with higher operation convenience and improved user experience. In addition, using the above method, in the process of generating static images from the story text, the image description text is first generated, which can describe the semantic content included in the static image. The image description text not only conforms to the plot of the story, but also can guide the generated static image, so as to ensure that the content of the static image is consistent with the content of the story, improve the generation effect of the static image, and further make the picture book video more in line with the user's needs.

[0052] Next, the process of generating a story text based on the theme will be described.

[0053] In this embodiment, the user can input an audio mode through the front-end page, and can also upload custom character information or select pre-configured character information. The character information may include a character image and a character name, and all these information are added to the input information. It should be noted that other embodiments may lack character information.

[0054] Next, the electronic device executing the picture book video generation method obtains the input information, and then inputs the theme, audio mode, character name, and story prompt information into the large model to obtain the story text output by the large model.

[0055] For example, the story prompt information is used to guide the large model to generate a story text based on the theme when the audio mode is the reading mode; and the story prompt information is also used to guide the large model to generate a story text with a rhyming style based on the theme when the audio mode is the singing mode. The above story prompt information can be pre-configured according to actual needs, and this embodiment does not limit the story prompt information.

[0056] In this embodiment, the user can freely select the audio mode through the front-end page. The audio mode includes the reading mode and the singing mode. The reading mode represents reading the story text, and the singing mode represents singing the story text. The large model can generate story texts adapted to different audio requirements. For example, it outputs a story text with regular narration in the reading mode and automatically generates a more rhyming story text in the singing mode, making the story text more matched with the audio mode and improving the user experience.

[0057] Next, the process of generating static images based on the image description text will be described.

[0058] In this embodiment, the user can input or select a video ratio through the front-end page. The video ratio can represent the ratio of the length to the width of an image, and the video ratio can include 9:16, 16:9, 1:1, etc. The user can input or select a video style through the front-end page, and the video style can include a cartoon style, a watercolor style, a coloring style, etc. The user can also upload custom character information or select pre-configured character information. The character information can include a character image and a character name, and all this information is added to the input information. It can be understood that in other embodiments, the character information may be missing.

[0059] Next, the electronic device executing the picture book video generation method obtains the input information, and then inputs the image description text, the video ratio, the video style, the character image, and the picture prompt information into the large model to obtain a static image output by the large model. The picture prompt information is used to guide the large model to generate a static image based on the input information.

[0060] In this embodiment, the user can freely select the video ratio and the video style to meet various needs of the user. By inputting parameters such as the image description text, the video style, and the video ratio into the large model, an image with a style and ratio that both adapt to the requirements can be generated, reducing the difficulty of image generation.

[0061] Next, the process of generating a picture book video based on the static image will be described.

[0062] In one embodiment, the user can input or select a video effect through the front-end page, and the video effect can be a still image effect.

[0063] The transition effect can be pre-configured, or the corresponding transition effect can be selected from multiple effects according to the video style, the video effect, etc. The transition effect can be a book page turning effect, etc. This embodiment does not limit the transition effect. After generating multiple static images, multiple static images can be spliced based on the transition effect to obtain the image frame sequence included in the picture book video. For example, multiple static images are displayed in sequence, and the order of the static images is consistent with the order of the story sub-text, and a transition is made between two adjacent static images through the transition effect.

[0064] In this embodiment, multiple static images are spliced through the transition effect, improving the smoothness of the picture book video, alleviating the difficulty of the user using traditional video editing software for the transition between images, and thus reducing the difficulty and workload of generating the picture book video. In addition, by corresponding the order of the story sub-text with the image frame sequence of the static images, the coherence of the story narrative logic is ensured.

[0065] In another embodiment, the user can input or select a video effect through the front-end page, and the video effect can be an animated image effect.

[0066] After generating multiple static images, based on a large model, animated image prompt information can be generated according to the static images and story text. For example, the image, story text, and pre-configured prompt information are input into the large model to obtain the animated image prompt information output by the large model. The animated image prompt information can be text, and it is used to guide the large model on how to convert the static image into a dynamic image. For example, the animated image prompt information describes the blinking of the little girl's eyes, the falling of leaves, etc. Next, the static image and the animated image prompt information can be input into the large model, and the large model adjusts the objects in the static image based on the animated image prompt information so that the objects move between multiple frames of images to form a dynamic effect, obtaining the dynamic image output by the large model. Next, multiple dynamic images can be spliced based on the transition effect to obtain the sequence of image frames included in the picture book video. For example, multiple dynamic images are displayed in sequence, and the order of the dynamic images is consistent with the order of the story sub-text, and a transition is made between two adjacent dynamic images through the transition effect.

[0067] In this embodiment, the user can select the animated image effect. The electronic device can use the large model to generate the animated image prompt information based on the story text and the static image, and then generate the dynamic image through the animated image prompt information, realizing the conversion from the static image to the dynamic image and making the generated picture book video more vivid. In addition, the animated image prompt information can more accurately guide the large model to control the process of image dynamicization, ensuring that the actions in the image conform to the content of the story. In addition, there is no need for the user to manually design the dynamic change method of the image, reducing the difficulty of generating the picture book video with the animated image effect.

[0068] According to another embodiment of the present disclosure, the generated picture book video not only includes pictures, which can be the image frame sequence described above, but also can include audio. The user can input or select a voice color through the front-end page. The voice color can include the voice color extracted from a piece of audio uploaded by the user, and can also include multiple pre-configured candidate voice colors, such as cute female voice, kind female voice, etc. The user can also not select a voice color, so as to use the default voice color or a random voice color as the voice color in the input information. The user can also input or select an audio mode through the front-end page. The audio mode includes a reading mode and a singing mode. The reading mode represents reading the story text, and the singing mode represents singing the story text using a target melody. Among them, the target melody can be selected by the user on the front-end page, can also be the default, can also be randomly selected, or can be matched according to the lengths of the respective story sub-texts in the story text. The story text can be converted into the audio included in the picture book video by using the voice color and the audio mode in the input information. For example, if the user selects a cute female voice and the reading mode, the generated audio is to read the story text with the voice color of the cute female voice. If the user selects their own voice color and the singing mode, the generated audio is to sing the story text using the voice color uploaded by the user and the target melody. In this embodiment, the user can use their own voice color or a pre-configured voice color to read or sing the story text, which can make the picture book video more in line with the user's needs and improve user satisfaction.

[0069] According to another embodiment of the present disclosure, the generated picture book video not only includes pictures, which can be the image frame sequence described above, but also can include subtitles. The user can input or select a language category through the front-end page. The language category includes Chinese, English, Japanese, etc. The story text can be displayed in the language indicated by the language category in the input information as the subtitles included in the picture book video. For example, if the language category selected by the user is Chinese, the Chinese story text is used as the subtitles. In this embodiment, the subtitles can be displayed according to the language category selected by the user, thereby improving user satisfaction.

[0070] Figures 3A - 3B It is a schematic diagram of the front-end page of the picture book video generation method according to the embodiment of the present disclosure.

[0071] Next, the picture book video generation process provided in this embodiment will be described.

[0072] As Figure 3AAs shown in the figure, the user can input a theme through the front-end page 301 and select some options, such as selecting role settings, picture dubbing, picture ratio, video effects, picture language, video style, etc. The above role settings can include custom role information uploaded by the user or pre-configured role information. The role information can include a role image and a role name. The above picture dubbing can include timbre and audio mode. If the user creates an audio through the "created by me" option, the timbre can be extracted based on the audio. If the user selects a pre-configured option in the narration dubbing, it means that the user has selected the reading mode and the timbre in the option. If the user selects a music video, it means that the user has selected the singing mode. On this basis, the user can select a custom timbre or a pre-configured timbre. The above picture ratio refers to the aspect ratio of the image in the picture book video, that is, the video ratio in the above text. The picture ratio can include 9:16, 16:9, 1:1, etc. The video effects can include still image effects and animated image effects. The picture language is the language category in the above text. The picture language can include Chinese, English, Japanese, etc., indicating the language used for the subtitles. The video style can include cartoon style, watercolor style, coloring style, etc. For each of the above options, the user can make a selection according to actual needs. If the user does not make a selection, the content of the option can be empty, or the default option can be used, or an option can be randomly selected. It should be noted that if the user selects to upload custom role information and audio, the user is aware of and consents to the acquisition and use of such information, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0073] Next, when the user clicks the "create immediately" option, the front-end page 301 can encapsulate the theme and option content input by the user into input information and send it to the electronic device that executes the picture book video generation method, and the electronic device generates the picture book video. The electronic device can be a client or a server. In addition, information for interacting with the user can be displayed through the front-end page, such as displaying "Okay, we are about to create a picture book outline for you. Please wait a moment."

[0074] As Figure 3B shown in the figure, next, the electronic device that executes the picture book video generation method first generates a story text based on the theme. For example, the theme, the role name in the role information, the audio mode in the picture dubbing, and the first prompt information prompt1 are input into the large model, and the large model outputs a story text including multiple story sub-texts. The first prompt information prompt1 can be the story prompt information in the above text. The story text includes 5 story sub-texts. Then the story text can be returned to the front-end page 302 for display. The user can modify the story text according to actual needs. When the user determines that the story text does not need to be adjusted, the user can select the "produce picture book" option.

[0075] Next, the electronic device implementing the picture book video generation method responds to the user's front-end operation and generates an image description text based on the story text. For example, the story text, the character name in the character information, and the second prompt information prompt2 are input into a large model, and the large model outputs the image description text. The story text includes multiple story sub-texts, and the story sub-texts and the image description texts can correspond one by one.

[0076] Next, a picture book video can be generated based on the story text. It can be understood that since it takes time to generate a picture book video, interactive information can be displayed through the front-end page. For example, "Generating picture book" can be displayed, and information such as the generation progress can also be displayed.

[0077] Next, the image description text, the character image in the character information, the video ratio, the video style, and the third prompt information prompt3 can be input into a large model, and the large model outputs static images. Each image description text can be used to generate a static image. In the actual generation process, multiple image description texts can be processed in parallel.

[0078] Next, a picture book video can be generated based on the static images. If the input information includes a video effect and the video effect is a still image effect, the generated multiple static images can be spliced into an image frame sequence through a transition effect, so as to obtain the picture in the picture book video. In addition, the picture book video can also include audio and subtitles. For example, the story text can be read or sung in the tone of voice in the voice-over for the picture, so as to obtain the audio in the picture book video. The story text can also be displayed in the language in the language category as the subtitles in the picture book video. Next, the image frame sequence, audio, and subtitles can be combined into a picture book video and displayed.

[0079] It should be noted that after the picture book video is generated, a single image in the picture book video is a static image. If the user hopes to convert it into a dynamic image, they can select "Convert to dynamic video" through the front-end page, and then the electronic device can continue to convert the static image into a dynamic image, and then regenerate and display the picture book video.

[0080] If the input information includes a video effect and the video effect is a GIF effect, a static image can be generated first in the manner described above, and then the static image, the story text, and the fourth prompt information prompt4 are input into the large model. The large model outputs the fifth prompt information prompt5, which can be the GIF prompt information in the above text and is used to guide the large model to convert the static image into a dynamic image. Then, the generated static image and the fifth prompt information prompt5 are input into the large model, and the large model outputs a dynamic image. Each static image can be used to generate a dynamic image, and then multiple dynamic images can be spliced into an image frame sequence through a transition effect to obtain the pictures in the picture book video. Audio and subtitles can also be generated, and the image frame sequence, audio, and subtitles are combined into a picture book video and displayed.

[0081] It should be noted that after the picture book video is generated, the images in the picture book video are dynamic images. If the user hopes to convert them into static images, they can select "Convert to Static Video" through the front-end page, and then the electronic device can continue to splice the static images used to generate the dynamic images with the transition effect to obtain an image frame sequence, and then reprocess the image frame sequence, audio, and subtitles into a picture book video and display it.

[0082] In this embodiment, the user inputs a theme through the front-end page, selects options such as video effects, and clicks Start Creation to generate a picture book video that meets the requirements, with high operation convenience.

[0083] It should be noted that in the picture book video generation method provided in this embodiment, the large model is used in multiple steps. For example, the processes of generating the story text, generating the image description text based on the story text, generating the static image based on the image description text, generating the fifth prompt information prompt5 based on the static image, and generating the dynamic image all use the large model. The same large model or different large models can be used in each step, as long as the large model used can process the input data and output the corresponding data. The large model can be a large language model, a multi-modal large model, etc. The structure and working principle of the large model are not limited in this embodiment.

[0084] Figure 4 It is a schematic structural block diagram of a picture book video generation device according to an embodiment of the present disclosure.

[0085] As Figure 4 shown, the large model-based picture book video generation device 400 may include: an acquisition module 410, a dynamic image generation module 420, and a picture book video generation module 430.

[0086] The acquisition module 410 is used to acquire input information, and the input information includes a theme.

[0087] The dynamic image generation module 420 is used to generate multiple dynamic images based on a large model according to a theme; wherein, the dynamic images include multiple frames of images, and the objects in the dynamic images change among the multiple frames of images.

[0088] The picture book video generation module 430 is used to generate a picture book video according to multiple dynamic images.

[0089] According to another embodiment of the present disclosure, the input information further includes a video effect; the dynamic image generation module includes: a static image generation sub-module and a conversion sub-module. The static image generation sub-module is used to generate multiple static images based on a large model according to a theme. The conversion sub-module is used to convert multiple static images into multiple dynamic images respectively based on a large model when the video effect is a gif effect.

[0090] According to another embodiment of the present disclosure, the conversion sub-module includes: a gif prompt information generation unit and a processing unit. The gif prompt information generation unit is used to generate gif prompt information based on a large model according to the static image and the story text generated based on the theme when the video effect is a gif effect, and the gif prompt information is used to guide the way for the large model to convert the static image into a dynamic image. The processing unit is used to input the static image and the gif prompt information into the large model to obtain the dynamic image output by the large model.

[0091] According to another embodiment of the present disclosure, the static image generation sub-module includes: a story text generation unit, an image description text generation unit, and a static image generation unit. The story text generation unit is used to generate story text based on a large model according to a theme; the story text includes multiple story sub-texts. The image description text generation unit is used to generate multiple image description texts respectively for the multiple story sub-texts based on a large model according to the multiple story sub-texts. The static image generation unit is used to generate multiple static images according to the multiple image description texts.

[0092] According to another embodiment of the present disclosure, the input information further includes an audio mode; the story text generation unit includes: a first input sub-unit, which is used to input the theme, the audio mode, and the story prompt information into the large model to obtain the story text output by the large model; wherein, the story prompt information is used to guide the large model to generate a story text with a rhyming style based on the theme when the audio mode is the singing mode representing singing the story text.

[0093] According to another embodiment of the present disclosure, the input information further includes: a video ratio and a video style; the static image generation unit includes: a second input sub-unit, which is used to input the image description text, the video ratio, and the video style into the large model to obtain the static image output by the large model.

[0094] According to another embodiment of the present disclosure, the input information further includes an audio mode, and the device further includes: a first audio generation module configured to generate audio for reading the story text when the audio mode is a reading mode; a second audio generation module configured to generate audio for singing the story text with a target melody when the audio mode is a singing mode; wherein, the picture book video includes audio.

[0095] According to another embodiment of the present disclosure, the input information further includes a language category, and the device further includes: a subtitle generation module configured to display the story text generated based on the theme in the language indicated by the language category as subtitles included in the picture book video.

[0096] According to another embodiment of the present disclosure, the picture book video generation module includes: a splicing sub-module configured to splice a plurality of dynamic images based on a transition effect to obtain a sequence of image frames included in the picture book video.

[0097] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above video generation method.

[0098] According to an embodiment of the present disclosure, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above video generation method.

[0099] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product including a computer program, and the computer program implements the above video generation method when executed by a processor.

[0100] Figure 5 is a structural block diagram of an electronic device for implementing the picture book video generation method of the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0101] As Figure 5As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 502 or computer programs loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of device 500 can also be stored. The computing unit 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0102] Multiple components in device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disc, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0103] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the video generation method. For example, in some embodiments, the video generation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the video generation method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the video generation method by any other appropriate means (e.g., by means of firmware).

[0104] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0105] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0106] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0107] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0108] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0109] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other.

[0110] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0111] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A picture book video generation method based on a large model, comprising: Obtaining input information, where the input information includes a theme; Based on the large model, generating a plurality of dynamic images according to the theme; wherein, the dynamic images include multiple frames of images, and the objects in the dynamic images change among the multiple frames of images; Generating a picture book video according to the plurality of dynamic images.

2. The method according to claim 1, wherein, The input information further includes a video effect; the generating a plurality of dynamic images based on the large model according to the theme includes: Based on the large model, generating a plurality of static images according to the theme; When the video effect is a GIF effect, based on the large model, converting the plurality of static images into the plurality of dynamic images respectively.

3. The method according to claim 2, wherein, The converting the plurality of static images into the plurality of dynamic images respectively based on the large model when the video effect is a GIF effect includes: When the video effect is the GIF effect, based on the large model, generating GIF prompt information according to the static images and the story text generated based on the theme, where the GIF prompt information is used to guide the way for the large model to convert the static images into dynamic images; Inputting the static images and the GIF prompt information into the large model to obtain the dynamic images output by the large model.

4. The method according to claim 2, wherein The generating a plurality of static images based on the large model according to the theme includes: Based on the large model, generating a story text; the story text includes a plurality of story sub - texts; Based on the large model, generating a plurality of image description texts respectively for the plurality of story sub - texts; and Generating the plurality of static images according to the plurality of image description texts.

5. The method according to claim 4, wherein The input information further includes an audio mode; the generating a story text based on the large model according to the theme includes: Inputting the theme, the audio mode, and story prompt information into the large model to obtain the story text output by the large model; wherein, the story prompt information is used to guide the large model to generate a story text with a rhyming style based on the theme when the audio mode is a singing mode representing singing the story text.

6. The method according to claim 4, wherein, The input information further includes: a video ratio and a video style; the generating the plurality of static images according to the plurality of image description texts includes: Inputting the image description texts, the video ratio, and the video style into the large model to obtain the static images output by the large model.

7. The method according to any one of claims 1 to 6, wherein, The input information further includes an audio mode, and the method further includes: When the audio mode is a reading mode, generating an audio for reading the story text; When the audio mode is a singing mode, generating an audio for singing the story text using a target melody; wherein, the picture book video includes the audio.

8. The method according to claim 1, wherein The input information further includes a language category, and the method further includes: Displaying the story text generated based on the theme in the language indicated by the language category as the subtitle included in the picture book video.

9. The method according to claim 1, wherein The generating a picture book video according to the plurality of dynamic images includes: Stitch the multiple dynamic images based on the transition effect to obtain the image frame sequence included in the picture book video.

10. A picture book video generation device based on a large model, comprising: An acquisition module, configured to acquire input information, where the input information includes a theme; A dynamic image generation module, configured to generate multiple dynamic images based on the large model according to the theme; wherein, the dynamic image includes multiple frames of images, and the objects in the dynamic image change among the multiple frames of images; A picture book video generation module, configured to generate a picture book video according to the multiple dynamic images.

11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program, where the computer program implements the method according to any one of claims 1 to 9 when executed by a processor.

Citation Information

Patent Citations

  • Multifunctional network interesting transliteration system

    CN101196882A

  • Video synthesis method and device, electronic equipment and storage medium

    CN114466222A

  • Vehicle-mounted desktop dynamic wallpaper generation method and device, equipment, medium and vehicle

    CN117215706A

  • Story video generation corresponding to user input using generative models

    CN118212328A

  • Picture book generation method and device, electronic equipment, storage medium and product

    CN118447131A