Method for generating video, electronic device, medium, and program product

CN121970362APending Publication Date: 2026-05-01DOUYIN VISION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2024-08-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Current technologies for generating videos of written works are time-consuming and difficult, especially for literary works with many words and slow plot progression, making it difficult to quickly generate high-quality narration or viewing videos.

Method used

By determining multiple texts based on text and generating corresponding image frames based on the texts and object features, a concise and visually appealing video is ultimately generated, combining images and audio using a generative artificial intelligence model.

Benefits of technology

It significantly shortens video generation time, improves video quality, and enables users to experience audiovisual content in addition to reading, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970362A_ABST
    Figure CN121970362A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method for generating a video, electronic equipment, a storage medium and a computer program product. The method comprises the following steps: determining a plurality of texts based on a text; based on each of the plurality of documents and the object features of the object associated with each document, generating an image frame corresponding to each document; and generating a video corresponding to the text based on each document and the corresponding image frame. According to the method for generating the video provided by the embodiment of the invention, the literal works can be modified to quickly generate a winning video (such as an explanation video or a browsing video) with refined content and exquisite pictures, the video generation time is remarkably shortened, the quality of the generated video is improved, and the user experience is improved. Therefore, the user can feel audiovisual experience besides reading, and the user experience is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, electronic devices, media, and program products for generating video. Technical Field

[0001] This disclosure generally relates to the field of computers, and more specifically to methods for generating video, electronic devices, computer-readable storage media, and computer program products. Background Technology

[0002] With the continuous advancement of internet technology, various types of cultural products can be easily disseminated through the internet and provided to users via terminal devices, enabling users to browse various types of cultural products on their devices, thereby enriching their work and life and enhancing their user experience.

[0003] Summary of the Invention

[0004] According to exemplary embodiments of this disclosure, a method for generating video, an electronic device, a computer storage medium, and a computer program product are provided.

[0005] In a first aspect of this disclosure, a method for generating a video is provided, comprising: determining multiple texts based on text; generating an image frame corresponding to each text based on object features of each text and an object associated with each text; and generating a video corresponding to the text based on each text and the corresponding image frame.

[0006] In a second aspect of this disclosure, an electronic device is provided, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method described in the first aspect of this disclosure when executed by the at least one processing unit.

[0007] In a third aspect of this disclosure, a computer-readable storage medium is provided having machine-executable instructions stored thereon, which, when executed by a device, cause the device to perform the method described in the first aspect of this disclosure.

[0008] A fourth aspect of this disclosure provides a computer program product including computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method described in the first aspect of this disclosure.

[0009] The summary section is provided to introduce a series of concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or essential features of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a schematic diagram of an example system in which embodiments of the present disclosure can be implemented;

[0012] Figure 2 shows a flowchart of a method for generating video according to an embodiment of the present disclosure;

[0013] Figure 3 illustrates a schematic process for generating a video for text according to an embodiment of the present disclosure;

[0014] Figure 4 illustrates a schematic process of obtaining modified text based on text performed by a model according to an embodiment of the present disclosure;

[0015] Figure 5 shows a flowchart of an exemplary method for generating image frames based on text and object features according to an embodiment of the present disclosure;

[0016] Figures 6A-6B illustrate schematic diagrams of generating corresponding image frames for each text according to embodiments of the present disclosure;

[0017] Figure 7 shows a schematic block diagram of an example apparatus according to some embodiments of the present disclosure; and

[0018] Figure 8 shows a block diagram of an example device that can be used to implement embodiments of the present disclosure. Detailed Implementation

[0019] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0020] Various types of cultural products can be widely disseminated through internet technology, enriching users' work and lives and enhancing user experience. Among the many types of cultural products, written works, such as literature, are widely popular. Written works convey content to users through text and have been widely welcomed and loved.

[0021] Currently, users typically browse written works through reading. As terminal devices offer increasingly richer functionalities, it is hoped that they can also provide users with more diverse browsing methods. For example, in addition to reading, it is desirable for terminal devices to offer explanatory videos (videos explaining the written work) or browsing videos (for example, in browsing videos, the main content of the written work can be presented to the user in video form through video footage, narration, and character dialogue), allowing users to understand the main content of the written work through these explanatory or browsing videos.

[0022] While videos introducing written works already exist, they are primarily explanatory videos, generated mainly through manual editing of the text, hand-drawn graphics, and voice-over. This process is time-consuming, and adapting written works is challenging, especially for lengthy literary pieces with slow-paced plots. Therefore, there is a pressing need for a video generation method that can significantly shorten generation time and improve video quality for introducing written works (e.g., explanatory or browsing videos).

[0023] In view of this, embodiments of the present disclosure provide a method for generating video. The method may include: determining multiple text snippets based on text; generating image frames corresponding to each text snippet based on object features of each text snippet and an object associated with each text snippet; and generating a video corresponding to the text snippet based on each text snippet and the corresponding image frames. By employing the method for generating video according to embodiments of the present disclosure, concise, visually appealing, and engaging videos (e.g., narration videos or browsing videos) can be generated quickly, significantly reducing video generation time while improving the quality of the generated videos, enabling users to experience audiovisual experiences beyond reading, and further enhancing the user experience.

[0024] Embodiments of the present disclosure will now be described in further detail with reference to the accompanying drawings, wherein FIG1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. The example environment 100 includes a computing device 110. The computing device 110 may be deployed with models, which may include various types of intelligent models such as generative artificial intelligence models.

[0025] The computing device 110 may include, but is not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, wearable electronic devices, smart home devices, minicomputers, mainframe computers, edge computing devices, and distributed computing systems that include any one of the above systems or devices.

[0026] The computing device 110 can receive text 150, such as written works, process the received text 150, and generate a video 160 corresponding to the text 150 (e.g., an explanatory video or browsing video for the text 150). In some embodiments, the computing device 110 can determine a set of texts based on the text 150. For example, the computing device 110 can use a model to determine multiple texts based on the text 150. The texts can refer to text played by speech and can have a textual form for easy display in the video frame. The computing device 110 can generate an image frame corresponding to each text based on the object features of each text and the object associated with each text. For example, the computing device 110 can use a model to generate image frames. Further, the computing device 110 can generate a video 160 corresponding to the text 150 based on each text and the corresponding image frame. For example, the generated video 160 can be an explanatory video for the text 150, used to explain the text 150. Alternatively, the generated video 160 can be a browsing video for the text 150. For example, in browsing the video, the main content of the text work can be presented to the user in video form through video footage, narration, and dialogue.

[0027] The method for generating videos according to embodiments of this disclosure can adapt written works to quickly generate concise, visually appealing, and engaging videos (e.g., explanatory or browsing videos), significantly shortening video generation time while improving the quality of the generated videos, enabling users to experience audiovisual content beyond reading, and further enhancing the user experience.

[0028] The block diagram above, with reference to FIG1, illustrates an example environment 100 in which embodiments of the present disclosure can be implemented. The method for generating video according to embodiments of the present disclosure is described below with reference to FIG2. FIG2 shows a flowchart of a method 200 for generating video according to embodiments of the present disclosure. Method 200 can be executed on computing device 110 in FIG1 and any suitable computing device. It should be understood that the numbers in the flowchart of method 200 do not indicate the order in which these steps are performed; some or all of these steps can be performed in parallel, or the order of execution can be interchanged, and the present disclosure does not limit this. Furthermore, method 200 in FIG2 may also include additional steps not shown and / or the steps shown may be omitted, and the scope of the present disclosure is not limited in this respect.

[0029] In box 202, computing device 110 can determine multiple texts based on text 150. In some embodiments, text 150 can be a written work expressed in text form, such as a novel, a natural science work, etc., and this disclosure does not limit the specific type of written work. In some embodiments, the text can refer to text played by voice and can be in text form for easy display in video footage. Multiple texts combined together can form an introduction to text 150. In other words, introducing text 150 can be achieved by playing multiple texts by voice. In some embodiments, for narrated videos, the text can include narration text. In some embodiments, for browsing videos, the text can include narration text and / or dialogue text.

[0030] According to embodiments of this disclosure, determining multiple text snippets based on text can be performed using a model, thereby improving the language quality of the determined text and enhancing its plot appeal and character immersion, thus facilitating the generation of engaging videos. The text generation implementation will be described in detail below.

[0031] In box 204, computing device 110 can generate an image frame corresponding to each text based on the object features of each text in the multiple texts and the objects associated with each text.

[0032] In some embodiments, an object may refer to an object in text. Depending on the text type, the object may have different types. In some embodiments, an object may include one or more of the following: a person, an animal, a plant, an object, or a scene. For example, for fictional texts, the object may be one or more of the following: a person, an animal, a plant, an object, or a scene; for natural science texts, the object may be one or more of the following: an animal, a plant, an object, or a scene. This disclosure does not limit the specific type of the object.

[0033] In some embodiments, the computing device 110 can extract object features of at least one object in the text. For example, the computing device 110 can use a model to extract object features of at least one object in the text. In some embodiments, the model can extract object features of the main objects in the text. For example, for a novel text with a large number of words and a complex plot, the model can extract multiple character objects that appear frequently in the novel text and contribute significantly to the development of the plot. In other embodiments, the model can extract object features of all objects in the text. In some embodiments, objects may include scene objects, and accordingly, the model can extract object features of multiple scene objects in the text.

[0034] In some embodiments, each object has object characteristics. Object characteristics may include the object's name and other relevant features. Taking a person object as an example, the object characteristics of a person object may include feature elements such as person's name, identity, age, gender, physical appearance, and clothing. Through these feature elements, the characteristics and image of the object can be represented. Taking an object object as an example, the object characteristics of an object object may include feature elements such as object name, size, shape, and location. The object characteristics of different objects are different from each other.

[0035] In some embodiments, object features may have one or more versions, and at least one corresponding feature element is different in different versions of object features. Taking a person object as an example, different versions of person features can represent the image of the person object at different stages or in different scenarios. For example, for person 1, its object features may include feature elements: identity, age, gender, and clothing. Person 1 may have two versions of features, as shown below:

[0036] Character 1:

[0037] Version 1: Apprentice; 18 years old; male; wearing coarse cloth clothing;

[0038] Version 2: Master; 38 years old; male; wearing a long gown.

[0039] In some embodiments, for extracted object features, the computing device 110 may store the object features in at least one of the following ways: saving the object features as a model file of an object model generated based on the object features; or saving the object features as a language description used to describe the object features. The computing device 110 may generate a model file based on the object features and save the corresponding object features by saving the model file. Alternatively, the computing device 110 may generate a language description based on the object features and save the corresponding object features by saving the generated language description.

[0040] The computing device 110 can generate an image frame corresponding to each text based on the object features of each text in a set of multiple texts and the objects associated with each text. The model can generate image frames corresponding to texts. The detailed generation process of the image frames will be described in detail below, and will not be repeated here.

[0041] In box 206, computing device 110 can generate a video corresponding to the text based on each piece of text and the corresponding image frame. In some embodiments, the generated video can be an explanatory video for the text 150, used to explain the text 150. Alternatively, the generated video can be a browsing video for the text 150; for example, in a browsing video, the main content of the written work can be presented to the user in video form through video footage, narration, and dialogue.

[0042] In some embodiments, the text may include words. The computing device 110 can use text-to-speech (TTS) technology with a model to convert the text of the text into speech segments. The computing device 110 can convert each piece of text into a corresponding speech segment; that is, each speech segment corresponds to one piece of text. Since each piece of text corresponds to a generated image frame, correspondingly, each speech segment also corresponds to an image frame. For example, a piece of text T1, after conversion, corresponds to a speech segment Audio 1; text T1 corresponds to the generated image frame Image1, and correspondingly, the speech segment Audio 1 corresponds to the image frame Image 1.

[0043] In some embodiments, for generating browsing videos, since the browsing videos also contain dialogue, the text can also include dialogue text accordingly. For each piece of text, the computing device 110 can further determine the object initiating the dialogue text if it is determined that the text is dialogue text. For example, the object initiating the dialogue is object A. The computing device 110 can select a timbre corresponding to the object from a preset repository based on the object characteristics of the object. For example, the computing device 110 can select a timbre matching the object characteristics of object A from a pre-established timbre library based on the object characteristics of object A. The computing device 110 can generate a speech segment corresponding to the dialogue text based on the selected timbre. The computing device 110 can perform the above operations for each piece of text, thereby determining a matching timbre for the object of the dialogue text and further generating an audio segment corresponding to the text. Thus, the generated video can present the text content to the user in the form of a dialogue.

[0044] After generating corresponding audio segments for each piece of text, the computing device 110 can determine the audio length corresponding to each audio segment. Furthermore, the determined audio length of the audio segment can be less than or equal to the frame duration of the corresponding image frame in the generated video. That is, the frame duration of the corresponding image frame in the generated video can be greater than or equal to the audio length of the audio segment. For example, if the audio length of the audio segment is t1, then the frame duration (i.e., display duration) of the corresponding image frame in the generated video can be greater than or equal to t1.

[0045] The computing device 110 can also sequentially acquire each image frame corresponding to each text in the order of the text within the multiple texts. Thus, the order of the acquired image frames corresponds to the order of the texts. The computing device 110 can also associate each image frame with each corresponding audio segment. Using the example above as an illustration, the computing device 110 can determine that audio segment Audio 1 corresponds to image frame Image 1, and the computing device 110 can associate image frame Image 1 with audio segment Audio 1. Furthermore, the frame duration (display duration) of each image frame in the generated video can be greater than or equal to the audio length of the associated audio segment.

[0046] Therefore, by associating each text, the corresponding generated image frame, and the audio segment converted from that text, a video 160 corresponding to text 150 can be generated. Furthermore, by selecting a voice tone that matches the object for the dialogue text, the text content can be presented in the form of dialogue during video viewing, increasing the richness of the video.

[0047] In some embodiments, when displaying text in image frames, at least one of multiple texts can be split into multiple sub-texts, and each of the split sub-texts can be displayed at a different frame time in the image frame corresponding to the at least one text. For example, for texts with a relatively large number of words, considering the aesthetics of the generated video images, the text can be split into multiple sub-texts, such that each sub-text can include a smaller number of words. In some embodiments, the text can be split according to the delimiters (e.g., commas) included in the text. For example, for the text "On a clear morning after the rain, I walked alone on a quiet path," the text can be split into two sub-texts, "On a clear morning after the rain" and "I walked alone on a quiet path," based on the commas in the text, and these two sub-texts can be displayed at different frame times in the image frame (e.g., image frame X1) corresponding to the text "On a clear morning after the rain, I walked alone on a quiet path." For example, at frame time t1, the subtext "On a clear morning after the rain" can be displayed at a predetermined position in image frame X1, and at frame time t2, the subtext "I walk alone on a quiet path" can be displayed at a predetermined position in image frame X1. This allows for a relatively small number of characters to be displayed in each image frame, thus not affecting the aesthetics of the image.

[0048] Furthermore, in some embodiments, the computing device 110 can also select at least one image frame from multiple image frames corresponding to multiple text snippets. For example, the selected image frame can be a keyframe among the generated multiple image frames. The computing device 110 can move one or more selected image frames. For example, the computing device 110 can move one or more selected image frames vertically (e.g., from top to bottom or from bottom to top), horizontally (e.g., from left to right or from right to left), etc. And the movement time of each image frame is no greater than the frame duration of that image frame. By selecting and moving the selected image frames, the generated video can exhibit dynamic effects, thereby attracting more user attention and improving the user experience.

[0049] In some embodiments, for each text, the computing device 110 may also determine the object associated with the narration text if the text includes a voice-over text. If the object associated with the narration text is a person object, the computing device 110 may generate a video clip associated with the image frame based on the content of the text and the corresponding image frame of the text. For example, the computing device 110 may input the image frame corresponding to the text and the text into a model (e.g., a video generation model in the model) to generate a video clip corresponding to the image frame. For example, if the narration text is "Xiaoming stood up from his seat," and the object of the narration text is the person "Xiaoming," to enhance the dynamic effect of the video, the image frame corresponding to the text and the text may be input into a model, and the model may generate a video clip. For this example text, the generated video clip may show the dynamic process of the person ("Xiaoming") moving from a sitting position to standing up.

[0050] Furthermore, to further enhance the effect of the generated video, the computing device 110 can also select background music for the video 160 and associate the background music with the video. Thus, when a user plays the video 160, in addition to seeing the text, generating image frames, and hearing audio clips of the text, they can also hear the background music. In some embodiments, the computing device 110 can select background music based on the text 150, and the audio length of the background music corresponds to the duration of the generated video. In some embodiments, the computing device 110 can select background music based on the content of the text 150. For literary novels, the computing device 110 can select soothing music as background music; for comedic novels, the computing device 110 can select upbeat music as background music, and so on.

[0051] Therefore, the method for generating video according to the embodiments of this disclosure can adapt written works to quickly generate concise, visually appealing, and engaging videos (e.g., narration videos or browsing videos), significantly shortening video generation time while improving the quality of the generated videos, enabling users to experience audiovisual experiences beyond reading, and further enhancing the user experience.

[0052] The implementation process of generating video 160 for text 150 according to an embodiment of the present disclosure will be described below with reference to FIG3. FIG3 illustrates a schematic process 300 of generating video 160 for text 150 according to an embodiment of the present disclosure.

[0053] Text 150 can be a textual work that includes text. Computing device 110 can extract object features 310 from text 150 based on text 150. The specific implementation process of computing device 110 extracting object features 310 can be understood by referring to the description of box 204 in Figure 2 above. For the sake of brevity, it will not be repeated here.

[0054] In some embodiments, computing device 110 can determine multiple texts based on text. As shown in FIG3, computing device 110 can obtain modified text 320 (320') corresponding to text 150 based on text 150, and further determine multiple texts 330 based on modified text 320 (320').

[0055] In some embodiments, the generated video 160 can be an explanatory video for the text 150, used to explain the text 150. When video 160 is an explanatory video, modifying the text can include modified text 320 for the explanatory video. Alternatively, the generated video 160 can be a browsing video for the text 150. For example, in a browsing video, the main content of the written work can be presented to the user in video form through video footage, narration, and dialogue. When video 160 is a browsing video, modifying the text can include modified text 320' for the browsing video.

[0056] In some embodiments, the computing device 110 may obtain the modified text 320 (or modified text 320') corresponding to the text 150 based on the text 150, and determine multiple texts 330 based on the modified text 320 (or modified text 320').

[0057] The following describes the specific implementation process of obtaining modified text 320 for generating narration videos and determining multiple texts based on modified text 320.

[0058] In some embodiments, when generating the narration video 160, the modified text 320 can be a narration script generated based on text 150. This narration script has a significantly reduced word count compared to text 150. In other words, the narration script is a distillation of the content of text 150. Through this narration script, users can clearly understand the content and key points of text 150 without losing any important information.

[0059] In some embodiments, when generating modified text 320 for the narration video, the computing device 110 can generate modified text 320 corresponding to the text 150 based on the text 150 and prompt information for generating the modified text 320. In some embodiments, the prompt information provides information for generating the modified text 320, and the prompt information includes modification criteria for modifying the text 150.

[0060] In some embodiments, the modified text 320 (i.e., the narration script for the video) generated based on text 150 may conform to a specified standard (i.e., a modification standard). Since the narration script is a refinement of text 150, and to further enhance the user experience, it is desirable for the narration script to conform to certain standards. In some embodiments, this standard may include:

[0061] - The video should be engaging from the outset. For example, this can be achieved by directly constructing a background or introducing key objects, thus grabbing the user's attention from the very beginning. Furthermore, based on the text content type, other methods can be used to make the video opening more appealing, which are not limited in this disclosure.

[0062] - The narrative in the explanatory text is engaging. For example, the explanatory text can provide users with the main content of text 150 through a strong sense of rhythm, good transitions, and rich imagery, thereby providing users with an engaging plot. Understandably, for texts 150 with a lot of text (e.g., novel text), due to their length and slow plot progression, by having the computing device 110 output explanatory text containing the main content of text 150 with a strong sense of rhythm, users can understand the content of the text in a shorter time, enhancing the appeal of the explanatory text.

[0063] - High-quality language expression. Based on this standard, the narration script can possess high fluency, high accuracy, and high conciseness. This results in multiple scripts generated based on the narration script being fluent, accurate, and concise, thereby further improving the quality of the generated video.

[0064] - Role immersion. In some embodiments, the narration can be written in the first person. For example, the most important subject in text 150 can be described in the first person in the narration, which can make the generated video more immersive for the user.

[0065] It is understood that the above standards are merely illustrative. You can select one or more of these standards depending on the needs of the generated narration video. Furthermore, other standards can be set to ensure the generated narration video meets expectations.

[0066] In some embodiments, the computing device 110 may receive prompting information for generating the modified text 320, and utilize a model to generate the corresponding modified text 320, i.e., explanatory text, from the text 150. In some embodiments, the prompting information may include one or more of the standards listed above described in natural language, and the prompting information may also include examples of generating the modified text 320 to facilitate model learning. The model may generate the corresponding modified text 320 for the text 150 based on the prompting information. Table 1 shows an example of the prompting information.

[0067] Table 1

[0068] The prompts in Table 1 are merely illustrative and for purposes of explanation. For simplicity, the text A to C and the output explanatory text A' to C' in the examples are represented by symbols, but it is understood that the text A to C and the output explanatory text A' to C' are actually represented by words.

[0069] In some embodiments, the model can generate corresponding modified text 320 based on text 150. The model can be trained as follows: labeled training samples are obtained, including sample text and corresponding explanatory text; the model is trained using the training samples. After obtaining the trained model, its performance is evaluated on a validation dataset. If the model performance meets expectations, it is saved for subsequent prediction processing. If the model performance does not meet expectations, training continues with updated samples until a model that meets expectations is obtained.

[0070] The above describes the implementation process of generating modified text 320 for narration videos. The following describes the implementation process of determining multiple texts 330 based on the modified text 320. In some embodiments, after the computing device 110 generates the corresponding modified text 320 (i.e., narration text) based on the text 150, the computing device 110 can determine multiple texts 330 based on the modified text 320, as illustrated in the block diagram in FIG3.

[0071] In some embodiments, computing device 110 can determine multiple segmentation symbols in modified text 320. Segmentation symbols can be punctuation marks in the modified text 320, such as periods, question marks, exclamation marks, etc., used to separate characters and indicate the end of a sentence. Computing device 110 can segment the modified text 320 based on multiple segmentation symbols to obtain multiple segmented texts. For example, computing device 110 can segment the modified text 320 according to segmentation symbols and determine the characters preceding the segmentation symbols that have not yet been determined as segmented texts as segmented texts. Computing device 110 can determine each text in multiple texts 330 based on the number of characters in each segmented text.

[0072] Specifically, for each segmented text, in response to the fact that the number of characters in the segmented text is not greater than a threshold number, the computing device 110 can treat the segmented text as a single text. For example, the computing device 110 can identify segmentation symbols such as periods, question marks, and exclamation marks in the modified text 320, and identify the characters before the segmentation symbols that have not yet been identified as segmented text as segmented text. For each segmented text, in response to the fact that the number of characters in the segmented text is not greater than a threshold number, the computing device 110 can treat the segmented text as a single text. In some embodiments, the threshold number can be predetermined, for example, 40 characters, and this disclosure does not limit it.

[0073] For each segmented text, in response to the number of characters in the segmented text exceeding a threshold, the computing device 110 can perform at least one segmentation on the set of characters in the segmented text, and assign multiple subsets of characters after at least one segmentation to multiple texts respectively. During the segmentation process, the computing device 110 can segment the text according to the delimiters present in the segmented text and the number of characters between adjacent delimiters. In some embodiments, delimiters may include symbols such as commas, ellipses, pauses, colons, and quotation marks.

[0074] For example, consider a segmented text with 60 characters and a threshold of 40. The computing device 110 determines that the number of characters in the segmented text exceeds the threshold; therefore, the segmented text needs to be segmented again. The computing device 110 can further segment the character set exceeding the threshold based on the delimiters between symbols in the segmented text. The computing device 110 can identify the delimiters in the character set and divide the character set into multiple character subsets according to the delimiters. The computing device 110 can cluster the multiple character subsets based on the number of characters in each character subset to achieve further segmentation of the character set exceeding the threshold. For example, the computing device 110 determines that the number of characters in the multiple character subsets are 2, 7, 11, 14, 5, 8, and 13, respectively. Accordingly, the computing device can use the character subsets with 2, 7, 11, or 14 characters in the character subset as the first group of characters, and the character subsets with 5, 8, or 13 characters in the character subset as the second group of characters, and use the first group of characters and the second group of characters as two separate texts.

[0075] The following will describe an exemplary implementation process for obtaining multiple texts 330 for a narration video, with examples. For instance, during the generation of the narration video, the computing device 110 generates the following modified text M for the text 150:

[0076] "After drifting for more than ten days, I finally reached the shore by boat. I saw a forest and walked straight into it. It was a dense forest with tall and straight trees on both sides. Finally, I reached the end of the forest, where I found a house. Would anyone live here?"

[0077] The computing device 110 can determine segmentation marks in the modified text M, such as periods and question marks. The computing device 110 can then segment the modified text M according to these segmentation marks to obtain multiple segmented texts. For example, for the modified text M, the computing device 110 can obtain the following multiple segmented texts:

[0078] Segmented text 1: "After more than ten days of drifting, I finally arrived at the shore by boat."

[0079] Segmented copy 2: "I saw a forest, so I walked straight into it."

[0080] Segmenting the text 3: "This is a dense forest, with tall and straight trees on both sides."

[0081] Segmented copy 4: "Finally, I reached the end of the forest, where I found a house. Would anyone live here?"

[0082] The computing device 110 can determine the number of characters in each segmented text. For example, the computing device 110 can determine that the number of characters in segmented texts 1 to 4 are 20, 18, 20, and 33, respectively. Assume the threshold number is 40. The computing device 110 can determine that the number of characters in each segmented text is less than the threshold number. Therefore, the computing device 110 can use the currently determined segmented texts 1 to 4 as texts 1 to 4 respectively.

[0083] The modified text M above is merely illustrative and for explanation purposes. It is understood that the modified text M will differ depending on the specific text 150, and the resulting segmented text will also differ.

[0084] Alternatively, the computing device 110 can also utilize a model to segment the modified text 320 to obtain multiple pieces of text. The computing device 110 can also utilize a model trained to generate multiple pieces of text based on the modified text to segment the modified text 320 to obtain multiple pieces of text.

[0085] The above examples illustrate the specific implementation process of obtaining modified text 320 and determining multiple text snippets based on modified text 320 for generating narration videos. The following examples will describe in detail the specific implementation process of obtaining modified text 320' and determining multiple text snippets based on modified text 320' for generating browsing videos.

[0086] In some embodiments, when generating the browsing video 160, the modified text 320' can be copy generated based on text 150 (e.g., script). In the browsing video, the main content of the written work can be presented to the user in video form through video footage, narration, and character dialogue. Accordingly, on the one hand, the modified text 320' generated based on text 150 has a significantly reduced number of words compared to text 150. That is, the modified text 320' is a distillation of the content of text 150. Through the modified text 320', the user can clearly understand the content and key points of text 150 without losing any important information. On the other hand, the modified text 320' can include narration and / or dialogue. Narration can be associated with one or more objects to explain, prompt, or illustrate those objects. Dialogue can be a conversation initiated by an object (e.g., a character object) in text 150.

[0087] In some embodiments, when generating modified text 320' of a viewed video, the computing device 110 can generate modified text 320' corresponding to the text 150 based on the text 150 and prompting information for generating the modified text 320', using a model. In some embodiments, the prompting information provides information for generating the modified text 320', and the prompting information includes generation prompting information instructing the model to generate the modified text 320'. In some embodiments, the generation prompting information may include at least one of the following:

[0088] -Scene Analysis: The indicator model divides text 150 into multiple scenes based on scene objects in text 150.

[0089] - Person / Object Recognition: Instructs the model to acquire the object features of multiple person / objects in the text 150 and the relationships between different person / objects;

[0090] - Dialogue Extraction: The model is instructed to extract multiple dialogues from text 150 and determine at least one associated feature for each dialogue. For each dialogue, the associated feature may include at least one of the following: the scene in which the dialogue takes place, the person who initiates the dialogue, the emotion associated with the dialogue, or the content of the dialogue.

[0091] - Narration Transformation: Instruct the model to process text 150 to extract multiple narration segments. For example, instruct the model to transform statements in text 150, such as background information, plot explanations, or statements expressing the author's viewpoint, into multiple narration segments; or

[0092] - Organize the structure of the modified text: Instruct the model to combine one or more of the object features and relationships of multiple scenes, multiple characters, multiple dialogues, or multiple narrations, and use the combined result as the modified text 320'.

[0093] In some embodiments, the prompt information may include one or more of the generated prompt information listed above, described in natural language, and may also include an example of generating modified text 320' to facilitate model learning. The model can generate modified text 320' for browsing the video based on the prompt information for text 150. Table 2 shows an example of generating modified text 320' for browsing the video.

[0094] Table 2

[0095] The prompts in Table 2 are merely illustrative and for informational purposes. Various types of generated prompts can be set based on expectations or needs regarding text modification.

[0096] For scene analysis, the model can extract scene-related descriptions from the text and transform the extracted scene descriptions into scene settings in the modified text 320'. Scene settings can include scenes and features related to those scenes. In some embodiments, the model can classify and number the extracted scenes according to chronological order and plot development. For example, text 150 may contain multiple different scenes (e.g., scene S1, scene S2, and scene S3, with scene S1 preceding scene S2 and scene S2 preceding scene S3 in chronological order) appearing in different chapters. The model can determine the scene order in the scene settings as: scene S1, scene S2, and scene S3, based on the chronological order of the scenes in text 150. In some embodiments, when a scene includes multiple sub-scenes, the model can also identify these sub-scenes and display them along with their related features in the output modified text 320'.

[0097] For character object recognition, the model can extract object features from the text 150 and create a corresponding profile (i.e., a set of object features) for that character. For example, for character A, the model can create a profile for character A based on character A's features (e.g., name, age, personality traits, appearance, etc.), such as "The protagonist, Xiaoming, is a brave and resourceful young man with deep eyes and a firm gait." In some embodiments, the model can also determine the relationships between different character objects. For example, character A and character B are classmates. In some embodiments, the model can also consider the scene in which the character is located when creating a profile for a character object. That is, the profile of a character object can be associated with the scene, so that the same character can have different profiles in different scenes, thereby enabling the character object to better adapt to the scene and making the generated modified text 320' more coherent and fluent.

[0098] For dialogue extraction, the model can extract multiple dialogues from text 150. The model can group dialogues according to different scenarios and plots. For example, dialogues in scenario S1, dialogues in scenario S2, etc. The model can also adjust the format of the dialogues and label the speaker, the content of the dialogue, and the emotion expressed. For example, consider the following dialogue: "Xiaohong: (Calmly) Don't panic, let's look around for any clues." Thus, the model can determine at least one association feature for each dialogue in the multiple dialogues. For each dialogue, the association feature can include at least one of the following: the scenario in which the dialogue occurs, the speaker who initiates the dialogue, the emotion associated with the dialogue, or the content of the dialogue.

[0099] For narration conversion, the model can transform statements in text 150 that provide background information, explain the plot, or express the author's viewpoint into narration statements, making it easier for users to understand quickly. In some embodiments, the language style of the narration can be consistent with the overall style of text 150. For example, if text 150 is a humorous novel, the narration can have a humorous and lighthearted style; if text 150 is a suspense novel, the narration can have a suspenseful style.

[0100] Regarding the structure of the modified text, the model can combine one or more of the following based on prompts: object features and relationships from multiple scenes and characters, dialogues, or narration. The combined result is then used as the modified text 320'. The model can arrange the structure reasonably to ensure that the generated modified text 320' is attractive and fluent.

[0101] It is understood that the above-mentioned prompts are merely illustrative. You can select one or more of the above prompts depending on your needs for the generated video. Furthermore, you can set other prompts to ensure the generated video meets your expectations.

[0102] Figure 4 illustrates a schematic process of obtaining modified text based on text performed by a model according to an embodiment of the present disclosure. As shown in Figure 4, the model can receive text 150. Text 150 may include text of the form of a novel, etc. The model can process the input text 150 according to the prompts used to generate the modified text 320'. In some embodiments, the model can perform a scene segmentation operation 410 on the text 150 according to the prompts, and obtain multiple scenes 440: scene 440-1, scene 440-2, scene 440-3...scene 440-N. The model can set the order of the multiple scenes 440: scene 440-1, scene 440-2, scene 440-3...scene 440-N according to the chronological order of the scenes 440 in the text 150. In some embodiments, during the process of scene segmentation of text 150, the model can also obtain a text summary 401 (as shown by the dotted line in Figure 4) and combine text 150 and the text summary 401 of text 150 to perform scene segmentation operation 410 on text 150.

[0103] The model can also perform person object recognition operation 420 on the text 150 to obtain object features 450 of multiple person objects in the text 150 and the correlation between different person objects.

[0104] The model can also extract dialogue from text 150 (430) and obtain dialogue (460). In some embodiments, dialogue (460) may include conversation (462) and narration (464). Specifically, the model can extract multiple conversations from text 150 and determine the association features of each conversation, wherein, for each conversation, the association features include at least one of the following: the scene in which the conversation takes place, the character who initiates the conversation, the emotion associated with the conversation, or the content of the conversation. The model can also convert statements in text 150 that provide background information, explain the plot, or express the author's point of view into narration statements to extract multiple narrations (464) from text 150.

[0105] As shown in Figure 4, the model can also organize the structure of the modified text. Based on the prompts, the model can combine one or more of the following: multiple scenes 440, object features 450 and relationships of multiple characters, multiple dialogues 462, and multiple narrations 464, and use the combined result as the modified text 320'. The model can arrange the structure reasonably to ensure that the generated modified text 320' is attractive and fluent.

[0106] In some embodiments, the modified text 320' may include multiple scene scripts: scene script 1, scene script 2, scene script 3... scene script 4, and each scene script corresponds to each of multiple scenes 440-1, scene 440-2, scene 440-3... scene 440-N. For example, scene script 1 corresponds to scene 440-1, scene script 2 corresponds to scene 440-2, scene script 3 corresponds to scene 440-3, and so on. Each scene script may include at least one dialogue and / or at least one narration, and the multiple scene scripts are ordered in the modified text according to the order in which the corresponding scenes appear in the text 150.

[0107] In some embodiments, the model can generate corresponding modified text 320' based on text 150. The model can be trained as follows: labeled training samples are obtained, including sample text and the corresponding scene script; the model is trained using the training samples. After obtaining the trained model, its performance is evaluated on a validation dataset. If the model performance meets expectations, it is saved for subsequent prediction processing. If the model performance does not meet expectations, training continues with updated samples until a model that meets expectations is obtained.

[0108] The above describes the implementation process of generating modified text 320' for browsing videos. The following describes the implementation process of determining multiple texts 330 based on the modified text 320'. In some embodiments, after the computing device 110 generates corresponding modified text 320' (i.e., multiple scene scripts) based on text 150, the computing device 110 can determine multiple texts 330 based on the modified text 320', as illustrated in the block diagram in Figure 3.

[0109] In some embodiments, for each scene script in the scene script, the computing device 110 can distinguish between narration and dialogue within the scene script. For narration in the scene script, the computing device 110 can segment the narration based on the object associated with it to obtain at least one narration text. For dialogue in the scene script, the computing device 10 can segment the dialogue based on the object initiating the dialogue to obtain at least one dialogue text. The computing device 110 can use each of the at least one dialogue text and the at least one narration text as each of multiple texts.

[0110] The following will describe an exemplary implementation process for obtaining multiple text snippets 330 used for browsing videos, with examples. For instance, during the generation of a video for browsing, the computing device 110 generates a scenario script for the text 150 as follows:

[0111] "A light came on in the secret room. Xiaoming raised a torch and shouted, 'Is anyone there?'"

[0112] The computing device 110 can distinguish between narration and dialogue in a scene script. Narration in a scene script is language used to explain or describe events, such as, "A light came on in the locked room, and Xiaoming raised a torch," while dialogue is what the characters say, such as, "Is anyone there?"

[0113] For the narration "A light came on in the locked room, and Xiaoming raised a torch" in the scene script, the computing device 110 can segment the narration based on the objects associated with it to obtain at least one narration text. The narration includes two associated objects, such as the locked room where the light came on and Xiaoming raising a torch. Accordingly, the computing device can segment the narration into two narration texts:

[0114] Narration 1: "A light appeared in the secret room"; and

[0115] Narration 2: "Xiao Ming raised the torch."

[0116] For dialogue in a scene script, computing device 10 can segment the dialogue based on the object initiating the dialogue to obtain at least one dialogue text. For example, for the dialogue "Is anyone there?", if the initiator of the dialogue is Xiaoming, computing device 110 can identify this text as a dialogue text.

[0117] Dialogue 1: "Is anyone there?"

[0118] The computing device 110 can use each of the dialogue script 1 and the narration scripts 1 and 2 as one of multiple scripts. That is, by segmenting the text in the scene script, the computing device 110 obtains three scripts:

[0119] Narration 1: "A light appeared in the secret room";

[0120] Narration 2: "Xiao Ming raised the torch"; and

[0121] Dialogue 1: "Is anyone there?"

[0122] The scenario scripts above are merely illustrative and for explanation purposes. It is understood that the scenario scripts will differ depending on the specific text (150), and the segmented text will also vary.

[0123] Alternatively, computing device 110 can utilize a model to segment the modified text 320' to obtain multiple pieces of text. Computing device 110 can also utilize a model trained to generate multiple pieces of text based on the modified text to segment the modified text 320' to obtain multiple pieces of text.

[0124] Returning to Figure 3, after acquiring multiple texts 330, the computing device 110 can generate an image frame corresponding to each text based on each text in the multiple texts 330 and the object features of the object associated with each text. Each text corresponds to one image frame. Accordingly, the computing device 110 can obtain multiple image frames 340 based on the multiple texts 330 and object features 310, wherein each image frame corresponds to each text and the object features of the object associated with each text.

[0125] The following will describe in detail the implementation process of generating image frames 340 based on text 330 and object features 310. Figure 5 shows a flowchart of an exemplary method 500 for generating image frames based on text and object features according to an embodiment of this disclosure.

[0126] In some embodiments, the computing device 110 may generate a first prompt message corresponding to each of the multiple texts 330 at block 502, wherein the first prompt message provides prompt information for generating an image frame corresponding to the text.

[0127] During the process of generating the first prompt information for each piece of text, the computing device 110 can obtain the second prompt information and, by calling the model and based on the second prompt information and each piece of text, generate the first prompt information corresponding to that piece of text. In some embodiments, the second prompt information provides information for generating the first prompt information and may include at least one of the following: the output format of the first prompt information; the operation to be performed to generate the first prompt information; the rules for generating the first prompt information; or an example for generating the first prompt information. Table 3 shows an example of the second prompt information.

[0128] Table 3

[0129] It is understood that the second prompt information shown in Table 3 is merely exemplary and for illustrative purposes.

[0130] Taking the example of the text above as an example, by adopting a method based on each piece of text and the second prompt information, the computing device 110 can generate the first prompt information corresponding to each piece of text by calling the model. For example, for text 1: "After more than ten days of drifting, I finally arrived at the shore by boat," the generated first prompt could be {{Character A; facing the camera; disembarking; seaside scene; long shot}; {boat; seaside scene; long shot}}; for text 2: "I saw a forest and walked straight towards it," the generated first prompt could be {{Character A; back to the camera; walking towards the forest; forest scene; long shot}; {forest; forest scene; long shot}}; for text 3: "This is a dense forest, with tall and straight trees on both sides," the generated first prompt could be {forest; forest scene; close-up}; for text 4: "Finally, I reached the end of the forest, where I found a house. Would anyone live here?" the generated first prompt could be {{Character A; back to the camera; standing; forest end scene; long shot}; {house; forest end scene; close-up}}.

[0131] In some embodiments, during the generation of the first prompt, the model can refer to contextual information to determine the scene in the first prompt. For example, when the text contains a scene cue word, the model can determine the scene in the first prompt based on that cue word. When the text does not contain a cue word, the model can refer to the scene in the first prompt corresponding to the previous text and set the scene in the first prompt corresponding to the current text as the scene in the previous text, thereby ensuring the scene coherence of the generated video.

[0132] For example, considering the narration 1: "A light appeared in the locked room"; narration 2: "Xiao Ming raised a torch"; and dialogue 1: "Is anyone there?", the model can generate the first prompt information {{light; locked room; close-up}} for narration 1: "A light appeared in the locked room". Since the scene information is not determined in narration 2, the model can refer to the scene of the first prompt information in narration 1 and determine the scene in the first prompt information of narration 2 as the scene "locked room" of the first prompt information in narration 1. Therefore, the first prompt information for narration 2 can be {Xiao Ming; facing the camera; raising a torch; locked room; close-up}. Since the scene information cannot be determined in dialogue text 1, the model can refer to the scene of the first prompt in the narration text 2 adjacent to dialogue text 1, and determine the scene in the first prompt of dialogue text 1 as the scene "locked room" of the first prompt of narration text 2. Thus, the first prompt of dialogue text 1 can be {Xiaoming; facing the camera; saying "Is anyone there; locked room; close-up}.

[0133] The above describes in detail the process of generating the first prompt message for each piece of text. Now, returning to Figure 5, at box 504, the computing device 110 can obtain the object characteristics of the object associated with each first prompt message. In some embodiments, the computing device 110 can obtain the object characteristics associated with the object name in the first prompt message. In some embodiments, the object characteristics can be the object characteristics after saving the extracted object characteristics.

[0134] In some embodiments, for the extracted object features, the computing device 110 may save the object features in at least one of the following ways: by saving the object features as a model file of an object model generated based on the object features; or by saving the object features as a language description used to describe the object features.

[0135] Based on the object name in the first prompt information, the computing device 110 can obtain the object features associated with that object name. For example, depending on how the object features are stored, the computing device 110 can obtain the model file or language description associated with the object name. For example, if the object name in the first prompt information is "Person A", the computing device 110 can obtain the model file corresponding to Person A.

[0136] Returning to Figure 5, at box 506, the computing device 110 can generate an image frame corresponding to each piece of text based on each first prompt message and the object features associated with that first prompt message. In some embodiments, the computing device 110 can attach object features to each corresponding first prompt message to obtain each attached first prompt message. The computing device 110 can input each attached first prompt message into a model (e.g., an image generation model within the model) to generate each image frame corresponding to each piece of text.

[0137] Figures 6A-6B illustrate schematic diagrams of generating corresponding image frames for each piece of text according to embodiments of the present disclosure. Figure 6A shows a schematic diagram of image frames for an explanatory video generated based on text used to generate the explanatory video. Figure 6B shows a schematic diagram of image frames for a browsing video generated based on text used to generate the browsing video. As can be seen from Figures 6A-6B, each piece of text corresponds to one image frame. Furthermore, each corresponding image frame is associated with that text.

[0138] Returning to Figure 3, in Figure 3, the computing device 110 can use a model to convert the text of multiple documents into corresponding speech segments using text-to-speech (TTS) technology. That is, the computing device 110 can use the model to convert each document into a corresponding speech segment; in other words, each speech segment corresponds to one document. Since each document corresponds to a generated image frame, correspondingly, each speech segment also corresponds to an image frame. For example, after conversion, a document T1 corresponds to a speech segment Audio 1, document T1 corresponds to the generated image frame Image 1, and correspondingly, the speech segment Audio 1 corresponds to the image frame Image 1.

[0139] In some embodiments, for generating browsing videos, since the browsing videos also contain dialogue, the computing device 110 can, for each text, further determine the object initiating the dialogue if it is determined to be a dialogue text. For example, the initiator of the dialogue is object A. The computing device 110 can select a timbre corresponding to the object from a preset repository based on the object characteristics of the object. For example, the computing device 110 can select a timbre matching the object characteristics of object A from a pre-established timbre library based on the object characteristics of object A. The computing device 110 can generate a speech segment corresponding to the dialogue text based on the selected timbre. The computing device 110 can perform the above operations for each text, thereby determining a matching timbre for the object of the dialogue text and further generating an audio segment corresponding to the text. Thus, the generated video can present the text content to the user in the form of a dialogue.

[0140] After generating corresponding audio segments for each piece of text, the computing device 110 can acquire multiple audio segments 350. Further, the computing device 110 can determine the audio length corresponding to each audio segment. Moreover, the determined audio length of the audio segment can be less than or equal to the frame duration of the corresponding image frame in the generated video. That is, the frame duration of the corresponding image frame in the generated video can be greater than or equal to the audio length of the audio segment. For example, if the audio length of the audio segment is t1, then the frame duration (i.e., display duration) of the corresponding image frame in the generated video can be greater than or equal to t1.

[0141] The computing device 110 can also sequentially acquire each image frame corresponding to each text in the order of the text within the multiple texts. Thus, the order of the acquired image frames corresponds to the order of the texts. The computing device 110 can also associate each image frame with each corresponding audio segment. Using the example above as an illustration, the computing device 110 can determine that audio segment Audio 1 corresponds to image frame Image 1, and the computing device 110 can associate image frame Image 1 with audio segment Audio 1. Furthermore, the frame duration (display duration) of each image frame in the generated video can be greater than or equal to the audio length of the associated audio segment.

[0142] Therefore, by associating each text, the corresponding generated image frame, and the audio segment converted from the text, a video 160 (e.g., a narration video or a browsing video) corresponding to the text 150 can be generated.

[0143] In some embodiments, when text is displayed in an image frame, at least one of the multiple texts can be split into multiple sub-texts, and the split sub-texts can be displayed in the image frames corresponding to the at least one text at different frame times. For example, for texts with a relatively large number of words, considering the aesthetics of the generated video images, the text can be split into multiple sub-texts, so that each sub-text can include a smaller number of words. In some embodiments, the text can be split according to the delimiters (e.g., commas) included in the text. This can be understood by referring to the specific implementation methods described above; for simplicity, it will not be repeated here.

[0144] Furthermore, in some embodiments, the computing device 110 can also select at least one image frame from multiple image frames corresponding to multiple text snippets. For example, the selected image frame can be a keyframe among the generated multiple image frames. The computing device 110 can move one or more selected image frames. For example, the computing device 110 can move one or more selected image frames vertically (e.g., from top to bottom or from bottom to top), horizontally (e.g., from left to right or from right to left), etc. And the movement time of each image frame is no greater than the frame duration of that image frame. By selecting and moving the selected image frames, the generated video can exhibit dynamic effects, thereby attracting more user attention and improving the user experience.

[0145] In some embodiments, for each text, the computing device 110 may also determine the object associated with the narration text if the text includes a voice-over text. If the object associated with the narration text is a person object, the computing device 110 may generate a video clip associated with the image frame based on the content of the text and the corresponding image frame of the text. For example, the computing device 110 may input the image frame corresponding to the text and the text into a model (e.g., a video generation model in the model) to generate a video clip corresponding to the image frame. For example, if the narration text is "Xiaoming stood up from his seat," and the object of the narration text is the person "Xiaoming," to enhance the dynamic effect of the video, the image frame corresponding to the text and the text may be input into a model, and the model may generate a video clip. For this example text, the generated video clip may show the dynamic process of the person ("Xiaoming") moving from a sitting position to standing up.

[0146] Furthermore, to further enhance the effect of the generated video, as shown in Figure 3, the computing device 110 can also select background music 370 for the video 160 and associate the background music 370 with the video 160. Thus, when the user plays the video 160, in addition to seeing the text, the generated image frames, and hearing the audio clips of the text, they can also hear the background music 370. In some embodiments, the computing device 110 can select background music based on the text 150, and the audio length of the background music corresponds to the duration of the video. In some embodiments, the computing device 110 can select background music based on the content of the text 150. For literary novels, the computing device 110 can select soothing music as background music; for comedic novels, the computing device 110 can select upbeat music as background music, and so on.

[0147] Furthermore, as shown by the dashed line in Figure 3, the computing device 110 can also select background music 370 based on the modified text 320 (320'), and use the background music 370 with multiple image frames 340 and multiple audio segments 350 to generate video 160.

[0148] Therefore, the method for generating video according to the embodiments of this disclosure can adapt written works to quickly generate concise, visually appealing, and engaging videos (e.g., narration videos or browsing videos), significantly shortening video generation time while improving the quality of the generated videos, enabling users to experience audiovisual experiences beyond reading, and further enhancing the user experience.

[0149] Figure 7 shows a schematic block diagram of an example apparatus 700 according to some embodiments of the present disclosure. Apparatus 700 can be implemented by software, hardware, or a combination of both. As shown in Figure 7, apparatus 700 includes a determining module 710, a first generating module 720, and a second generating module 730.

[0150] In some embodiments, the determining module 710 can determine multiple text snippets based on the text. The first generating module 720 can generate an image frame corresponding to each text snippet based on the object features of each text snippet and the object associated with each text snippet. The second generating module 730 can generate a video corresponding to the text snippet based on each text snippet and each image frame.

[0151] The device 700 in Figure 7 can be used to implement the process described above in conjunction with Figures 1 to 6B, which will not be repeated here for the sake of brevity.

[0152] The division of modules or units in the embodiments of this disclosure is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods. Furthermore, the functional units in the disclosed embodiments may be integrated into one unit, exist as separate physical entities, or two or more units may be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional unit.

[0153] Figure 8 shows a block diagram of an example device 800 that can be used to implement embodiments of the present disclosure. It should be understood that the device 800 shown in Figure 8 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. For example, device 800 can be used to correspond to computing device 110 described herein in conjunction with Figure 1 and can be used to perform the processes described above in Figures 1 through 7.

[0154] As shown in Figure 8, device 800 is in the form of a general-purpose computing device. Components of computing device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage devices 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit 810 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 800.

[0155] Computing device 800 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof). Storage device 830 can be removable or non-removable media and may include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within computing device 800.

[0156] The computing device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG8, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.

[0157] The communication unit 840 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 800 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0158] Input device 850 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 800 can also communicate with one or more external devices (not shown) via communication unit 840 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 800, or with any device that enables computing device 800 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0159] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.

[0160] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0161] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0162] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0163] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0164] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating video, the method comprising: Multiple copy pieces are determined based on the text; Based on the object features of each of the multiple texts and the objects associated with each text, an image frame corresponding to each text is generated; as well as Based on each text and the corresponding image frame, a video corresponding to the text is generated.

2. The method according to claim 1, wherein the object includes at least one of a person, animal, plant, object, and scene.

3. The method according to claim 1, further comprising: Extract the object features of at least one object from the text.

4. The method of claim 3, wherein the object feature of each of the at least one object includes the object name of the object, and the object feature has at least one version.

5. The method of claim 4, wherein the object feature includes a plurality of feature elements, and wherein at least one corresponding feature element is different in different versions of the object feature.

6. The method of claim 1, wherein the object features are stored in at least one of the following ways: A model file of the object model generated based on the object features; or A language description used to describe the characteristics of the object.

7. The method of claim 1, wherein determining multiple text entries based on text comprises: Based on the text, obtain the modified text corresponding to the text; as well as Based on the modified text, the multiple text entries are determined.

8. The method according to claim 7, wherein obtaining the modified text corresponding to the text comprises: Based on the text and the prompt information used to generate the modified text, the modified text corresponding to the text is generated, wherein the prompt information provides a means for generating the modified text. The information describes the text to be modified, and the prompt information includes the modification criteria for modifying the text.

9. The method of claim 7, wherein determining the plurality of texts based on the modified text comprises: Identify multiple segmentation symbols in the modified text; The modified text is segmented based on the multiple segmentation symbols to obtain multiple segmented texts; as well as Each piece of text in the multiple segmented texts is determined based on the number of characters in each segmented text.

10. The method of claim 9, wherein determining each of the plurality of texts comprises: In response to the fact that the number of characters in each segmented text is not greater than a threshold number, each segmented text is treated as a single text. or In response to the fact that the number of characters in each segmented text is greater than the threshold number, the set of characters in the segmented text is segmented at least once, and multiple subsets of the characters after the at least one segmentation are respectively assigned to multiple texts.

11. The method of claim 7, wherein obtaining the modified text corresponding to the text comprises: Based on the text and the prompt information for generating the modified text, the modified text corresponding to the text is generated, wherein the prompt information provides information for generating the modified text, and the prompt information includes at least one of the following: The indicator model divides the text into multiple scenes based on scene objects within the text; The model is instructed to acquire the object features of multiple character objects in the text and the relationships between different character objects; The model is instructed to extract multiple dialogue segments from the text and determine at least one association feature for each dialogue segment, wherein, for each dialogue segment, the at least one association feature includes at least one of the following: the dialogue occurred. The scene, the person who initiates the conversation, the emotions associated with the conversation, or the content of the conversation; The model is instructed to process the text to extract multiple narration segments; or The model is instructed to combine one or more of the following: the multiple scenes, the object features of the multiple character objects, the correlation, the multiple dialogues, or the multiple narrations, and to use the combination result as the modified text.

12. The method of claim 11, wherein the modified text includes a plurality of scene scripts corresponding to the plurality of scenes respectively, and each scene script includes at least one dialogue and / or at least one narration, and the plurality of scene scripts are sorted in the modified text according to the order in which the corresponding scenes appear in the text.

13. The method of claim 12, wherein determining the plurality of texts based on the modified text comprises: For each of the plurality of scene scripts, perform the following operations: In response to the inclusion of narration in the scene script, the narration is segmented based on the object associated with the narration to obtain at least one narration text; In response to the dialogue included in the scene script, the dialogue is segmented based on the object that initiates the dialogue to obtain at least one dialogue text. as well as Each of the at least one dialogue text and the at least one narration text shall be used as each of the multiple texts.

14. The method of claim 1, wherein generating an image frame corresponding to each of the plurality of texts based on object features of each text and an object associated with each text includes: Based on each of the texts, a first prompt message corresponding to each of the texts is generated, and the first prompt message provides prompt information for generating the corresponding image frame; Obtain the object characteristics of the object associated with the first prompt information; as well as Based on the first prompt information and the object characteristics, an image frame corresponding to each piece of text is generated.

15. The method of claim 14, wherein generating a first prompt message corresponding to each piece of text based on each piece of text includes: Based on each piece of text and the received second prompt information, a first prompt information corresponding to each piece of text is generated, wherein the second prompt information provides information for generating the first prompt information, and wherein the second prompt information includes at least one of the following: The output format of the first prompt message; Generate the operation to be performed in the first prompt message; The rules for generating the first prompt message; Regarding the constraints for generating the first prompt message; or Example used to generate the first prompt message.

16. The method of claim 14, wherein obtaining the object characteristics of the object associated with the first prompt information includes: Based on the first prompt information, determine the object name of the object included in the first prompt information; as well as Based on the object name, obtain the object characteristics associated with the object name.

17. The method according to claim 16, wherein generating the image frame corresponding to each text message based on the first prompt information and the object features, comprising: The object features are attached to the first prompt information to obtain the attached first prompt information; as well as The attached first prompt information is input into the image generation model, so that the image generation model generates the image frame corresponding to each piece of text.

18. The method according to claim 1, further comprising: For each piece of copy, perform the following operations: In response to the fact that the text is narration text, identify the objects associated with the narration text; In response to the fact that the object associated with the narration is a character object, a video clip associated with the image frame is generated based on the content of the narration and the corresponding image frame.

19. The method according to claim 1, further comprising: Select at least one image frame from the multiple image frames corresponding to the multiple texts; as well as Move the selected at least one image frame. The time for each of the at least one selected image frames to move is no greater than the frame time length of each selected image frame.

20. The method of claim 1, wherein at least one of the plurality of texts can be split into a plurality of sub-texts, and each of the plurality of sub-texts is displayed in an image frame corresponding to the at least one text at a different frame time.

21. The method according to claim 1, further comprising: Each text message is converted into a corresponding audio segment, wherein the audio segment corresponds to the image frame; as well as Determine the audio length corresponding to the speech segment.

22. The method of claim 21, wherein converting each text message into a corresponding audio segment includes: For each piece of copy, perform the following operations: In response to the fact that the text is a dialogue text, the object that initiated the dialogue text is determined; Based on the object's characteristics, select the corresponding timbre from a preset repository; as well as Based on the timbre, a speech segment corresponding to the dialogue text is generated.

23. The method of claim 21, further comprising: According to the order of the texts in the multiple texts, the image frames corresponding to each text are obtained sequentially; as well as Associate the image frame with the corresponding audio segment. The frame time length of the image frame is greater than or equal to the audio length corresponding to the speech segment.

24. The method of claim 23, further comprising: Background music is selected based on the text, and the audio length of the background music corresponds to the duration of the video.

25. The method of claim 1, wherein the text comprises fictional text.

26. An electronic device comprising: At least one processing unit; At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 25 when executed by the at least one processing unit.

27. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 25.

28. A computer program product having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 25.