Video script generation method, video generation method and related device

CN118945414BActive Publication Date: 2026-08-11BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-10
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

但是,广大的中小型用户因缺灵感、缺制作能力、缺预算、缺资源等等原因无法低成本且高效地生成满足用户需求的视频文案或者视频

Benefits of technology

[0029] As can be seen, the video script generation method and related equipment provided in this disclosure can determine the video structure information of a target video based on user-input promotional information using a pre-established video structure library, and then directly and quickly generate the video script of the target video based on the determined video structure information and promotional information using a pre-established corpus, thereby improving the user experience. Furthermore, since the video script of the target video in the embodiments of this disclosure is generated using a pre-established video structure library and corpus, it can meet the video delivery requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118945414B_ABST
    Figure CN118945414B_ABST
Patent Text Reader

Abstract

This disclosure provides a method for generating video scripts, including: obtaining promotional information corresponding to a target video; generating video structure information of the target video based on the promotional information; and generating video scripts for the target video based on the video structure information and the promotional information. Based on the above video script generation method, this disclosure also provides a video generation method, a video script generation device, a video generation apparatus, an electronic device, a storage medium, and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of Internet technology, and in particular to a video script generation method, a video generation method, a video script generation device, a video generation device, an electronic device, a storage medium, and a program product. Background Technology

[0002] Currently, with the rapid development of short video applications, users generally have a wide range of needs for video generation and distribution on short video content platforms. However, many small and medium-sized users are unable to generate video scripts or videos that meet user needs in a low-cost and efficient manner due to a lack of inspiration, production capabilities, budget, resources, and other reasons. Summary of the Invention

[0003] In view of this, embodiments of this disclosure provide a video script generation method that can directly and quickly generate video scripts that meet user needs based on promotional information input by the user, thereby improving the user experience. Furthermore, embodiments of this disclosure also provide a video generation method that can directly generate videos that meet user needs based on promotional information input by the user and original materials uploaded by the user, thereby greatly improving the user experience.

[0004] According to some embodiments of this disclosure, the above-described video script generation method may include: obtaining promotional information corresponding to a target video; generating video structure information of the target video based on the promotional information; wherein the video structure information includes the type and duration information of at least one video segment contained in the target video; and generating video script of the target video based on the video structure information and the promotional information.

[0005] In embodiments of this disclosure, generating video structure information of the target video based on the promotional information includes: determining the tags corresponding to the target video based on the promotional information; and selecting video structure information matching the tags of the target video from a pre-established video structure library as the video structure information of the target video; wherein the video structure library stores video structure information of various video segmentation combinations and their corresponding tags.

[0006] In embodiments of this disclosure, generating video structure information of the target video based on the promotional information includes: generating video structure information of the target video based on a video structure generation model and the promotional information; wherein the video structure generation model is trained and generated based on a pre-established video structure library; and the video structure library stores video structure information of various video segmentation combinations and their corresponding tags.

[0007] In embodiments of this disclosure, the method further includes: acquiring multiple videos to be processed within a predetermined time interval; performing the following operations for each video to be processed: determining the tag corresponding to the video to be processed based on the feature information of the product associated with the video to be processed, target user information, time node information, and delivery target information; deconstructing the video to be processed to obtain video structure information corresponding to the video to be processed; and adding the information pair consisting of the video structure information corresponding to the video to be processed and the tag corresponding to the video to be processed to the video structure library.

[0008] In embodiments of this disclosure, generating the video script for the target video based on the video structure information and the promotional information includes: determining the tags corresponding to the target video based on the promotional information; for each video segment in the at least one video segment, selecting corpus information corresponding to the video segment from a pre-established corpus based on the type and duration information of the video segment; wherein the corpus stores multiple corpus information corresponding to each video segment and their corresponding tags; and combining the corpus information corresponding to each video segment to obtain the video script for the target video.

[0009] In embodiments of this disclosure, generating video script for the target video based on the video structure information and the promotional information includes: determining tags corresponding to the target video based on the promotional information; for each video segment in the at least one video segment, selecting corpus information corresponding to the video segment from a pre-established corpus based on the type and duration information of the video segment; wherein the corpus stores multiple corpus information corresponding to each video segment and their corresponding tags; expanding the corpus information corresponding to each video segment based on a script generation model to obtain script information corresponding to each video segment; wherein the script generation model is trained and generated based on the corpus; and combining the script information corresponding to each video segment to obtain the video script for the target video.

[0010] In embodiments of this disclosure, generating video script for the target video based on the video structure information and the promotional information includes: for each video segment in the at least one video segment, generating script corresponding to each video segment based on the type and duration information of the video segment using a script generation model and the promotional information; wherein the script generation model is trained and generated based on the corpus; and combining the scripts corresponding to each video segment to obtain the video script for the target video.

[0011] In embodiments of this disclosure, the method further includes: acquiring at least one original material corresponding to the target video; and determining features corresponding to the at least one original material; wherein generating text corresponding to each video segment based on the text generation model and the promotional information includes: generating text corresponding to each video segment based on the text generation model, the promotional information, and features corresponding to the at least one original material.

[0012] Based on the above-described video script generation method, embodiments of this disclosure also provide a video generation method, comprising: acquiring promotional information corresponding to a target video and at least one original material; generating video structure information of the target video based on the promotional information; wherein the video structure information includes the type and duration information of at least one video segment contained in the target video; generating video script of the target video based on the video structure information and the promotional information; and editing the at least one original material based on the video structure information, the video script, and the promotional information to obtain an edited video file.

[0013] In embodiments of this disclosure, editing the at least one original material based on the video structure information, the video script, and the promotional information includes: editing the at least one original material according to the video structure information, the quantity, length, and quality of the at least one original material to obtain at least one backup material; determining the display subject of each of the at least one backup material; editing the at least one backup material based on the display subject of the at least one backup material, the promotional information, and the video structure information to obtain a video corresponding to each video segment; and merging the videos corresponding to each video segment to obtain the edited video file.

[0014] In embodiments of this disclosure, editing the at least one original material based on the video structure information, the video text, and the promotional information further includes: performing the following operations on the video corresponding to each video segment: performing video content recognition on the video corresponding to the video segment to obtain a first recognition result; performing text content recognition on the video text corresponding to the video segment to obtain a second recognition result; and editing the video of the video segment based on the first recognition result and the second recognition result.

[0015] In embodiments of this disclosure, the video generation method may further include: generating a voiceover for the target video based on the video text and the promotional information; and adding the voiceover to the edited video file to generate the target video.

[0016] In embodiments of this disclosure, generating the voiceover for the target video based on the video text and the promotional information includes: determining the tag corresponding to the target video according to the promotional information; selecting a target virtual voice that matches the tag corresponding to the target video from a pre-set set of virtual voices based on the tag corresponding to the target video; and generating the voiceover for the target video based on the voice pack of the video text and the target virtual voice.

[0017] Based on the above method, embodiments of this disclosure provide a video script generation device, which may include:

[0018] The information acquisition module is used to acquire promotional information corresponding to the target video;

[0019] A video structure information determination module is used to generate video structure information of the target video based on the promotional information; wherein, the video structure information includes the type and duration information of at least one video segment contained in the target video; and

[0020] The video script generation module is used to generate video scripts for the target video based on the video structure information and the promotional information.

[0021] Based on the above method, embodiments of this disclosure provide a video generation apparatus, which may include:

[0022] The information acquisition module is used to acquire promotional information corresponding to the target video;

[0023] A video structure information determination module is used to generate video structure information of the target video based on the promotion information; wherein, the video structure information includes the type and duration information of at least one video segment contained in the target video;

[0024] The video script generation module is used to generate video scripts for the target video based on the video structure information and the promotional information; and

[0025] The video editing module is used to edit at least one original material based on the video structure information, the video script, and the promotional information to obtain an edited video file.

[0026] Furthermore, embodiments of this disclosure also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described video text generation method or video generation method.

[0027] Embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the above-described video text generation method or video generation method.

[0028] Embodiments of this disclosure also provide a computer program product, including computer program instructions that, when executed on a computer, cause the computer to perform the aforementioned video text generation method or video generation method.

[0029] As can be seen, the video script generation method and related equipment provided in this disclosure can determine the video structure information of a target video based on user-input promotional information using a pre-established video structure library, and then directly and quickly generate the video script of the target video based on the determined video structure information and promotional information using a pre-established corpus, thereby improving the user experience. Furthermore, since the video script of the target video in the embodiments of this disclosure is generated using a pre-established video structure library and corpus, it can meet the video delivery requirements.

[0030] Furthermore, the video generation method and related equipment provided in this disclosure can directly and quickly generate target videos that meet user needs based on the video text of the target video, the promotional information input by the user, and the original materials uploaded by the user, thereby further improving the user experience. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 The internal structure of the video generation system described in some embodiments of this disclosure is shown.

[0033] Figure 2 The implementation flow of the video text generation method described in some embodiments of this disclosure is shown.

[0034] Figure 3 The implementation flow of the method for generating video structure information of a target video based on promotional information, as described in some embodiments of this disclosure, is shown.

[0035] Figure 4 The implementation flow of the method for generating video text for a target video based on video structure information and promotional information, as described in some embodiments of this disclosure, is shown.

[0036] Figure 5 The implementation flow of the method for generating video text for a target video based on video structure information and promotional information, as described in other embodiments of this disclosure, is shown.

[0037] Figure 6 The internal structure of the video text generation engine 110 described in some embodiments of this disclosure is shown.

[0038] Figure 7 This disclosure illustrates the implementation flow of a method for generating video text for a target video based on video structure information and promotional information, as described in some embodiments of the present disclosure.

[0039] Figure 8 The implementation flow of the method for automatically building a video structure library according to an embodiment of this disclosure is shown.

[0040] Figure 9 The implementation flow of the video generation method described in the embodiments of this disclosure is shown.

[0041] Figure 10 The implementation flow of the method for editing at least one original material based on video structure information, video script, and promotional information, as described in the embodiments of this disclosure, is shown.

[0042] Figure 11 The implementation flow of the method for generating voice-over for a target video based on video text and promotional information, as described in an embodiment of this disclosure, is shown.

[0043] Figure 12 The internal structure of the video text generation apparatus described in some embodiments of this disclosure is shown.

[0044] Figure 13 The internal structure of the video generation apparatus described in some embodiments of this disclosure is shown.

[0045] Figure 14 A schematic diagram of a more specific electronic device hardware structure described in some embodiments of this disclosure is shown. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0047] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0048] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.

[0049] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.

[0050] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0051] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0052] As mentioned earlier, currently, most small and medium-sized users on short video content platforms are unable to generate video scripts or videos that meet their needs in a low-cost and efficient manner. Furthermore, it is understood that in various short video applications, the video script is a crucial component. Whether the video script is attractive to users is one of the key factors influencing the popularity of a short video. Therefore, this disclosure provides a video generation system. On the one hand, this system can directly and quickly generate video scripts that meet user needs based on user-input promotional information, thereby improving the user experience. On the other hand, the aforementioned video generation system can also directly generate videos that meet user needs based on user-input promotional information and user-uploaded original materials, further enhancing the user experience.

[0053] Figure 1 The internal structure of the video generation system described in this disclosure embodiment is shown. For example... Figure 1 As shown, the video generation system described above may include: a video structure library 102, a corpus 104, an interface module 106, a video structure generation engine 108, a video script generation engine 110, and a video editing engine 112.

[0054] In the embodiments of this disclosure, the video structure library 102 is used to store video structure information and corresponding tags for combinations of various video segments. The video structure information includes the type and duration information of at least one video segment contained in a video.

[0055] Specifically, in the embodiments of this disclosure, various video segment types can be pre-determined based on the video's function and content. For example, based on the video's function, three types of video segments can be set, including: pre-roll segment, main video segment, and post-roll segment. The pre-roll segment typically refers to the portion of the video spliced ​​before the main video; while the post-roll segment typically refers to the portion of the video spliced ​​after the main video. Furthermore, based on the video's content, various video segment types can be further set for the main video segment, including: eye-catching copywriting segment, product introduction segment, product description segment, benefit description segment, and guidance information segment, etc. By analyzing and deconstructing the visuals, scenes, background music, and copywriting of a large number of videos on short video content platforms, each video can typically be decomposed into a combination of one or more of the above-mentioned video segment types, thereby obtaining the video structure information of each video. For example, by analyzing the visuals, scenes, background music, copywriting, and other content of a video, this video can be decomposed into three video segments: pre-roll segment, main video segment, and post-roll segment. Furthermore, the main video segment can be further broken down into four parts: product introduction segment, product description segment, benefit description segment, and guidance information segment. This yields the video structure: intro segment + product introduction segment + product description segment + benefit description segment + guidance information segment + outro segment. In addition to the above video structure, the duration of each video segment can also be obtained. Furthermore, the combined structure of the video segments and the duration of each segment can be used to obtain the overall video structure information. In other words, the video structure information of a video defines the type of at least one video segment and the duration of each segment. Thus, given a video structure, the type and duration of each video segment can be determined, essentially providing a structural template for the video.

[0056] Furthermore, in the embodiments of this disclosure, each video on a short video content platform typically corresponds to one or more tags. These tags can usually correspond to different dimensions, describing a video from different perspectives. In the embodiments of this disclosure, the dimensions of these tags can include: product feature dimensions, target user dimensions, time node dimensions, and targeting dimensions, etc. The product feature dimension typically includes the product name, product category, and product selling points, etc. For example, a video related to men's waterproof athletic shoes could include the following tags: Product Name: ABC; Product Category: Men's Shoes; Product Selling Point: Waterproof; Target User: Young Users; and Target: Product Promotion, etc. Additionally, the video can also correspond to a time node, such as a holiday or solar term, etc.

[0057] Based on the above configuration, in the embodiments of this disclosure, the video structure library 102 stores various combined structures of video segments and one or a set of tags corresponding to each combined structure of video segments. It should be noted that the meanings of the tags corresponding to each combined structure of video segments are consistent with the meanings of the tags corresponding to the videos, and will not be repeated here. Since the tags can contain multiple tags across multiple dimensions, the one or a set of tags corresponding to the combined structure of video segments can also be called a tag matrix.

[0058] In some embodiments of this disclosure, the video structure library 102 can be manually created by the creative team developing and designing the video. In other embodiments, the video structure library 102 can also be automatically created based on videos already published on short video content platforms. In still other embodiments, the video structure library 102 can be created using a combination of manual and automatic methods. Specific methods for automatically creating the video structure library 102 will be described in detail later and will be omitted here.

[0059] In the embodiments of this disclosure, the corpus 104 is used to store multiple corpora corresponding to each video segment and their corresponding tags.

[0060] Specifically, in the embodiments of this disclosure, the available corpus and the corresponding tags for each video segment can be pre-recorded according to their characteristics, thereby establishing the aforementioned corpus 104. Based on the aforementioned corpus 104, the text for each video segment can be generated according to the promotional information. It should be noted that the meanings of the tags corresponding to the corpus and the tags corresponding to the videos are basically the same, and will not be repeated here. Alternatively, in addition to the dimensions of video tags such as product feature dimension, target user dimension, time node dimension, and placement target dimension, the tags corresponding to the corpus can also include a video structure dimension to identify the type of video segment corresponding to the corpus. That is, for the same product feature dimension, target user dimension, time node dimension, and placement target dimension tags, different video segment types can also correspond to different corpora. In some embodiments of this disclosure, the aforementioned corpus 102 can be manually established by the creative team of video development and design.

[0061] In the embodiments of this disclosure, the interface module 106 is mainly used to receive promotional information input by the user or to receive original materials uploaded by the user, and also to output video scripts or videos generated by the system to the user.

[0062] In the embodiments of this disclosure, the video structure generation engine 108 is mainly used to determine the video structure of the target video based on the promotional information input by the user. The video script generation engine 110 is mainly used to generate video scripts based on the promotional information input by the user and the video structure information of the target video output by the video structure generation engine 108. The video editing engine 112 is mainly used to generate a video based on the original materials uploaded by the user, the video structure information of the target video output by the video structure generation engine 108, and the video script output by the video script generation engine 110. The specific operation methods of the video structure generation engine 108, the video script generation engine 110, and the video editing engine 112 will be described in detail later and will be omitted here.

[0063] In the embodiments of this disclosure, the video generation system described above can directly and quickly generate video scripts or videos that meet user needs based on user-input promotional information and original materials, thereby greatly improving the user experience.

[0064] The specific operation process of the above-mentioned video structure generation engine 108, video text generation engine 110, and video editing engine 112 will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0065] Figure 2 The implementation flow of the video text generation method described in the embodiments of this disclosure is illustrated. It should be noted that this method can be implemented collaboratively by the aforementioned video structure generation engine 108 and the aforementioned video text generation engine 110. Typically, the aforementioned video structure generation engine 108 and the aforementioned video text generation engine 110 may include hardware devices with data processing capabilities and / or the necessary software to drive the operation of such hardware devices. Figure 2 As shown in the embodiments of this disclosure, the video script generation method may include the following steps:

[0066] In step 202, the promotional information corresponding to the target video is obtained.

[0067] In step 204, the video structure information of the target video is generated based on the promotional information described above.

[0068] In step 206, the video script for the target video is generated based on the video structure information and the promotional information described above.

[0069] The following section will provide a detailed explanation of each step in the above video script generation method, with reference to the accompanying diagrams and specific examples.

[0070] Regarding step 202 above, in the embodiments of this disclosure, the target video refers to the video that the user wishes to be generated by the video generation system based on promotional information and original materials.

[0071] Furthermore, as mentioned above, the main objective of this disclosure is to provide a method for directly generating video scripts that meet user needs based on promotional information provided by users, thereby serving users of short video content platforms. Therefore, in the embodiments of this disclosure, the generated video scripts are typically used to promote a specific target product of the short video content platform user. In other words, the generated video scripts will generally correspond to a specific target product.

[0072] Based on the above, in the embodiments of this disclosure, the promotional information corresponding to the target video may include one or more of the following: characteristic information of the target product, target user information, time node information, and target information for ad placement. The characteristic information of the target product typically includes: the product name, product category, and product selling points, etc.

[0073] In some embodiments of this disclosure, the promotional information corresponding to the target video can be information input by the user through the interface module 106. In specific applications, the interface module 106 can provide the user with a graphical user interface and offer multiple options for promotional information across different dimensions to facilitate the user's input of the corresponding promotional information. For example, the interface module 106 can provide the user with input boxes or drop-down menus for the product name, product category, product selling points, target user information, time node information, and target information for information input or selection.

[0074] Furthermore, in some other embodiments of this disclosure, assuming that a user has already opened a shop on a short video content platform and the target product is a certain item sold in that shop, then the step 202 above, which involves obtaining the promotional information corresponding to the target video, may include: first, obtaining the identifier corresponding to the target product on the short video content platform; and then, determining the promotional information corresponding to the target video based on the identifier. Specifically, through the identifier, the video structure generation engine 108 can obtain product information corresponding to the identifier from the short video content platform, and further extract one or more of the following information from the product information: product name, product category, product selling points, target users, and targeting. Alternatively, through the identifier, the video structure generation engine 108 can obtain a product detail image corresponding to the identifier from the short video content platform, and extract one or more of the following information from the product detail image: product name, product category, product selling points, target users, and targeting. In addition, the time node information in the promotional information can be estimated based on the current time.

[0075] Regarding step 204 above, in some embodiments of this disclosure, the specific implementation method for generating the video structure information of the target video based on the aforementioned promotional information can be as follows: Figure 3 As shown, the main steps include the following:

[0076] In step 302, the tags corresponding to the target video are determined based on the above promotional information.

[0077] In step 304, video structure information that matches the tags of the target video is selected from the pre-established video structure library as the video structure information of the target video.

[0078] As mentioned above, in the embodiments of this disclosure, the promotional information may include one or more of the following: characteristic information of the target product corresponding to the target video, target user information, time node information, and delivery target information. The characteristic information of the target product typically includes information such as the product name, product category, and product selling points. Furthermore, in the embodiments of this disclosure, the tags corresponding to the video may correspond to different dimensions. Specifically, the dimensions of the tags may include: product characteristic dimension, target user dimension, time node dimension, and delivery target dimension, etc. Thus, in the embodiments of this disclosure, the dimensions of the information contained in the promotional information are essentially consistent with the dimensions of the tags. Therefore, in step 302 above, the tags of the target video in one or more dimensions can be directly determined based on the promotional information.

[0079] Furthermore, as mentioned above, in the embodiments of this disclosure, the video structure library stores video structure information and corresponding tags for various combinations of video segments. Thus, in step 304, the tags corresponding to the target video can be matched with the tags corresponding to the various combinations of video segments stored in the video structure library, and the video structure information that best matches the tags of the target video can be selected as the video structure information of the target video. In a specific example, when performing tag matching, corresponding weight values ​​can be pre-set for tags of different dimensions, where the weight values ​​can be the same or different; then, scores are assigned based on the similarity of the tags corresponding to the target video and the tags corresponding to the combinations of various video segments in each dimension, obtaining the similarity values ​​of the two tags in each dimension; next, the similarity values ​​of the two tags in each dimension are weighted and summed to obtain the similarity between the target video and the combinations of various video segments; finally, one or more combinations of video segments with the highest similarity to the target video are selected from the combinations of various video segments, and their corresponding video structure information is used as the video structure information of the target video. As mentioned earlier, the video structure information defines the types of one or more video segments contained in the target video, as well as the duration of each video segment. For example, a single piece of video structure information can define the target video as including five video segments: a 5-second introductory segment, a 10-second product introduction segment, a 35-second product description segment, a 5-second benefit description segment, and a 5-second guidance information segment.

[0080] In other embodiments of this disclosure, in step 204 above, video structure information of the target video can be generated based on the video structure generation model and the aforementioned promotional information. In these embodiments, the video structure generation model can be trained based on the aforementioned pre-established video structure library. As mentioned earlier, the video structure library stores video structure information and its corresponding tags for various combinations of video segments. Therefore, the video structure generation model can be obtained by training the machine learning model using the video structure information and its corresponding tags stored in the video structure library as training data. The video structure generation model can output the video structure information with the highest matching degree to the input promotional information based on the input promotional information.

[0081] Regarding step 206 above, in some embodiments of this disclosure, the method for generating video scripts for target videos based on video structure information and promotional information can be specifically as follows: Figure 4 As shown, it includes the following steps:

[0082] In step 402, the tags corresponding to the target video are determined based on the above promotional information.

[0083] In step 404, for each video segment included in the above video structure information, corpus matching the tags of the target video is selected from the pre-established corpus according to the type and duration information of the video segment as the corpus information corresponding to the video segment.

[0084] In step 406, the corpus information corresponding to each of the above video segments is combined to obtain the video text of the target video.

[0085] Specifically, the implementation method of step 402 above can be referred to step 302 above, and will not be repeated here.

[0086] Furthermore, as mentioned earlier, the corpus stores multiple corpora corresponding to each video segment and their corresponding tags. Therefore, regarding step 404, for a given video segment, the tags corresponding to the target video can be matched with the tags corresponding to multiple corpora under that video segment stored in the corpus, and the corpus whose tags best match the target video's tags can be selected as the corpus information corresponding to the video segment. Of course, in practical applications, the type of the video segment can also be directly used as the video structure tag of the corpus and matched with the tags of the corpus in the corpus along with other tags. Further, during the matching process, corpora that are too long or too short can be pre-filtered based on the duration information, ensuring that the duration of the corpora to be matched roughly matches the duration information of the video segments. In a specific example, during matching, corresponding weight values ​​can be pre-set for labels in different dimensions. Then, scores are assigned based on the similarity of the labels corresponding to the target video and the labels corresponding to each corpus in each dimension, resulting in the similarity value of the two labels in each dimension. Next, the similarity values ​​of the two labels in each dimension are weighted and summed to obtain the similarity between the target video and each corpus. Finally, one or more corpora with the highest similarity to the target video are selected from multiple corpora and used as the corpus information corresponding to the above video segment.

[0087] As can be seen, the above-mentioned video script generation method can generate video scripts based on a pre-established corpus, and can achieve the goal of quickly generating video scripts that meet user needs.

[0088] Regarding step 206 above, in some other embodiments of this disclosure, the method for generating video scripts for target videos based on video structure information and promotional information may be as follows: Figure 5 As shown, it includes the following steps:

[0089] In step 502, for each video segment included in the above video structure information, a corresponding text is generated for each video segment based on the text generation model and the above promotional information, according to the type and duration information of the video segment.

[0090] In step 504, the text corresponding to each of the above video segments is combined to obtain the video text of the target video.

[0091] Regarding step 502 above, in the embodiments of this disclosure, the copywriting generation model is trained and generated based on the aforementioned corpus. As mentioned earlier, the corpus stores multiple corpora corresponding to each video segment and their corresponding tags. Therefore, by using the multiple corpora corresponding to each video segment and their corresponding tags stored in the corpus as training data to further train the large language model, the copywriting generation model can be obtained. The trained copywriting generation model not only possesses the ability of a large language model for natural language understanding and generating natural language text, but also has the ability to intelligently generate video copy suitable for product promotion based on the user's promotional needs (product characteristics, target users, promotional goals, etc.). Therefore, in the embodiments of this disclosure, the copywriting for each video segment of the target video can be directly generated based on promotional information through the copywriting generation model, and then combined to obtain the complete video copy for the target video, achieving the goal of quickly generating video copy that meets user needs.

[0092] In some embodiments of this disclosure, the video scripts obtained in step 504 can typically be multiple. In practical applications, post-processing can be performed on the multiple video scripts obtained after step 504. This post-processing may include, on the one hand, verifying whether the multiple video scripts meet the requirements of the video structure information and filtering out those that do not meet the requirements; on the other hand, it may also include predicting the popularity of each of the multiple video scripts and selecting one or more video scripts with higher predicted popularity values ​​as the final video script output to the user. The popularity prediction operation can be implemented using a trained video script popularity prediction model, which will not be detailed here.

[0093] Figure 6 An example of the internal structure of a video text generation engine 110 described in some embodiments of this disclosure is shown. For example... Figure 6 As shown, in some embodiments of this disclosure, the video text generation engine 110 may include: a text generation model 602 and a text post-processing module 604.

[0094] The aforementioned copywriting generation model 602 can be trained based on the established corpus 104, and after receiving promotional information and video structure information input by the user, it outputs multiple alternative video copywriting options.

[0095] The above-mentioned copy post-processing module 604 can post-process the multiple alternative video copy output by the above-mentioned copy generation model 602 to obtain one or more video copy that can be finally output to the user.

[0096] Regarding step 206 above, in some other embodiments of this disclosure, the method for generating video scripts for target videos based on video structure information and promotional information can be specifically as follows: Figure 7 As shown, it includes the following steps:

[0097] In step 702, the tags corresponding to the target video are determined based on the promotion information.

[0098] In step 704, for each video segment included in the above video structure information, corpus matching the tags of the target video is selected from the pre-established corpus according to the type and duration information of the video segment as the corpus information corresponding to the video segment.

[0099] In step 706, the corpus information corresponding to each video segment is expanded based on the copywriting generation model to obtain the copywriting corresponding to each video segment.

[0100] In step 708, the text corresponding to each of the above video segments is combined to obtain the video text of the target video.

[0101] Specifically, the implementation methods of steps 702-704 can refer to steps 402-404 above, and the implementation method of step 706 can also refer to step 502 above, and will not be repeated here.

[0102] In a specific example of this disclosure, assuming the determined structural information of the target video indicates that the target video includes the following video segments: introductory segment + product introduction segment + product description segment + benefit description segment + guidance information segment + end segment, then in the above method, for each video segment, the copywriting library can be called using the promotional information or based on tags determined by the promotional information to obtain the corpus that best matches the promotional information for each video segment. Then, for each video segment, the promotional information and the corpus matched from the corpus are input into the copywriting generation model, which expands and outputs the video copywriting corresponding to each video segment, and combines them to obtain the complete video copywriting. In some cases, the duration of certain video segments may be very short, therefore, the amount of text required for the copywriting of such video segments is very small, for example, the introductory segment usually has a very short duration. In the above method, for such short video segments, the copywriting generation model can be skipped from expanding the corpus extracted from the copywriting library, and the copywriting for this video segment can be generated directly using the obtained corpus.

[0103] The above Figure 7 The method shown combines corpus-based video script generation with model-based automatic video script generation. This not only enables the rapid generation of video scripts that meet user needs but also improves the accuracy of the generated video scripts.

[0104] In other embodiments of this disclosure, in addition to promotional information, users can also upload one or more original materials through the aforementioned interface module 106. These original materials can specifically be multimedia files in one or more forms, such as videos or images. It should be noted that the embodiments of this disclosure do not limit the quantity, length, content, or format of the aforementioned original materials. In these embodiments, in the process of generating video scripts, in addition to referring to the aforementioned promotional information, the content of one or more original materials uploaded by the user can also be further referenced. In these embodiments, in the aforementioned... Figure 5 The video text generation method shown can be further enhanced by: acquiring at least one original material corresponding to the target video; and determining the features corresponding to the at least one original material. The features corresponding to the original material are typically subject information and / or content information related to the original material obtained after content recognition, such as information about the product or person corresponding to the original material, and the content the original material intends to express. In practical applications, content recognition models or other methods can be used for multimedia files such as images and videos; the specific implementation methods disclosed in this embodiment are not limited.

[0105] Based on this, step 502 above, which describes generating the text corresponding to each video segment based on the text generation model and the promotional information according to the type and duration information of the video segments, may include: generating the text corresponding to each video segment based on the text generation model, the promotional information, and the features corresponding to at least one of the original materials, according to the type and duration information of the video segments.

[0106] As can be seen, the above method, in the process of generating video scripts, not only refers to the promotional information input by the user, but also to the content of one or more original materials uploaded by the user. This allows the generated scripts to not only match the promotional information, but also fit with the original materials uploaded by the user, resulting in a better fit between the final video scripts and the visuals.

[0107] As can be seen from the above, the video script generation method provided in this disclosure can determine the video structure information of a target video based on user-input promotional information using a pre-established video structure library, and then directly and quickly generate the video script of the target video using a pre-established corpus based on the determined video structure information and promotional information. Furthermore, since the video script of the target video is generated using a pre-established video structure library and corpus, it can meet the video delivery requirements.

[0108] As mentioned earlier, the video structure library can be created manually, automatically based on videos already published on short video content platforms, or by combining these two methods. The following section will provide a detailed explanation of the method for automatically creating the video structure library. Figure 8 The implementation flow of the method for automatically building a video structure library according to embodiments of this disclosure is shown. For example... Figure 8 The method described above may include the following steps:

[0109] In step 802, multiple videos to be processed within a predetermined time interval are acquired.

[0110] Perform the following operations for each video to be processed:

[0111] In step 804, relevant information about the video to be processed is determined.

[0112] In step 806, the tags corresponding to the video to be processed are determined based on the relevant information of the video to be processed.

[0113] In step 808, the video to be processed is deconstructed to obtain the video structure information corresponding to the video to be processed.

[0114] In step 810, the information pair consisting of the video structure information corresponding to the video to be processed and the tags corresponding to the video to be processed is added to the video structure library.

[0115] Specifically, in the embodiments of this disclosure, the predetermined time interval can be a periodically set time interval, such as daily, weekly, or even monthly. That is to say, the above... Figure 8 The method described can be executed periodically, meaning it can periodically retrieve a subset of highly popular videos from short video content platforms for analysis to update the video structure library, thereby achieving continuous iteration, upgrading, and optimization of the video structure library. Furthermore, the aforementioned... Figure 8 The method shown can be used not only to build video structure libraries, but also to further update and improve video structure libraries that have already been built manually.

[0116] Typically, to achieve better execution results and obtain more attractive video scripts and videos, the videos to be processed are usually selected from those with high popularity values ​​published on short video content platforms. The embodiments of this disclosure do not limit how the popularity value of a video is determined. It can generally be based on a comprehensive evaluation of multiple aspects, such as video quality, video appeal to users, and the input and / or output of the video. For example, one or more indicators such as completion rate, conversion rate, return on investment, and consumption can be selected to evaluate the quality of the video, thereby obtaining the popularity value of each video on the short video content platform. Videos with higher popularity values ​​(e.g., videos with popularity values ​​higher than a preset threshold or those ranked high in popularity) are then selected as the videos to be processed.

[0117] In step 804 above, the relevant information of the video to be processed includes: characteristic information of the product associated with the video, target user information, time node information, and delivery target information, or any combination thereof. The characteristic information of the product associated with the video, target user information, and delivery target information can be extracted from the additional information of the video. Specifically, videos on short video content platforms typically contain corresponding additional information, which may carry characteristic information of the product associated with the video, target user information, and delivery target information. Therefore, this information can be extracted from the additional information of the video. Alternatively, the above information can also be determined through methods such as content recognition of the video.

[0118] As can be seen, the dimensions of the information contained in the above-mentioned video to be processed are basically consistent with the dimensions of the video tags. Therefore, in step 806 above, the tags corresponding to the video to be processed can be determined directly based on the relevant information of the video to be processed.

[0119] In step 808 above, information from multiple dimensions, such as the video frame, scene, background music, and text, can be analyzed. Based on the analysis results, the video is decomposed into multiple video segments, and the type and duration of each video segment are determined based on the analysis results, thereby obtaining the video structure information corresponding to the video to be processed. In practical applications, the specific analysis of information from multiple dimensions, such as the video frame, scene, background music, and text, can also be performed using content recognition models or other methods; the specific implementation method is not limited in the embodiments disclosed herein.

[0120] As can be seen from the above method of establishing a video structure library, by deconstructing popular videos on short video platforms, we can obtain video deconstruction information of popular videos. If we further use this to establish a video structure library, the target videos determined based on the video structure library can largely inherit the "narrative" structure of popular videos, thereby improving the quality and attractiveness of the generated video scripts.

[0121] Based on the above-described video text generation method, embodiments of this disclosure also provide a video generation method. Figure 9 The implementation flow of the video generation method described in this embodiment is illustrated. This video generation method can be implemented collaboratively by the aforementioned video structure generation engine 108, the aforementioned video text generation engine 110, and the aforementioned video editing engine 112. The aforementioned video structure generation engine 108, the aforementioned video text generation engine 110, and the aforementioned video editing engine 112 typically include hardware devices with data processing capabilities and / or the necessary software to drive the operation of such hardware devices. Figure 9 As shown, the above video generation method mainly includes the following steps:

[0122] In step 902, promotional information corresponding to the target video and at least one original material are obtained.

[0123] In step 904, the video structure information of the target video is generated based on the above promotional information.

[0124] In step 906, a video script for the target video is generated based on the aforementioned video structure information and promotional information.

[0125] In step 908, based on the above video structure information, the above video script, and the above promotional information, at least one of the above original materials is edited to obtain an edited video file.

[0126] It should be noted that the above steps 902-906 can be referred to the aforementioned... Figures 2-8The method shown is implemented as described, and will not be explained again here.

[0127] Regarding step 908 above, the specific implementation method for editing at least one of the original materials based on the video structure information, the video script, and the promotional information can be found in [reference needed]. Figure 10 Specifically, it can include:

[0128] In step 1002, the at least one original material is screened and edited according to the video structure information, the quantity, length and quality of the at least one original material, to obtain at least one spare material.

[0129] Specifically, in step 1002 above, when the amount of original material is large and the length is long enough to meet the total duration requirement of the target video, the original material can be screened, filtered, or even edited based on the quality of the original material, such as the image quality, the stability of the picture, and the clarity, etc., which are used to measure video quality. The original material with poor quality can be filtered out or the poor quality parts can be removed by editing to obtain the backup material that can be used to generate the target video.

[0130] In step 1004, the display subject of at least one of the above-mentioned backup materials is determined.

[0131] In the embodiments of this disclosure, the display subject of at least one of the above-mentioned backup materials can be determined by content recognition, such as what product or person a certain video shows, etc., so that it can be matched with the target product corresponding to the target video in subsequent steps.

[0132] In step 1006, based on the display subject of the at least one backup material, the promotional information, and the type and duration information of at least one video segment contained in the target video, the at least one backup material is edited to obtain the video corresponding to each video segment.

[0133] In the embodiments of this disclosure, in step 1006 above, the target product corresponding to the target video can be determined based on the aforementioned promotional information, and the target product corresponding to the target video can be matched with the display subject of the aforementioned backup materials. Video materials with a high degree of matching are selected, or the highly matched portions of the backup materials are retained through editing. Further, the selected backup materials are edited according to the type and duration information of at least one video segment contained in the target video to obtain the video corresponding to each video segment. For example, for a promotional video related to men's waterproof sports shoes, the display subject of the backup materials is identified and analyzed, retaining materials or portions of materials where the display subject is men's sports shoes, and editing to obtain the video corresponding to each video segment.

[0134] In step 1008, the videos corresponding to each video segment are combined to obtain the edited video file.

[0135] In the embodiments of this disclosure, after obtaining the videos corresponding to each video segment, the videos corresponding to each video segment can be spliced ​​together in the order of each video segment to obtain the edited video file.

[0136] Furthermore, in some other embodiments of this disclosure, based on obtaining the videos corresponding to each video segment, the videos corresponding to each video segment can be further edited according to the generated video text to achieve a better match between the video footage and the text. Specifically, the above method may further include: performing the following operations for each video segment: performing content recognition on the video corresponding to the above video segment to obtain a first recognition result; performing content recognition on the video text corresponding to the above video segment to obtain a second recognition result; and further editing the video corresponding to the above video segment based on the first and second recognition results. Since the above method requires secondary editing based on the degree of fit between the video content and the text content, the length of the video corresponding to each video segment output in step 1006 should be longer than the requirement of the duration information corresponding to the video segment. Taking a promotional video related to men's waterproof sports shoes as an example, in the above method, content recognition can be performed on the video corresponding to each video segment, and content recognition can also be performed on the text corresponding to each video segment. The text with a high degree of matching in the content recognition results can be matched with the video and then edited, so that the content of the video and the content of the text are basically consistent. For example, suppose we identify a video script describing the waterproof performance of shoes, and simultaneously identify a video showing water being splashed onto the shoes. We can then match the identified script and video together, and edit them by incorporating factors such as script and video lengths to ensure their durations match. It's important to note that the aforementioned content recognition of both video and script can be achieved using one or more content recognition models. These models will possess the ability to understand the content of video and / or text. Therefore, the editing method described above can use content understanding capabilities to match video elements in video footage with text elements in video scripts, achieving precise matching and automatic editing of video and text elements, thus achieving "audio-visual synchronization" in the editing logic.

[0137] In some embodiments of this disclosure, the above Figure 9 The video generation method shown may further include the following steps, such as... Figure 9 As shown in the dashed box in the image.

[0138] In step 910, a voiceover for the target video is generated based on the aforementioned video script and promotional information.

[0139] In step 912, the aforementioned voiceover is added to the edited video file to generate the target video.

[0140] For step 910 above, the specific method for generating voiceover for the target video based on the video script and promotional information can be found in [reference needed]. Figure 11 Specifically, it can include:

[0141] In step 1102, the tags corresponding to the target video are determined based on the above promotional information.

[0142] In step 1104, based on the tags corresponding to the target video, a target virtual timbre that matches the tags corresponding to the target video is selected from the pre-set virtual timbres.

[0143] In step 1106, a voiceover for the target video is generated based on the generated video text and the voice pack of the aforementioned target virtual voice.

[0144] Specifically, in the embodiments of this disclosure, the pre-set virtual voice will also have one or more tags to describe the characteristics of the virtual voice, such as the characteristics of the matching product and the characteristics of the target user. In this way, by matching the tags corresponding to the target video with the tags of the virtual voice, suitable virtual voices for broadcasting the target video can be found as much as possible, thereby improving the user experience.

[0145] Regarding step 912 above, after generating the video script, the edited video file, and the voice-over, the voice-over can be added to the edited video file to generate the final version of the target video.

[0146] Furthermore, subtitles corresponding to the video script can be added to the target video. When adding subtitles, it's necessary to consider not only the phrasing of natural language but also factors such as the playback speed of the video footage and audio, as well as the display area of ​​the subtitles, to achieve consistency between the subtitles and the visuals and audio. In practical applications, video compositing engines can be used to synthesize video files, audio, and subtitles. That is, after inputting the edited video file, the voice pack or virtual voice identifier for the audio, and the video script into the video compositing engine, the engine can directly output the synthesized target video.

[0147] As can be seen from the above, the video generation method provided in this disclosure can automatically match editing strategies based on the user's promotion needs (product characteristics, target users, promotion goals, etc.) to generate target videos that meet the video delivery requirements, thereby greatly improving the user experience.

[0148] Corresponding to the above method, embodiments of this disclosure also disclose a video script generation device and a video generation device.

[0149] Figure 12 The internal structure of the video text generation apparatus described in this embodiment is shown. Figure 12 As shown, the device may include:

[0150] Information acquisition module 1202 is used to acquire promotional information corresponding to the target video;

[0151] The video structure information determination module 1204 is used to generate video structure information of the target video based on the promotion information; and

[0152] The video script generation module 1206 is used to generate video scripts for the target video based on the video structure information and the promotion information.

[0153] Figure 13 The internal structure of the video generation apparatus described in this embodiment is shown. For example... Figure 13 As shown, the device may include:

[0154] Information acquisition module 1302 is used to acquire promotional information corresponding to the target video;

[0155] The video structure information determination module 1304 is used to generate video structure information of the target video based on the promotion information.

[0156] The video script generation module 1306 is used to generate a video script for the target video based on the video structure information and the promotion information.

[0157] The video editing module 1308 is used to edit the at least one original material based on the video structure information, the video script, and the promotional information to obtain an edited video file.

[0158] Furthermore, the aforementioned video generation device may further include the following modules, such as... Figure 13 As shown in the dashed box in the image.

[0159] Dubbing module 1310 is used to generate dubbing for the target video based on the video script and the promotional information; and

[0160] The compositing module 1312 is used to add the dubbing to the edited video file to generate the target video.

[0161] The specific implementation of each of the above modules can be found in the aforementioned methods and accompanying drawings, and will not be repeated here. For ease of description, the above apparatus is described in terms of function, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0162] The apparatus described above is used to implement the corresponding video text generation method or video generation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0163] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the video text generation method or video generation method described in any of the above embodiments.

[0164] Figure 14 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 2010, a memory 2020, an input / output interface 2030, a communication interface 2040, and a bus 2050. The processor 2010, memory 2020, input / output interface 2030, and communication interface 2040 are interconnected internally via the bus 2050.

[0165] The processor 2010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0166] The memory 2020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 2020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 2020 and is called and executed by the processor 2010.

[0167] The input / output interface 2030 is used to connect input / output modules to enable information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0168] The communication interface 2040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (e.g., USB, Ethernet cable) or wireless means (e.g., mobile network, Wi-Fi, Bluetooth).

[0169] Bus 2050 includes a pathway for transmitting information between various components of the device, such as processor 2010, memory 2020, input / output interface 2030, and communication interface 2040.

[0170] It should be noted that although the above-described device only shows the processor 2010, memory 2020, input / output interface 2030, communication interface 2040, and bus 2050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0171] The electronic devices described above are used to implement the corresponding video generation methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0172] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the video text generation method or video generation method as described in any of the above embodiments.

[0173] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0174] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the task processing method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0175] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.

[0176] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0177] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0178] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating video scripts, comprising: Obtain promotional information corresponding to the target video; The video structure information of the target video is generated based on the promotional information; wherein, the video structure information includes the type and duration information of at least one video segment contained in the target video; the type of the at least one video segment is determined according to the function and content of the video segment, including: pre-roll segment, video backbone segment, and post-roll segment; The tags corresponding to the target video are determined based on the promotional information; For each of the at least one video segment, corpus data matching the tags of the target video is selected from a pre-established corpus based on the type and duration of the video segment as the corpus information corresponding to that video segment; wherein, the corpus stores multiple corpus data corresponding to each video segment and the tags corresponding to the multiple corpus data; and The corpus information corresponding to each video segment is combined to obtain the video text of the target video.

2. The video script generation method according to claim 1, wherein, The video structure information for generating the target video based on the promotional information includes: Based on the promotional information, determine the tags corresponding to the target video; and Video structure information matching the tags of the target video is selected from a pre-established video structure library as the video structure information of the target video; wherein, The video structure library stores video structure information and its corresponding tags for various combinations of video segments.

3. The video script generation method according to claim 1, wherein, The video structure information for generating the target video based on the promotional information includes: The video structure information of the target video is generated based on the video structure generation model and the promotional information; wherein... The video structure generation model is trained and generated based on a pre-established video structure library; and The video structure library stores video structure information and its corresponding tags for various combinations of video segments.

4. The video text generation method according to claim 2 or 3, wherein, The method further includes: Acquire multiple videos to be processed within a predetermined time interval; For each video to be processed, perform the following steps: Determine the relevant information of the video to be processed; wherein, the relevant information of the video to be processed includes: one or any combination of the characteristic information of the product associated with the video to be processed, the target user information, the time node information, and the delivery target information; The tags corresponding to the video to be processed are determined based on the relevant information of the video to be processed; The video to be processed is deconstructed to obtain the video structure information corresponding to the video to be processed; and The information pair consisting of the video structure information corresponding to the video to be processed and the tag corresponding to the video to be processed is added to the video structure library.

5. The video script generation method according to claim 1, wherein, The corpus information corresponding to each video segment is combined to obtain the video text of the target video, including: The text generation model expands the corpus information corresponding to each video segment to obtain the text corresponding to each video segment; wherein, the text generation model is trained and generated based on the corpus; and The text corresponding to each video segment is combined to obtain the video text of the target video.

6. The video text generation method according to claim 5, wherein, The method further includes: acquiring at least one original source material corresponding to the target video; and determining features corresponding to the at least one original source material; wherein... Expanding the corpus information corresponding to each video segment based on the copywriting generation model to obtain the copywriting corresponding to each video segment includes: expanding the corpus information corresponding to each video segment based on the type and duration information of the video segment, the copywriting generation model, the promotional information, and the features corresponding to at least one original material, to obtain the copywriting corresponding to each video segment.

7. A video generation method, comprising: Obtain promotional information corresponding to the target video and at least one original source material; The video structure information of the target video is generated based on the promotional information; wherein, the video structure information includes the type and duration information of at least one video segment contained in the target video; the type of the at least one video segment is determined according to the function and content of the video segment, including: pre-roll segment, video backbone segment, and post-roll segment; The tags corresponding to the target video are determined based on the promotional information; For each video segment in the at least one video segment, corpus matching the tags of the target video is selected from a pre-established corpus according to the type and duration information of the video segment as the corpus information corresponding to the video segment; wherein, the corpus stores multiple corpus corresponding to each video segment and the tags corresponding to the multiple corpus; The corpus information corresponding to each video segment is combined to obtain the video text of the target video; and Based on the video structure information, the video script, and the promotional information, the at least one original material is edited to obtain an edited video file.

8. The video generation method according to claim 7, wherein, Editing the at least one original material based on the video structure information, the video script, and the promotional information includes: Based on the video structure information, the quantity, length, and quality of the at least one original material, the at least one backup material is edited to obtain at least one backup material. Each of the at least one backup material is to be displayed in a specific subject; Based on the display subject of the at least one backup material, the promotional information, and the video structure information, the at least one backup material is edited to obtain videos corresponding to each video segment; and The videos corresponding to each video segment are combined to obtain the edited video file.

9. The video generation method according to claim 8, wherein, Editing the at least one original material based on the video structure information, the video script, and the promotional information further includes: Perform the following operations on the videos corresponding to each video segment: Perform video content recognition on the videos corresponding to the video segments to obtain a first recognition result; Text content recognition is performed on the video text corresponding to the video segments to obtain a second recognition result; and The video segments are edited based on the first and second recognition results.

10. The video generation method according to claim 7, wherein, The method further includes: The target video's voiceover is generated based on the video script and the promotional information; and The voiceover is added to the edited video file to generate the target video.

11. The video generation method according to claim 10, wherein, Generating the voiceover for the target video based on the video script and the promotional information includes: The tags corresponding to the target video are determined based on the promotional information; Based on the tags corresponding to the target video, select a target virtual timbre from pre-set virtual timbres that matches the tags corresponding to the target video; and The voiceover for the target video is generated based on the video text and the voice pack corresponding to the target virtual voice.

12. A video script generation device, comprising: The information acquisition module is used to acquire promotional information corresponding to the target video; A video structure information determination module is used to generate video structure information of the target video based on the promotional information; wherein, the video structure information includes the type and duration information of at least one video segment contained in the target video; the type of the at least one video segment is determined according to the function and content of the video segment, including: a pre-roll segment, a video backbone segment, and a post-roll segment; and The video script generation module is used to determine the tags corresponding to the target video based on the promotional information; for each video segment in the at least one video segment, it selects corpus information corresponding to the video segment from a pre-established corpus based on the type and duration information of the video segment; wherein, the corpus stores multiple corpus information corresponding to each video segment and the tags corresponding to the multiple corpus information; and combines the corpus information corresponding to each video segment to obtain the video script of the target video.

13. A video generation apparatus, comprising: The information acquisition module is used to acquire promotional information corresponding to the target video; A video structure information determination module is used to generate video structure information of the target video based on the promotion information; wherein, the video structure information includes the type and duration information of at least one video segment contained in the target video; the type of the at least one video segment is determined according to the function and content of the video segment, including: pre-roll segment, video backbone segment, and post-roll segment; A video script generation module is used to determine the tags corresponding to the target video based on the promotional information; for each video segment in the at least one video segment, it selects corpus information corresponding to the video segment from a pre-established corpus based on the type and duration of the video segment; wherein, the corpus stores multiple corpus information corresponding to each video segment and the tags corresponding to the multiple corpus information; and combines the corpus information corresponding to each video segment to obtain the video script of the target video. The video editing module is used to edit at least one original material based on the video structure information, the video script, and the promotional information to obtain an edited video file.

14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1-3 or 5-11.

15. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method according to any one of claims 1-3 or 5-11.

16. A computer program product comprising computer program instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-3 or 5-11.

Citation Information

Patent Citations

  • Processing method and device, electronic equipment and medium

    CN114697760A

  • Video generation method and device based on video structure information, equipment and medium

    CN114845161A

  • Video generation method and device and propaganda type video generation method and device

    CN115174824A