Method and apparatus for generating video template, and electronic device
By obtaining and completing material elements and combining pre-trained models to generate video templates, the problems of high labor costs and limited number of templates in the existing technology are solved, and efficient and low-cost video template generation and quality improvement are achieved.
Patent Information
- Application Number
- PCT/CN2024/074071
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-03
- Filing Date
- 2024-01-25
- Publication Date
- 2025-09-04
Smart Images

Figure CN2024074071_04092025_PF_FP_ABST
Abstract
Description
Method, device and electronic device for generating video template
[0001] This application claims priority to the Chinese invention patent application entitled “Method, device and electronic device for generating video templates” and application number 202310118673.6, filed on February 3, 2023. The entire contents of that application are incorporated by reference into this application. Technical Field
[0002] The present disclosure relates to the field of computer technology, and in particular to a method, device, and electronic device for generating a video template. Background Art
[0003] Internet users often need to edit their videos when creating them, a process that can be challenging and time-consuming. Video templates, with defined structural information, can render videos that are consistent in structure and look and feel with sample videos. This lowers the barrier to entry for internet users and improves both efficiency and quality.
[0004] Currently, video templates are mainly designed manually by designers. The number of such video templates is limited and they consume a lot of manpower costs.
[0005] Summary of the Invention
[0006] This disclosure section is provided to briefly introduce concepts that will be described in detail in the detailed description section below. This disclosure section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] In a first aspect, an embodiment of the present disclosure provides a method for generating a video template, the method comprising: obtaining a first category of material elements; completing a second category of material elements based on the first category of material elements; determining a target editing element based on the first category of material elements and the second category of material elements; and generating a target video template based on the first category of material elements, the second category of material elements and the target editing element.
[0008] In a second aspect, an embodiment of the present disclosure provides a device for generating a video template, comprising: an acquisition unit for acquiring a first category of material elements; a completion unit for completing a second category of material elements based on the first category of material elements; a determination unit for determining a target editing element based on the first category of material elements and the second category of material elements; and a generation unit for generating a target video template based on the first category of material elements, the second category of material elements and the target editing element.
[0009] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating a video template as described in the first aspect.
[0010] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for generating a video template as described in the first aspect.
[0011] The method, device and electronic device for generating a video template provided by the embodiments of the present disclosure obtain a first type of material element, wherein the first type of material element is an existing material element of the target video template (i.e., the video template to be generated); then, based on the first type of material element, the second type of material element is completed; then, based on the first type of material element and the second type of material element, the target editing element of the target video template is determined, wherein the target editing element can be used to edit the target material element; finally, based on the first type of material element, the second type of material element and the target editing element, the target video template is generated. In this way, any existing material element of the target video template can be used to generate the target video template, thereby improving the generation efficiency of the video template and reducing the generation cost of the video template. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0013] FIG1 is a flow chart of an embodiment of a method for generating a video template according to the present disclosure;
[0014] FIG2 is a flowchart of another embodiment of a method for generating a video template according to the present disclosure;
[0015] FIG3 is a flow chart of an embodiment of establishing a sound-image matching model in a method for generating a video template according to the present disclosure;
[0016] FIG4 is a schematic diagram of an application scenario of establishing an audio-visual matching model in the method for generating a video template according to the present disclosure;
[0017] FIG5 is a flowchart of yet another embodiment of a method for generating a video template according to the present disclosure;
[0018] FIG6 is a schematic structural diagram of an embodiment of an apparatus for generating a video template according to the present disclosure;
[0019] FIG7 is an exemplary system architecture in which the method for generating a video template according to an embodiment of the present disclosure may be applied;
[0020] FIG8 is a schematic diagram of a basic structure of an electronic device provided according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0022] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0023] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0024] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0025] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0026] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0027] Please refer to Figure 1, which shows a process of an embodiment of a method for generating a video template according to the present disclosure. As shown in Figure 1, the method for generating a video template includes the following steps:
[0028] Step 101: Obtain a first type of material element.
[0029] In this embodiment, the execution entity (eg, a server and / or a terminal device) may obtain the first type of material elements.
[0030] In some embodiments, the first type of material elements may include at least one of the following but is not limited to: picture material, audio material, and reference video template.
[0031] Here, the reference video template may be a video template used to provide a reference.
[0032] Here, the first type of material elements are usually existing material elements.
[0033] Step 102: Completing the second category of material elements based on the first category of material elements.
[0034] The second type of material elements are generally material elements that are missing from the target video template. The target material elements may include target picture materials and target audio materials.
[0035] Specifically, if the first type of material element is a picture material, the execution entity may obtain audio material as the second type of material element of the target video template. As an example, audio material may be selected from hot audio materials as the second type of material element.
[0036] If the first type of material element is audio material, the execution subject may obtain picture material as the second type of material element. As an example, picture material may be selected from hot picture materials as the second type of material element.
[0037] If the above-mentioned first-category material elements are audio materials and picture materials, the above-mentioned execution entity can obtain a reference video template as the above-mentioned second-category material elements.
[0038] Step 103: Determine a target editing element based on the first type of material elements and the second type of material elements.
[0039] Here, the target editing element can be used to edit the target material element.
[0040] In some embodiments, the target editing element may include at least one of the following but is not limited to: transition, animation, special effect, sticker, font and filter.
[0041] In some optional implementations of this embodiment, a transition generally refers to the transition or conversion between scenes in a video. Special effects generally refer to visual effects created by computer software and added to a recorded video. Stickers generally refer to decorative patterns added to a video. By adding editing elements such as transitions, animations, stickers, and filters to a target video template, the richness of the video template content can be enhanced.
[0042] In some embodiments, the execution entity may input the first and / or second types of editing elements into a pre-trained editing element recommendation model to obtain target editing elements for the target video template. The editing element recommendation model may be used to characterize the correspondence between the template's material elements and the template's editing elements.
[0043] Step 104: Generate a target video template based on the first type of material elements, the second type of material elements, and the target editing elements.
[0044] Here, the generated target video template may include the above-mentioned first-category material elements, second-category material elements and target editing elements.
[0045] In this embodiment, the execution entity may generate the target video template based on the first type of material elements determined in step 101 , the second type of material elements determined in step 102 , and the target editing elements determined in step 103 .
[0046] For example, if the target editing element includes a portrait filter, the execution entity may add the portrait filter to the picture element. If the target editing element includes a blessing sticker, the blessing sticker may be added to the picture element.
[0047] Here, after generating the target video template, the execution entity may package the target video template into a template resource package, and may recommend a title and appropriate copy for the target video template.
[0048] The method provided by the above embodiment of the present disclosure is through obtaining the first category of material elements, wherein the above first category of material elements are the existing material elements of the above target video template (i.e., the video template to be generated); then, based on the above first category of material elements, the second category of material elements are completed, wherein the above second category of material elements; then, based on the above first category of material elements and / or the second category of material elements, the target editing elements of the above target video template are determined, wherein the above target editing elements can be used to edit the above target material elements; finally, based on the above first category of material elements, the second category of material elements and the above target editing elements, the above target video template is generated. In this way, any existing material elements of the target video template can be used to generate the target video template, thereby improving the generation efficiency of the video template and reducing the generation cost of the video template.
[0049] In some embodiments, the method further includes: in response to a video generation instruction including a target video template identifier, acquiring original material, reading the target video template; and generating a video in combination with the target video template.
[0050] The target video template identifier may indicate the target video template. The original material may be material provided by a user for generating a video, such as several pictures or a piece of audio.
[0051] In some implementations, the first-category material elements, the second-category material elements, and the target editing elements in the target video template are stored in an agreed format. In response to the video generation instruction, the first-category material elements, the second-category material elements, and the target editing elements in the target video template are read according to the agreed format, and the read first-category material elements, the second-category material elements, and the target editing elements are combined with the acquired original material to generate a video.
[0052] Here, a method for generating a video using a target video template is provided. Specifically, the user provides the original material and reads the video template. The original material and the video template can be combined to generate the user's desired video. In this way, the user can quickly generate the desired video using the video template.
[0053] In some embodiments, when completing the second category of material elements based on the first category of material elements includes completing picture materials based on audio materials, the above step 102 may include: inputting the number of audio materials and picture materials into the audio-visual matching model, and outputting a picture group in the picture material library that matches the number of picture materials.
[0054] In some implementations, the first type of material elements and candidate materials in the picture material library may be input into a pre-established audio-visual matching model, and the audio-visual matching model may output a picture group that matches the audio material based on the degree of matching.
[0055] In some implementations, since there are multiple candidate materials, the audio-visual matching model outputs multiple matching degrees, representing the matching degrees between multiple candidate materials and the first-category material elements. Based on the matching degrees, the second-category material elements can be selected from the candidate materials. For example, the candidate material with the highest matching degree can be selected from the candidate materials as the second-category material element; alternatively, any candidate material from the candidate materials whose matching degree exceeds a preset matching degree threshold can be selected as the second-category material element.
[0056] By using the audio-visual matching model to output the matching degree between audio elements and picture elements, more suitable material elements can be selected based on the matching degree, thereby improving the generation quality of the video template.
[0057] In some optional implementations of this embodiment, the execution entity may complete the second-category material elements based on the first-category material elements in the following manner: the execution entity may select materials that match the first-category material elements from a preset material library as the second-category material elements. Here, the materials in the preset material library typically correspond to at least one pre-set tag, for example, the tags may be funny, unboxing, winter, travel, and so on. Here, for each material in the preset material library, the execution entity may determine the degree of match between at least one tag corresponding to the material and at least one tag corresponding to the first-category material elements; and then, the material with the highest degree of match may be selected from the preset material library as the second-category material element.
[0058] Afterwards, the execution entity may combine the first type of material elements with the second type of material elements. Since one of the first type of material elements and the second type of material elements is a picture material and the other is an audio material, the execution entity may divide the audio material according to the number of picture materials (also referred to as the number of slots), so that the number of audio segments after division is equal to the number of picture materials. Then, the multiple audio segments obtained after division may be matched with the sequence of picture materials. Card point matching is usually used to associate the playback time of the segment nodes with the preset positions in the sequence of picture materials, so that the corresponding picture material is presented while a certain audio segment is played. By selecting material elements from a preset material library to generate a video template, the generation quality of the video template can be guaranteed.
[0059] In some optional implementations of this embodiment, the above-mentioned execution subject may select materials that match the above-mentioned first-category material elements as the above-mentioned second-category material elements from the preset material library in the following manner: the above-mentioned execution subject may determine the candidate materials based on the categories of the materials in the preset material library and the categories of the above-mentioned first-category material elements. For example, the above-mentioned categories may be food, people, scenery, and so on. Here, the above-mentioned execution subject may select materials of the same category as the above-mentioned first-category material elements from the above-mentioned preset material library as candidate materials. As an example, if the above-mentioned first-category material elements are pictures of food, audio of food may be obtained. It should be noted that if the type of the candidate material is a picture, the determined candidate materials are usually multiple groups of picture materials, and the number of pictures in each group of picture materials may be set according to actual conditions.
[0060] In some optional implementations of this embodiment, the execution entity may also establish the material library. The material library establishment step may include at least one of the following:
[0061] The execution entity can identify the image information in the image and store the image information in association with the image. The image information can include at least one of the following: character relationships, image style, and image scene. Character relationships can also be called interpersonal relationships, which usually refer to the interconnected social relationships formed by interactions among social groups, such as couples, mothers and children, siblings, and friends. Image style can include but is not limited to at least one of the following: fresh, literary, private, fashionable, and black and white. Image scenes are usually the scenes presented in the image, such as dance, sports, parties, travel, and reunions.
[0062] As an example, the execution subject may input a picture into a pre-trained picture recognition model to obtain picture information of the picture. The picture recognition model may be used to characterize the correspondence between the picture and the picture information.
[0063] The above-mentioned execution entity can identify audio information in the audio and store the above-mentioned audio information in association with the above-mentioned audio. The above-mentioned audio information may include a musical style. The musical style can also be called audio style, audio melody, audio type, or audio genre, which usually refers to the representative and unique appearance of the audio work as a whole, such as folk, rock, Chinese style, etc. As an example, the above-mentioned execution entity can input the audio into a pre-trained audio recognition model to obtain the audio information of the audio. The above-mentioned audio recognition model can be used to characterize the correspondence between audio and audio information.
[0064] The above-mentioned execution entity can parse the editing element information in the reference template element, and the above-mentioned editing element information includes the spatiotemporal layout information of the editing element and the editing element identifier. The spatiotemporal layout information of the editing element is usually used to represent the timestamp corresponding to when the editing element appears in the reference template and the position presented in the reference template. The editing element identifier can be used to indicate the editing element. For example, the editing element identifier "1" can represent the editing element of transition, and the editing element identifier "2" can represent the editing element of animation.
[0065] In some embodiments, the target video template is generated based on the first category material elements, the second category material elements and the target editing elements, including: detecting the climax segment of the audio material; dividing the climax segment into segments according to lyrics information and / or beat information to obtain segment nodes; and performing card point matching on the climax segment divided into segments with the sequence of the target picture material, wherein the card point matching is used to associate the playback time of the segment node with a preset position in the sequence of the target picture material.
[0066] In some optional implementations of this embodiment, the above-mentioned execution subject can combine the above-mentioned first category material elements with the above-mentioned second category material elements in the following manner: the above-mentioned execution subject can detect the climax segment of the target audio material. The climax segment of the audio can be understood as the most emotional and infectious part of the audio, such as the chorus part, which is usually related to factors such as rhythm, beat, strength and speed. The climax segment can appear at any position in the audio. The above-mentioned execution subject can input the above-mentioned target audio material into a pre-trained climax segment detection model to obtain the climax segment of the above-mentioned target audio material. The above-mentioned climax segment detection model can be used to characterize the correspondence between audio and the climax segment of audio.
[0067] The climax segment can then be segmented based on the lyrics and / or beat information to obtain segment nodes. Specifically, the execution entity can segment the climax segment based on the lyrics, i.e., each line of lyrics corresponds to a segment. The execution entity can also segment the climax segment based on the beat. A beat is a unit of rhythm. In audio, a series of beats with varying strengths repeat at regular intervals. The execution entity can also segment the climax segment based on both the lyrics and the beat.
[0068] Then, the climax segment divided by the segments can be matched with the sequence of the above-mentioned target image material. Here, the above-mentioned card point matching is usually used to associate the playback time of the above-mentioned segment node with the preset position in the sequence of the above-mentioned target image material. As an example, if four segments are obtained after the climax segment is divided, and there are four pictures in the sequence of the above-mentioned target image material, the first audio segment can be associated with the first picture in the sequence in the order of playback time from first to last, and the first picture in the sequence can be presented while playing the first audio segment. Similarly, the second audio segment can be associated with the second picture in the sequence in terms of playback time, the third audio segment can be associated with the third picture in the sequence in terms of playback time, and the fourth audio segment can be associated with the fourth picture in the sequence in terms of playback time.
[0069] By using the climax segment of the target audio material to generate a video template, the generation quality of the video template is further improved; in addition, the climax segment is segmented according to the lyrics information and / or beat information, thereby ensuring the integrity of the lyrics and / or beat.
[0070] In some embodiments, the above step 103 may include: acquiring a reference sequence based on the editing elements and editing structure of the reference template; inputting the image material, audio material, and reference sequence into a multimodal model to obtain a target editing element sequence.
[0071] The reference template may include at least one editing element, which may be stored in multiple tracks according to a predetermined editing structure. The reference sequence may lack at least one editing element relative to the template editing element sequence. By inputting image material, audio material, and the reference sequence into a multimodal model, at least one editing element may be obtained. Thus, the at least one editing element obtained based on the multimodal model is combined with the reference sequence to obtain a target editing element sequence having a greater number of editing elements than the reference sequence.
[0072] Here, the clip structure mentioned above may include the clip structure of multiple tracks in the video template. For example, the clip structure of the music track, editing element track, picture track, etc. The clip structure of the track includes the temporal and spatial distribution information of different materials.
[0073] Here, the multimodal model can be used to input image materials, audio materials, and editing elements, and output one or more editing elements. Different editing elements (i.e., different editing elements expected to be output) can use the same multimodal model or different multimodal models.
[0074] As an example, if the target editing element to be determined is a transition, the image material and audio material of the target video template, as well as other editing elements in the reference template other than transitions (e.g., animations, filters, etc.) can be used to determine the transition of the target video template. If the target editing element to be determined is an animation, the image material and audio material of the target video template, as well as other editing elements in the reference template other than animations (e.g., transitions, filters, etc.) can be used to determine the animation of the target video template. Here, if the target editing element to be determined is a transition, the execution entity can input the first category of editing elements and / or the second category of editing elements and other editing elements in the reference template other than transitions into a pre-trained transition prediction model to obtain the transition of the target video template. The transition prediction model is generally used to predict the transition of a video template. If the target editing element to be determined is an animation, the execution entity can input the first category of editing elements and / or the second category of editing elements and other editing elements in the reference template other than animations into a pre-trained animation prediction model to obtain the animation of the target video template. The above animation prediction model is usually used to predict the animation of a video template. In this way, when determining the target editing element, other editing elements except the target editing element are fixed, thereby improving the accuracy of the determined target editing element.
[0075] In some embodiments, the above-mentioned step 104 may include: determining the target spatiotemporal layout information of the target editing element in the first category of material elements and / or the second category of material elements; according to the spatiotemporal layout information, associating the display timing and / or display position of the target editing element with the first category of material elements and / or the second category of material elements to obtain the target video template.
[0076] Here, the target spatiotemporal layout information includes target time layout information and / or target space layout information.
[0077] In some optional implementations of this embodiment, after determining the target editing element, the execution entity may determine target spatiotemporal layout information of the target editing element in the first-category material element and / or the second-category material element. The target spatiotemporal layout information may include target time layout information and / or target spatial layout information, i.e., the timestamp corresponding to the appearance of the target editing element in the first-category material element and / or the second-category material element and / or the position of the target editing element in the first-category material element and / or the second-category material element.
[0078] The above-mentioned execution entity can generate the above-mentioned target video template based on the above-mentioned first-category material elements and / or second-category material elements and the above-mentioned target editing elements in the following manner: the above-mentioned execution entity can generate the above-mentioned target video template based on the above-mentioned first-category material elements and / or second-category material elements, the above-mentioned target editing elements and the above-mentioned target spatiotemporal layout information.
[0079] Specifically, the above-mentioned execution entity can use the above-mentioned target spatiotemporal layout information to add the above-mentioned target editing element to the above-mentioned first-category material element and / or second-category material element. As an example, if the above-mentioned target spatiotemporal layout information indicates that the transition element appears at the moment of 1 minute and 20 seconds, the transition element can be associated with the picture presented at the moment of 1 minute and 20 seconds and the audio played. If the above-mentioned target spatiotemporal layout information indicates that the sticker element is presented in the upper left corner of the picture (which can be represented in the form of coordinates), the sticker element can be added to the upper left corner of the picture.
[0080] It should be noted that the presentation duration of transition elements and animation elements can be set according to actual conditions (e.g., user usage). In order to avoid blocking significant objects (e.g., faces) in the image, the presentation position of the subtitle information can be set. In addition, the candidate positions of the subtitle information can be scored in combination with an aesthetic scoring model to obtain the optimal presentation position of the subtitle information.
[0081] By determining the timestamp corresponding to when the target editing element appears in the first category material element and / or the second category material element and / or the position at which it is presented in the above-mentioned first category material element and / or the second category material element, the appearance time and presentation position of the editing element can be made more appropriate, thereby further improving the generation effect of the video template.
[0082] In some embodiments, the above step 104 may include: replacing the reference material elements in the reference template with other material elements in the first category and the second category except the reference template, and replacing the reference editing elements in the reference template with the target editing elements to obtain the target video template.
[0083] Optionally, the first-category material element may include a reference target, or the second-category material element may include a reference template. The reference target may include a reference material element. The image and audio materials in the reference template are replaced with the image and audio materials in the first-category material element and the second-category material element. The reference editing elements in the reference template are replaced with the target editing elements. A target video template may be obtained.
[0084] Therefore, a new target video template can be generated more quickly based on the original video template.
[0085] 2, which shows a process flow of another embodiment of a method for generating a video template according to the present disclosure. As shown in FIG2, the method for generating a video template includes the following steps:
[0086] Step 201: Obtain the first type of material elements.
[0087] In this embodiment, the first type of material elements may include target picture material and / or target audio material.
[0088] Step 202: Determine a reference template of a target video template based on the target picture material and / or target audio material.
[0089] In this embodiment, the execution subject of the method for generating a video template can determine a reference template for the target video template based on the target image material and / or the target audio material. Here, the image material, audio material, and reference template can have corresponding labels. The execution subject can select a reference template that matches the label of the target image material from a preset reference template library; can also select a reference template that matches the label of the target audio material from a preset reference template library; can also combine the labels of the target image material and the target audio material to select a matching reference template from a preset reference template library.
[0090] Step 203: Determine the target editing element of the target video template based on the target picture material, the target audio material and the reference template.
[0091] Step 204: Replace the reference material elements in the reference template with other material elements in the first and second categories except the reference template, and replace the reference editing elements in the reference template with the target editing elements to obtain the target video template.
[0092] As can be seen from Figure 2, compared to the embodiment corresponding to Figure 1, process 200 of the editing template generation method in this embodiment embodies the steps of determining a reference template for a target video template and replacing the material elements and editing elements in the reference template with the material elements and editing elements corresponding to the target video template. Thus, the solution described in this embodiment can learn from the creator's editing experience from existing reference templates, thereby improving the template quality of the target video template.
[0093] Please refer to Figure 3, which shows a process for establishing an audio-visual matching model in an embodiment of the method for generating a video template according to the present disclosure. As shown in Figure 3, the audio-visual matching model establishment method includes the following steps:
[0094] Step 301: Acquire an existing video template from an existing video template library, determine the audio in the existing video template as an audio sample, and determine the image in the existing video template as an image sample.
[0095] In this embodiment, the execution entity of the method for generating a video template can obtain an existing video template from an existing video template library. The existing video template library stores multiple existing video templates. Subsequently, the audio in the obtained existing video template can be determined as an audio sample, and the image in the existing video template can be determined as an image sample. Here, the audio samples and image samples belonging to the same existing video template generally match each other. The matching between the audio and the image can be understood as the audio and the image having the same or similar characteristics such as style and theme.
[0096] Step 302: Using the initial model, extract audio features of the audio sample and extract image features of the image sample.
[0097] In this embodiment, the execution entity may use the initial model to extract audio features of the audio sample determined in step 301, and extract image features of the image sample determined in step 301. The initial model may be any neural network capable of extracting audio features and image features, such as a convolutional neural network, a deep neural network, and the like.
[0098] Specifically, an audio sample can be input into the initial model to obtain audio features of the audio sample, and an image sample can be input into the initial model to obtain image features of the image sample. Here, a matching degree label is associated with the audio sample and the image sample, indicating the degree of match between the two. For example, the matching degree label can be represented as a numerical value, where a higher matching degree corresponds to a larger numerical value.
[0099] In this embodiment, the initial model may include a first initial sub-model and a second initial sub-model. An audio sample may be input into the first initial sub-model to obtain the audio features of the audio sample, and an image sample may be input into the second initial sub-model to obtain the image features of the image sample. Here, the second initial sub-model may be an encoder of an attention model. The attention model may be a model based on a multi-head attention mechanism. The attention mechanism was originally applied to image feature extraction tasks. For example, when a person observes an image, he or she does not observe every part of the image, but focuses on the important parts. The multi-head attention mechanism uses multiple attention mechanisms to perform separate calculations to obtain semantic information at more levels, and then splices and combines the results obtained by each attention mechanism to obtain the final result.
[0100] Step 303: determine the matching degree between the audio features and the image features, and compare the matching degree with the corresponding matching degree label, and use the feedback to modify the extraction parameters used to extract the audio features and the image features in the initial model to obtain an audio-visual matching model.
[0101] In this embodiment, the execution entity may determine the degree of match between the audio features and image features extracted using the initial model. Specifically, the audio features and image features extracted by the initial model may be input into a pre-trained matching prediction model to obtain the degree of match between the extracted audio features and image features. The matching prediction model may be used to predict the degree of match between the audio features and image features.
[0102] Afterwards, the above-mentioned matching degree can be compared with the corresponding matching degree label, and the feedback can be used to modify the extraction parameters used to extract audio features and image features in the initial model to obtain an audio-visual matching model. Specifically, after comparing the above-mentioned matching degree with the corresponding matching degree label, it can be determined based on the comparison result whether the above-mentioned initial model has achieved the preset target. The above-mentioned preset target can mean that the accuracy of the audio features and image features extracted by the above-mentioned initial model is greater than a preset accuracy threshold.
[0103] As an example, the difference between the above-mentioned matching degree and the corresponding matching degree label can be determined. If the difference is greater than a preset difference threshold, it can be considered that the above-mentioned initial model has not achieved the preset target. If the difference is less than or equal to the preset difference threshold, it can be considered that the above-mentioned initial model has achieved the preset target.
[0104] If the above goals are not achieved, the extraction parameters used to extract audio features and image features in the above initial model can be modified, and steps 301-303 are continued until the above initial model reaches the preset goals, thereby obtaining the above audio-visual matching model. As an example, the back propagation algorithm (BP algorithm) and the gradient descent method (e.g., the mini-batch gradient descent algorithm) can be used to adjust the extraction parameters of the above initial model. It should be noted that the back propagation algorithm and the gradient descent method are currently well-known technologies that are widely studied and applied, and will not be described in detail here.
[0105] The method provided by the above-mentioned embodiment of the present disclosure obtains audio samples and image samples from an existing video template library, extracts audio features from the audio samples, extracts image features from the image samples, determines the matching degree between the extracted audio features and image features, compares the matching degree with the corresponding matching degree label, and uses the comparison result to feedback and modify the extraction parameters used to extract audio features and image features in the initial model until the initial model meets the preset target, thereby learning the matching relationship between image and audio. In the application stage of the audio-visual matching model, by inputting audio and the required number of pictures, the audio can be matched with images in the candidate image group library (images with similar themes, styles, etc. can be grouped together), and the group of pictures with the highest matching degree is found for output.
[0106] Continuing with FIG4 , FIG4 is a schematic diagram of an application scenario for establishing an audio-visual matching model in the method for generating a video template according to this embodiment. In the application scenario of FIG4 , an audio sample 401 and an image sample 402 are matched with each other. An audio feature encoder 403 can be used to extract the audio features of the audio sample 401 (as shown in icon 405 ), and an image feature encoder 404 can be used to extract the image features of the image sample 402 (as shown in icon 406 ). Afterwards, the two modalities of audio and image can be aligned to learn the matching relationship between the audio and image.
[0107] Further referring to FIG5 , FIG5 is a process for generating a video template according to another embodiment of the present embodiment. The process may represent an optional process when the method for generating a video template is applied in a specific application scenario. The process is summarized as follows:
[0108] Preprocessing: Configurable parameters (style, number of slots, and type) are used to preprocess elements such as the candidate group image library, candidate audio library, and reference templates. The preprocessing steps include material understanding, material completion, and segmentation. Material understanding includes image label understanding, audio label understanding, and reference template parsing. Material completion includes image recommendation templates, image recommendation audio, and audio recommendation images. Segmentation includes audio card extraction, audio climax detection, and audio lyrics recognition.
[0109] The steps for recommending editing elements include: image matching stickers, image matching special effects, transition sequence generation, font recommendation, text recommendation, and animation recommendation.
[0110] The steps of space-time layout design include: sticker space layout, transition animation duration determination and motion effect timeline arrangement.
[0111] The post-processing steps include: generating resource packages, rendering sample videos, evaluating template quality, recommending titles, generating text, and outputting template details, thereby obtaining a video template and adding it to the video template library.
[0112] Further referring to FIG6 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for generating a video template. The device embodiment corresponds to the method embodiment shown in FIG1 , and the device can be specifically applied to various electronic devices.
[0113] As shown in Figure 6, the apparatus for generating a video template in this embodiment includes an acquisition unit 601, a completion unit 602, a determination unit 603, and a generation unit 604. The acquisition unit is configured to acquire a first-category material element; the completion unit is configured to complete a second-category material element based on the first-category material element; the determination unit is configured to determine a target editing element based on the first-category material element and the second-category material element; and the generation unit is configured to generate the target video template based on the first-category material element, the second-category material element, and the target editing element.
[0114] In this embodiment, the specific processing of the acquisition unit 601, the completion unit 602, the determination unit 603 and the generation unit 604 of the device for generating the video template and the technical effects brought about by them can be referred to the relevant descriptions of steps 101, 102, 103 and 104 in the corresponding embodiment of Figure 1 respectively, and will not be repeated here.
[0115] In some embodiments, the first category of material elements and the second category of material elements include at least one of the following: picture materials, audio materials, and reference video templates; the target editing elements include at least one of the following: transitions, animations, special effects, stickers, fonts, and filters.
[0116] In some embodiments, when the "completing the second category of material elements based on the first category of material elements" includes completing the picture material based on the audio material, the number of the audio material and the picture material is input into the audio-visual matching model, and a picture group in the picture material library that matches the number of the picture material is output.
[0117] In some embodiments, determining the target editing elements based on the first category of material elements and / or the second category of material elements includes: obtaining a reference sequence based on the editing elements and clipping structure of the reference template; inputting the image material, audio material, and reference sequence into a multimodal model to obtain a target editing element sequence.
[0118] In some embodiments, the generation of the target video template based on the first category of material elements, the second category of material elements and the target editing elements includes, in some embodiments, replacing the reference material elements in the reference template with other material elements in the first category of material elements and the second category of material elements except the reference template, and replacing the reference editing elements in the reference template with the target editing elements to obtain the target video template.
[0119] In some embodiments, the target video template is generated based on the first category of material elements, the second category of material elements and the target editing elements, including: determining the target spatiotemporal layout information of the target editing elements in the first category of material elements and / or the second category of material elements, wherein the target spatiotemporal layout information includes target time layout information and / or target space layout information; according to the spatiotemporal layout information, the display timing and / or display position of the target editing element is associated with the first category of material elements and / or the second category of material elements to obtain the target video template.
[0120] In some embodiments, the target video template is generated based on the first category material elements, the second category material elements and the target editing elements, including: detecting the climax segment of the audio material; dividing the climax segment into segments according to lyrics information and / or beat information to obtain segment nodes; and performing card point matching on the climax segment divided into segments with the sequence of the target picture material, wherein the card point matching is used to associate the playback time of the segment node with a preset position in the sequence of the target picture material.
[0121] In some embodiments, the device is also used for a sound and picture matching model establishment step, wherein the sound and picture matching model establishment step includes: obtaining an existing video template from an existing video template library, and determining the audio in the existing video template as an audio sample, and determining the image in the existing video template as an image sample, wherein the audio samples and image samples belonging to the same existing video template match each other; using the initial model, extracting audio features of the audio samples, and extracting image features of the image samples, wherein the audio samples and the image samples have a matching degree label indicating the degree of matching between the two; determining the matching degree between the audio features and the image features, and comparing the matching degree with the corresponding matching degree label, and feedback-modifying the extraction parameters used to extract audio features and image features in the initial model to obtain the sound and picture matching model.
[0122] In some embodiments, the apparatus is further configured to: in response to a video generation instruction including a target video template identifier, obtain original material, read the target video template; and generate a video in combination with the target video template.
[0123] Please refer to FIG. 7 , which shows an exemplary system architecture in which the method for generating a video template according to an embodiment of the present disclosure can be applied.
[0124] As shown in Figure 7, the system architecture may include terminal devices 701, 702, and 703, a network 704, and a server 705. The network 704 is used to provide a medium for communication links between the terminal devices 701, 702, and 703 and the server 705. The network 704 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0125] The terminal devices 701, 702, and 703 can interact with the server 705 via the network 704 to receive or send messages, etc. Various client applications can be installed on the terminal devices 701, 702, and 703, such as short video applications, video processing applications, and instant messaging software.
[0126] Terminal devices 701, 702, and 703 can be hardware or software. When terminal devices 701, 702, and 703 are hardware, they can be various electronic devices with display screens and support web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Group Audio Layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Group Audio Layer 4) players, laptop computers, and desktop computers, etc. When terminal devices 701, 702, and 703 are software, they can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules (for example, software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.
[0127] Server 705 can be a server that provides various services, such as a backend server that processes material elements obtained from terminal devices 701, 702, and 703. Server 705 can obtain the first category of material elements of the target video template, i.e., existing material elements; then, based on the first category of material elements, complete the second category of material elements of the target video template, i.e., missing material elements, to obtain the target material elements of the target video template; then, based on the target material elements of the target video template, determine the target editing elements of the target video template; finally, based on the target material elements and the target editing elements, generate the target video template.
[0128] It should be noted that the method for generating a video template provided by the embodiment of the present disclosure is usually executed by the server 705 . Accordingly, the device for generating a video template may be set in the server 705 .
[0129] It should be understood that the number of terminal devices, networks and servers in Figure 7 is only illustrative and any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0130] Referring now to Figure 8, a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure (e.g., the server in Figure 7) is shown. The electronic device shown in Figure 8 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0131] As shown in Figure 8, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0132] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although FIG8 illustrates an electronic device having various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0133] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0134] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0135] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0136] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0137] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device is enabled to: obtain a first category of material elements; complete a second category of material elements based on the first category of material elements; determine a target editing element based on the first category of material elements and the second category of material elements; and generate a target video template based on the first category of material elements, the second category of material elements and the target editing element.
[0138] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0140] The units described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the acquisition unit may also be described as a "unit for acquiring the first type of material element of the target video template."
[0141] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0142] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0143] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0144] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0145] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for generating a video template, comprising: Get the first type of material elements; Completing the second category of material elements based on the first category of material elements; Determining a target editing element based on the first category of material elements and the second category of material elements; A target video template is generated based on the first-category material elements, the second-category material elements, and the target editing elements.
2. According to the method of claim 1, the first category of material elements and the second category of material elements include at least one of the following: picture materials, audio materials, reference video templates; the target editing elements include at least one of the following: transitions, animations, special effects, stickers, fonts and filters.
3. The method according to claim 2, when the method of supplementing the second type of material elements according to the first type of material elements comprises supplementing the picture material according to the audio material, The number of the audio materials and the picture materials is input into the audio-visual matching model, and a picture group matching the number of the picture materials in the picture material library is output.
4. The method according to claim 2, wherein determining the target editing element based on the first type of material elements and the second type of material elements comprises: Acquire a reference sequence based on the editing elements and clipping structure of the reference template; The image material, audio material, and reference sequence are input into a multimodal model to obtain a target editing element sequence.
5. The method according to claim 2, wherein generating the target video template based on the first type of material elements, the second type of material elements, and the target editing element comprises: The reference material elements in the reference template are replaced with other material elements in the first and second categories except the reference template, and the reference editing elements in the reference template are replaced with the target editing elements to obtain the target video template.
6. The method according to claim 1, wherein generating the target video template based on the first type of material elements, the second type of material elements, and the target editing element comprises: Determining target spatiotemporal layout information of the target editing element in the first type of material elements and / or the second type of material elements, wherein the target spatiotemporal layout information includes target time layout information and / or target space layout information; According to the spatiotemporal layout information, the display timing and / or display position of the target editing element is associated with the first category material element and / or the second category material element to obtain the target video template.
7. The method according to claim 1, wherein generating the target video template based on the first type of material elements, the second type of material elements, and the target editing element comprises: Detecting the climax of the audio material; Dividing the climax segment into segments according to the lyrics information and / or the beat information to obtain segment nodes; The climax segment divided into segments is matched with a sequence of target picture materials at card points, wherein the card point matching is used to associate the segment node with a preset position in the sequence of the target picture materials for playback time.
8. The method according to claim 3, further comprising the step of establishing a sound-image matching model, wherein: The step of establishing the audio-visual matching model includes: Obtaining an existing video template from an existing video template library, and determining the audio in the existing video template as an audio sample, and determining the image in the existing video template as an image sample, wherein the audio sample and the image sample belonging to the same existing video template match each other; Extracting audio features of the audio sample and image features of the image sample using the initial model, wherein the audio sample and the image sample have a matching degree label indicating a matching degree between the two; Determine the degree of matching between the audio features and the image features, compare the degree of matching with the corresponding matching degree label, and use feedback to modify the extraction parameters used to extract the audio features and the image features in the initial model to obtain the audio-visual matching model.
9. The method according to claim 1, further comprising: In response to a video generation instruction including a target video template identifier, original material is acquired, the target video template is read, and the target video template is combined to generate a video.
10. A device for generating a video template, comprising: An acquisition unit, used for acquiring the first type of material elements; a completion unit, configured to complete the second type of material elements based on the first type of material elements; a determining unit, configured to determine a target editing element based on the first-category material elements and the second-category material elements; A generating unit is configured to generate a target video template based on the first-category material elements, the second-category material elements, and the target editing elements.
11. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.
12. A computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.