Video generation method, apparatus, device, storage medium and program product

Automating video clipping through a template-based approach reduces time and enhances video quality by eliminating manual operations.

JP7732004B2Active Publication Date: 2025-09-01BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023578709
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-10
Filing Date
2023-05-09
Publication Date
2025-09-01
Estimated Expiration
2043-05-09

AI Technical Summary

Technical Problem

Users lack clipping experience, leading to increased time and reduced quality in creating videos, especially when manually clipping video materials.

Method used

A method and device that apply a clipping operation from a pre-defined template to multimedia data generated from text, eliminating the need for manual video clipping.

Benefits of technology

Reduces time cost and improves video quality by automating the clipping process, ensuring higher efficiency and better video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007732004000001
    Figure 0007732004000001
  • Figure 0007732004000002
    Figure 0007732004000002
  • Figure 0007732004000003
    Figure 0007732004000003
Patent Text Reader

Abstract

The present disclosure relates to a video generating method, an apparatus, a device, a storage medium and a program product, the method including: generating initial multimedia data based on received text data; acquiring a target clipping template in response to a clipping template acquisition request; applying a clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data; and generating a target video based on the target multimedia data. The embodiments of the present disclosure generate a video by directly applying a clipping operation in the acquired clipping template to the multimedia data without a user manually clipping the video, which not only reduces the time cost of creating the video but also improves the quality of the created video.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Related Applications] This application claims priority to a Chinese patent application filed on May 10, 2022, entitled "Video generation method, apparatus, device, storage medium and program product," with application number 202210508063.2.

[0002] [Technical field] The present disclosure relates to the technical field of video processing, and in particular to video generation methods, apparatus, devices, storage media and program products. [Background technology]

[0003] With the rapid development of computer technology and mobile communication technology, various electronic device-based video platforms have become commonplace, greatly enriching people's daily lives, and more and more users are willing to share their video works on video platforms for other users to view.

[0004] In the related art, when creating a video, a user must first find all kinds of materials needed for the video by himself, then perform a series of complex video clipping operations on the materials, and finally generate a video production.

[0005] If a user lacks clipping experience, it will increase the time cost of creating a video and the quality of the created video will be low. Summary of the Invention

[0006] To solve the above technical problems, the embodiments of the present disclosure provide a video generation method, apparatus, device, storage medium and program product, which generate a video by directly applying the clipping operation in the acquired clipping template to multimedia data, eliminating the need for users to manually clip the video, which not only reduces the time cost of creating the video but also improves the quality of the created video.

[0007] According to a first aspect, an embodiment of the present disclosure provides a video generation method, the method comprising:

[0008] generating initial multimedia data based on the received text data, wherein the initial multimedia data includes a video image in which a spoken voice of the text data matches the text data, the initial multimedia data includes at least one multimedia fragment, each of the at least one multimedia fragment corresponding to at least one text fragment divided by the text data, a target multimedia fragment in the at least one multimedia fragment corresponding to a target text fragment in the at least one text fragment, the target multimedia fragment including a target video fragment and a target audio fragment, the target video fragment including a video image that matches the target text fragment, and the target audio fragment including a spoken voice that matches the target text fragment;

[0009] Obtaining a target clipping template in response to a clipping template obtainment request;

[0010] applying a clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data;

[0011] generating a target video based on the target multimedia data.

[0012] According to a second aspect, an embodiment of the present disclosure provides a video generation device, the device comprising: an initial multimedia data generation module for generating initial multimedia data based on the received text data, wherein the initial multimedia data includes a video image in which a reading voice of the text data matches the text data, the initial multimedia data includes at least one multimedia fragment, each of the at least one multimedia fragment corresponding to at least one text fragment divided by the text data, a target multimedia fragment in the at least one multimedia fragment corresponding to a target text fragment in the at least one text fragment, the target multimedia fragment including a target video fragment and a target audio fragment, the target video fragment including a video image matching the target text fragment, and the target audio fragment including a reading voice matching the target text fragment; a target clipping template acquisition module for acquiring a target clipping template in response to the clipping template acquisition request; a target multimedia data generation module for applying the clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data; a target video generation module for generating a target video based on the target multimedia data.

[0013] According to a third aspect, an embodiment of the present disclosure provides an electronic device, the electronic device comprising: one or more processors; a storage device for storing one or more programs; The one or more programs, when executed by the one or more processors, cause the one or more processors to perform the video generation method of any one of the first aspects above.

[0014] According to a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having stored thereon a computer program, the program causing a processor to perform the video generation method according to any one of the first aspects above when executed by the processor.

[0015] According to a fifth aspect, an embodiment of the present disclosure provides a computer program product including a computer program or instructions that, when executed by a processor, causes the computer program or instructions to perform the video generation method according to any one of the first aspects above.

[0016]

[0009] Embodiments of the present disclosure provide a video generation method, apparatus, device, storage medium, and program product, the method including: generating initial multimedia data based on received text data; acquiring a target clipping template in response to a clipping template acquisition request; applying clipping operations indicated by the target clipping template to the initial multimedia data to obtain target multimedia data; and generating a target video based on the target multimedia data. In the embodiments of the present disclosure, the clipping operations in the acquired clipping template are directly applied to the multimedia data to generate a video, eliminating the need for users to manually clip the video, thereby reducing the time cost of video creation and improving the quality of the created video. [Brief explanation of the drawings]

[0017] These and other features, advantages, and aspects of each example of the present disclosure will become more apparent with reference to the following specific embodiments in conjunction with the accompanying drawings. Identical or similar reference numerals refer to identical or similar elements throughout the accompanying drawings. It should be understood that the accompanying drawings are schematic, and that actual objects and elements are not necessarily drawn to scale. [Figure 1] FIG. 1 is an architecture diagram of a video production scenario provided by an embodiment of the present disclosure. [Figure 2]1 is a schematic flowchart of a video generation method in an embodiment of the present disclosure. [Figure 3] FIG. 10 is a schematic diagram of a trigger for a template theme control in an embodiment of the present disclosure. [Figure 4] FIG. 10 is a schematic diagram of a trigger for a template control in an embodiment of the present disclosure. [Figure 5] FIG. 10 is a schematic diagram of a template application prompt in an embodiment of the present disclosure. [Figure 6] 1 is a schematic structural diagram of a video generating device in an embodiment of the present disclosure; [Figure 7] 1 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure; DETAILED DESCRIPTION OF THE INVENTION

[0018] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although several embodiments of the present disclosure are illustrated in the accompanying drawings, it should be understood that the present disclosure can be realized in various forms and is not limited to the embodiments described herein, but rather these embodiments are provided for a more thorough and complete understanding of the present disclosure. It should be understood that the accompanying drawings and embodiments of the present disclosure are used for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0019] It should be noted that the steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the performance of illustrated steps. The scope of the present disclosure is not particularly limited in this respect.

[0020] As used herein, the term "comprises" and variations thereof are open-ended, meaning "including, but not limited to." The term "based on" means "based at least in part on." The term "in one embodiment" means "at least one embodiment," the term "in another embodiment" means "at least one other embodiment," and the term "in some embodiments" means "at least some embodiments." Relevant definitions of other terms are provided in the description below.

[0021] It should be noted that the concepts of "first", "second", etc. referred to in this disclosure are used to distinguish between different devices, modules or units, and are not used to define the order or interdependence of functions performed by these devices, modules or units.

[0022] It should be noted that the modifications "one" and "multiple" referred to in this disclosure are intended to be exemplary rather than limiting, and those skilled in the art will understand that "one or more" should be understood unless the context clearly indicates otherwise.

[0023] The names of messages or information interacted between multiple devices in the embodiments of the present disclosure are used for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0024] Before describing the embodiments of the present application in detail, application scenarios of the embodiments of the present application will be described first.

[0025] When users deal with documents, they are often presented in text format, which can be time-consuming for users to read. Converting text information into video eliminates the need for users to struggle to decipher the text. By listening to the audio and watching the video, users can clarify the information conveyed in the article, making it easier for users to acquire information. Alternatively, because a text is long and users find it time-consuming to read it, they may not have the energy to read it one by one. By converting the article into video, users can quickly understand the information conveyed in the article through the video, and then select the parts of the article that interest them and read them carefully. Furthermore, the diverse presentation formats of video make it easier to attract users' attention than reading boring text, encouraging them to read articles in this way.

[0026] In the related art, keywords are extracted from text data, and for each keyword, video images matching the keyword are searched for in a predetermined image library. Then, the text information and the video images are synthesized according to typographic rules to obtain a target video. However, in the related art, only a simple synthesis of the searched video images and the text data is performed, and the resulting video is not of high quality. After that, the user needs to manually clip the video, and if the user lacks clipping experience, the video quality will be affected.

[0027] In an embodiment of the present application, after generating initial multimedia data based on text data, a target clipping template is obtained, and the clipping operation indicated by the target clipping template is applied to the initial multimedia data to realize clipping processing of the initial multimedia data, so that the user does not have to manually clip the video, which not only reduces the time cost of video production but also improves the quality of the produced video. Figure 1 shows an architecture diagram of a video production scenario provided by an embodiment of the present disclosure.

[0028] As shown in FIG. 1 , the architecture diagram may include at least one client electronic device 101 and at least one server 102. The electronic device 101 establishes a connection and interacts with the server 102 via a network protocol, such as HyperText Transfer Protocol over Secure Socket Layer (HTTPS). Here, the electronic device 101 may include a device with communication capabilities, such as a mobile phone, tablet computer, desktop computer, laptop computer, in-vehicle terminal, wearable device, all-in-one computer, or smart home device, or a device simulated by a virtual machine or simulator. The server 102 may include a device with memory and computing capabilities, such as a cloud server or server cluster.

[0029] Based on the above architecture, a user can create a video within a designated platform on an electronic device 101, which may be a designated application program or a designated website. After creating a video, the user can send the video to a server 102 of the designated platform, which can receive the video sent from the electronic device 101, store the received video, and send the video to an electronic device that needs to play the video.

[0030] In an embodiment of the present disclosure, in order to reduce the time cost of creating a video and improve the quality of the created video, the electronic device 101 receives a user's request to obtain a clipping template for initial multimedia data. After receiving the clipping template request, the electronic device 101 can obtain a target clipping template, apply the clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data, and generate a target video based on the target multimedia data. In this way, by directly applying the clipping operation in the target clipping template obtained during the target video generation process to the initial multimedia data, the user does not need to manually clip the video, thereby reducing the time cost of creating the video and improving the quality of the created video.

[0031] Optionally, based on the above architecture, the electronic device 101 receives a clipping template acquisition request to acquire a target clipping template, applies the clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data, and generates a target video based on the target multimedia data, so that the electronic device 101 locally applies the clipping operation indicated by the target clipping template to the initial multimedia data to generate the target video, thereby further reducing the time cost of creating the video.

[0032] Optionally, based on the above architecture, after receiving the clipping template acquisition request, the electronic device 101 can also send a clipping template acquisition request including a template identifier to the server 102. After receiving the clipping template acquisition request including the template identifier sent from the electronic device 101, the server 102 can, in response to the clipping template acquisition request, acquire a target clipping template, apply clipping operations indicated by the target clipping template to the initial multimedia data to obtain target multimedia data, generate a target video based on the target multimedia data, and send the generated target video to the electronic device 101. The electronic device 101 can request the server 102 to acquire the target clipping template based on the clipping template acquisition request and apply clipping operations indicated by the target clipping template to the initial multimedia data to generate the target video, thereby further improving the quality of the created video and reducing the amount of data processing by the electronic device 101.

[0033] For example, electronic devices include mobile, fixed or portable terminals such as mobile phones, stations, units, devices, multimedia computers, multimedia tablets, Internet nodes, communicators, desktop personal computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, electronic book devices, gaming devices, or any combination thereof (including accessories and peripherals of these devices, or any combination thereof).

[0034] The server may be a physical server or a cloud server, and the server may be a single server or a server cluster.

[0035] The video generating method proposed by the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings.

[0036] FIG. 2 is a flowchart of a video generation method in an embodiment of the present disclosure, which is applicable to generating a video based on text information, and which can be performed by a video generation device, which can be realized in a software and / or hardware manner, and which can be implemented in the electronic device shown in FIG. 1.

[0037] As shown in FIG. 2, the video generating method provided by the embodiment of the present disclosure mainly includes steps S101 to S104.

[0038] S101, generating initial multimedia data based on received text data. In one embodiment of the present disclosure, the text data may be data input by a user to an electronic device using an input device, or may be data transmitted from another device to the electronic device.

[0039] In one embodiment of the present disclosure, before generating the initial multimedia data based on the received text data, the method further includes receiving text data in response to a user's data input operation, where the user's data input operation may include a text data addition operation or a text data input operation, and is not particularly limited in this embodiment.

[0040] In one embodiment of the present disclosure, the initial multimedia data includes a video image in which the reading audio of the text data matches the text data, the initial multimedia data includes at least one multimedia fragment, each of the at least one multimedia fragment corresponding to at least one text fragment divided by the text data, a target multimedia fragment in the at least one multimedia fragment corresponding to a target text fragment in the at least one text fragment, the target multimedia fragment including a target video fragment and a target audio fragment, the target video fragment including a video image matching the target text fragment, and the target audio fragment including the reading audio matching the target text fragment.

[0041] In one embodiment of the present disclosure, generating initial multimedia data based on received text data includes: dividing the received text data into at least one text fragment, where the text fragment includes multiple target text fragments; for each target text fragment, searching for a video image corresponding to the target text fragment from a preset image library based on the target text fragment, and processing the video image according to preset video effects to obtain a target video fragment corresponding to the target text fragment; obtaining speech audio matching the target text fragment and generating a target audio fragment; and synthesizing the target video fragment and the target audio fragment to obtain a target multimedia fragment; and for each target text fragment, obtaining multiple target multimedia fragments and synthesizing the multiple target multimedia fragments in a forward / backward order of the target text fragment to obtain initial multimedia data.

[0042] In one embodiment of the present disclosure, the video image includes subtitle text that matches the target text fragment.

[0043] In an embodiment of the present disclosure, subtitle text that matches a target text fragment is added to a video image to facilitate a user to intuitively view the subtitles corresponding to the spoken audio during the process of watching a video, and to improve the user's viewing experience.

[0044] S102, in response to the clipping template acquisition request, acquire a target clipping template.

[0045] In an embodiment of the present disclosure, the response to the clipping template acquisition request may be a response to the clipping template acquisition request after accepting a user operation of the electronic device, or a response to the clipping template acquisition request after detecting generation of initial multimedia data.

[0046] The target clipping template may be a clipping template selected based on a user's operation of the electronic device, or may be a clipping template that is automatically matched based on keywords in the text data.

[0047] In one embodiment of the present disclosure, obtaining the target clipping template includes the electronic device obtaining the target clipping template from a template database pre-stored locally.

[0048] In one embodiment of the present disclosure, obtaining the target clipping template includes the electronic device obtaining a template identifier corresponding to the target clipping template, sending a clipping template obtainment request including the template identifier to a server, the server responding to the clipping template obtainment request including the template identifier, obtaining the target clipping template based on the template identifier, and returning the obtained target clipping template to the electronic device.

[0049] In one embodiment of the present disclosure, if a target clipping template is not obtained, a prompt pop-up box is displayed on the display interface of the electronic device, and the prompt pop-up box is used to indicate that the target clipping template has not been obtained.

[0050] In one embodiment of the present disclosure, obtaining a target clipping template in response to a clipping template acquisition request includes determining, in response to a trigger operation of a template theme control, a clipping template corresponding to the trigger operation as the target clipping template, and obtaining the target clipping template.

[0051] In one embodiment of the present disclosure, at least one template theme control is displayed on an interaction interface of an electronic device, and in response to a user's triggering operation of the template theme control, a clipping template corresponding to the triggering operation is determined as a target clipping template.

[0052] As shown in FIG. 3, in response to a trigger operation of the template theme 1 control by the user, the clipping template corresponding to the template theme 1 control is determined as the target clipping template.

[0053] In the embodiment of the present disclosure, the target clipping template is selected by the user's trigger operation, which makes it easier for the user to select a clipping template that satisfies the user, thereby improving the user's usage experience.

[0054] In one embodiment of the present disclosure, the method further includes displaying a video editing area before responding to a trigger operation of a clipping template control, where the video editing area includes a template control, and displaying a mask area in response to a trigger operation of the template control, and displaying at least one template theme control in the mask area.

[0055] In an embodiment of the present disclosure, as shown in FIG. 4 , after generating the initial multimedia data, a video preview area 10 and a video editing area 20 are displayed on the display interface of the electronic device, and the video editing area 20 includes multiple editing controls, such as a template control, a screen control, a text control, a reading tone control, and a music control. Here, the template control is used to instruct a user to edit the initial multimedia data using an existing template. The screen control is used to instruct a user to edit the video image in the initial multimedia data. The text control is used to instruct a user to edit the subtitle text in the initial multimedia data. The reading tone control is used to instruct a user to edit the reading voice in the initial multimedia data. The music control is used to instruct a user to edit the background music in the initial multimedia data.

[0056] In one embodiment of the present disclosure, in response to a user's trigger operation of a template control, one mask area is displayed and multiple clipping template theme controls are displayed in the mask area, as shown in Fig. 4. In response to a left or right swipe operation on the mask area, multiple clipping template theme controls are displayed with a left or right swipe effect.

[0057] In the embodiment of the present disclosure, multiple template theme controls are displayed after responding to a user's trigger operation of the template control, which is simple and easy to understand and provides high user operation convenience.

[0058] S103, applying the clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data.

[0059] In one embodiment of the present disclosure, the target clipping template includes at least one clipping operation, which can be applied to the initial multimedia data to perform the clipping operation on the initial multimedia data.

[0060] In one embodiment of the present disclosure, as shown in FIG. 5 , in the process of applying the clipping operation indicated by the target clipping template to the initial multimedia data, since clipping the initial multimedia data takes a certain amount of time, an application prompt box is displayed on the display interface of the electronic device, and the application prompt box is used to instruct the user that the clipping process is being performed on the initial multimedia video with the clipping operation indicated by the clipping template.

[0061] In one embodiment of the present disclosure, if the clipping operation indicated by the target clipping template is successfully applied to the initial multimedia data, a prompt message indicating successful application of the clipping template is displayed; if the clipping operation indicated by the target clipping template is unsuccessfully applied to the initial multimedia data, a prompt message indicating unsuccessful application of the clipping template is displayed, prompting the user to reselect a clipping template.

[0062] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a video compositing operation, and applying the clipping operation indicated by the target clipping template to the initial multimedia data to obtain the target multimedia data includes compositing video fragments included in the target clipping template and multimedia fragments included in the initial multimedia data based on the video compositing operation to obtain the target multimedia data.

[0063] In an embodiment of the present disclosure, the target clipping template includes one or more video fragments, and the clipping operation indicated by the target clipping template, in the case of a video combining operation, includes combining one or more video fragments included in the target clipping template with multimedia fragments included in the initial multimedia data to obtain the target multimedia data.

[0064] In the embodiment of the present disclosure, a video fragment included in the target clipping template is added between any two video frames of the multimedia fragment. The above video fragment synthesis operation may be any existing video synthesis method, and is not particularly limited in this embodiment.

[0065] In the embodiments of the present disclosure, the video composition operation in the clipping template realizes the composition of multiple videos, avoiding the user from manually compositing the videos, reducing the time cost of creating the videos, and improving the quality of the created videos.

[0066] In one embodiment of the present disclosure, synthesizing a video fragment included in a target clipping template and a multimedia fragment included in initial multimedia data based on a video synthesis operation to obtain target multimedia data includes loading a video fragment included in the target clipping template into a set position of a multimedia fragment included in the initial multimedia data based on the video synthesis operation to obtain target multimedia data, wherein the set position includes before the first frame media data of the initial multimedia data and / or after the last frame media data of the initial multimedia data.

[0067] In an embodiment of the present disclosure, the target clipping template includes multiple video fragments and additional positions corresponding to each video fragment.

[0068] In one embodiment of the present disclosure, if the adding position corresponding to a video fragment included in the target clipping template is a prologue position, the video fragment is added before the first frame media data of the initial multimedia data as a target video prologue.

[0069] In one embodiment of the present disclosure, if the adding position corresponding to a video fragment included in the target clipping template is an epilogue position, the video fragment is added after the last frame media data of the initial multimedia data as a prologue of the target video.

[0070] In an embodiment of the present disclosure, if the text data includes a text theme, the text theme is added to the position of the text theme in the video fragment corresponding to the prologue, and the text theme is edited and rendered on the screen according to the text theme display effect included in the target clipping template.Furthermore, if the text data includes a text author, the text author is added to the position of the text author in the video fragment corresponding to the prologue, and the text author information is edited and rendered on the screen according to the text author display effect included in the target clipping template.

[0071] In one embodiment of the present disclosure, when video creator information is obtained, the video creator information is added to the creator's position in the video fragment corresponding to the epilogue, and the video creator information is edited and rendered on the screen according to the video creator display effect included in the target clipping template.

[0072] In the embodiments of the present disclosure, the video compositing operation in the clipping template realizes the operation of adding a prologue and / or an epilogue, thereby avoiding the user from manually adding a prologue or an epilogue, reducing the time cost of creating a video, and improving the quality of the created video.

[0073] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a transition setting operation, and applying the clipping operation indicated by the target clipping template to the initial multimedia data to obtain the target multimedia data includes adding a transition effect to the multimedia fragments included in the initial multimedia data based on the transition setting operation to obtain the target multimedia data.

[0074] In one embodiment of the present disclosure, the initial multimedia data includes multiple video images corresponding to text data, and the process of switching between the multiple video images necessarily involves image transition settings. In the related art, a user needs to manually set the transition effect between two adjacent video images, which increases the time cost of video production.

[0075] In one embodiment of the present disclosure, the transition effect includes one or more of a cut-in animation effect, a blink animation effect, a gradient animation effect, a cross-dissolve animation effect, a zoom animation effect, and the like.

[0076] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a transition setting operation, and the transition setting operation includes multiple transition effect types, and based on application of the multiple transition effect types included in the transition setting operation to the multimedia fragments, each multimedia fragment has a corresponding transition effect.

[0077] In one embodiment of the present disclosure, if a transition setting operation includes a transition effect type, applying the transition effect type to a multimedia fragment causes the multimedia fragment to have the same transition effect.

[0078] In an embodiment of the present disclosure, a transition effect is added to a multimedia fragment through a transition setting operation in a clipping template, which avoids users from manually setting the transition effect, reduces the time cost of creating a video, and improves the quality of the created video.

[0079] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a virtual object addition operation, and applying the clipping operation indicated by the target clipping template to the initial multimedia data to obtain the target multimedia data includes adding a virtual object included in the target clipping template to a preset position of the initial multimedia data by the virtual object addition operation to obtain the target multimedia data.

[0080] In one embodiment of the present disclosure, the virtual subject includes various subjects such as a target video fragment, a virtual sticker, a virtual object, a virtual card, etc. Selectable features may include facial decoration features, hair decoration features, clothing features, clothing accessory features, etc.

[0081] In one embodiment of the present disclosure, the virtual object stored in the target clipping template may be directly added to the preset position of the initial multimedia data. Optionally, specific parameters of the preset position may be stored in the target clipping template. A flash effect sticker stored in the target clipping template may be added to the third-width video image.

[0082] In one embodiment of the present disclosure, additional locations of virtual objects may be determined based on keywords presented in the text information, and the virtual objects may be selectively added to the video images corresponding to the keywords.

[0083] In the embodiments of the present disclosure, a virtual object is added to a multimedia fragment through a virtual object addition operation in a clipping template, thereby avoiding the user from manually adding a virtual object, reducing the time cost of creating a video, and improving the quality of the created video.

[0084] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a background audio addition operation, and applying the clipping operation indicated by the target clipping template to the initial multimedia data to obtain the target multimedia data includes mixing the background audio included in the target clipping template and the speech audio included in the initial multimedia data based on the background audio addition operation to obtain the target multimedia data.

[0085] In one embodiment of the present disclosure, the target clipping template includes background audio, and an add background audio operation mixes the background audio and the speech audio to obtain target multimedia data based on a timestamp corresponding to the background audio and a timestamp corresponding to the speech audio.

[0086] In one embodiment of the present disclosure, playback parameters of the background audio are adjusted based on playback parameters of the speech audio to better blend the two.

[0087] In the embodiments of the present disclosure, background music is added to multimedia fragments through the operation of adding background audio in the clipping template, which avoids users from manually adding background music, reduces the time cost of video creation, and improves the quality of the created video.

[0088] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a keyword extraction operation, and applying the clipping operation indicated by the target clipping template to the initial multimedia data includes, for at least one target text fragment, extracting keywords in the target text fragment and adding the keywords to the target multimedia fragment corresponding to the target text fragment.

[0089] In one embodiment of the present disclosure, keywords include dates, numbers, names of people, proper names, place names, plants, animals, and the like.

[0090] In one embodiment of the present disclosure, if the target text fragment is “Zhang San paid 200,000 yuan in cash to Li Si on the same day,” and the keyword extracted from the target text fragment is “200,000 yuan,” the keyword “200,000 yuan” is added to the target multimedia fragment corresponding to the target text fragment.

[0091] In an embodiment of the present disclosure, the target clipping module further includes a keyword parameter, where the keyword parameter includes the color, font, additional effect, etc. of the keyword. Set the display information of the keyword in the target multimedia fragment according to the keyword parameter.

[0092] In the embodiment of the present disclosure, through the keyword extraction operation in the clipping template, keywords can be added to the multimedia fragment, so that users can understand the key information of the text fragment more clearly.

[0093] In one embodiment of the present disclosure, adding a keyword to a target multimedia fragment corresponding to a target text fragment includes obtaining key text information that matches the keyword, and adding the keyword and the key text information to a target multimedia fragment corresponding to the target text fragment.

[0094] In the embodiment of the present disclosure, keywords are extracted from the target text fragment, and then key information corresponding to the keywords is obtained based on the keywords. For example, the keyword is "Wang Wu," and the key information corresponding to the keyword is "Wang Wu is an actor, and his representative works are TV Series A and Movie B." In this case, "Wang Wu" is the keyword, and "actor" and "representative works TV Series A and Movie B" are added as key text information to the target multimedia fragment. In addition, if the keyword is "embezzlement in the course of official duties" and the corresponding key text information is "embezzlement in the course of official duties refers to the act of a person belonging to a company, enterprise, or other unit taking advantage of his or her position to dishonestly obtain a relatively large amount of money from the unit's assets," then "embezzlement in the course of official duties" is the keyword, and "embezzlement in the course of official duties refers to the act of a person belonging to a company, enterprise, or other unit taking advantage of his or her position to dishonestly obtain a relatively large amount of money from the unit's assets" is added as key text information to the target multimedia fragment.

[0095] In one embodiment of the present disclosure, different display parameters may be set for keywords and key text information.

[0096] In one embodiment of the present disclosure, the key text information that matches the above keywords may be text information extracted from text data, or may be text information obtained from the Internet or a preset knowledge base. The method of obtaining the key text information is not particularly limited in this embodiment.

[0097] In the embodiment of the present disclosure, key text information is extracted by keywords, and the keywords and key text information are added to the video, helping users quickly understand the knowledge related to the keywords and understand the content of the text data.

[0098] S104, generating a target video based on the target multimedia data.

[0099]

[0009] Embodiments of the present disclosure provide a video generation method, apparatus, device, storage medium, and program product, the method including: generating initial multimedia data based on received text data, where the initial multimedia data includes a video image in which speech audio of the text data matches the text data; the initial multimedia data includes at least one multimedia fragment, each of the at least one multimedia fragment corresponding to at least one text fragment separated by the text data; a target multimedia fragment in the at least one multimedia fragment corresponds to a target text fragment in the at least one text fragment, the target multimedia fragment including a target video fragment and a target audio fragment, the target video fragment including a video image that matches the target text fragment, and the target audio fragment including speech audio that matches the target text fragment; acquiring a target clipping template in response to a clipping template acquisition request; applying clipping operations indicated by the target clipping template to the initial multimedia data to obtain target multimedia data; and generating a target video based on the target multimedia data. In embodiments of the present disclosure, generating a video by directly applying the clipping operations in the acquired clipping template to the multimedia data not only reduces the time cost of video creation without requiring a user to manually clip the video, but also improves the quality of the created video.

[0100] FIG. 6 is a flowchart of a video generation method in an embodiment of the present disclosure, which is applicable to generating a video based on text information, and the method can be performed by a video generation device, which can be implemented in software and / or hardware, and which can be provided in an electronic device.

[0101] As shown in FIG. 6 , the video generating device 60 provided by the embodiment of the present disclosure mainly includes: an initial multimedia data generating module 61, a target clipping template obtaining module 62, a target multimedia data generating module 63, and a target video generating module 64.

[0102] Here, the initial multimedia data generation module 61 is used to generate initial multimedia data based on the received text data, where the initial multimedia data includes a video image whose reading audio of the text data matches the text data, the initial multimedia data includes at least one multimedia fragment, each of the at least one multimedia fragment corresponding to at least one text fragment divided by the text data, a target multimedia fragment in the at least one multimedia fragment corresponds to a target text fragment in the at least one text fragment, the target multimedia fragment includes a target video fragment and a target audio fragment, the target video fragment includes a video image that matches the target text fragment, and the target audio fragment includes reading audio that matches the target text fragment, the target clipping template acquisition module 62 is used to acquire a target clipping template in response to the clipping template acquisition request, the target multimedia data generation module 63 is used to apply a clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data, and the target video generation module 64 is used to generate a target video based on the target multimedia data.

[0103] In one embodiment of the present disclosure, the video image includes subtitle text that matches the target text fragment.

[0104] In one embodiment of the present disclosure, the target clipping template acquisition module 62, which is used to acquire a target clipping template in response to a clipping template acquisition request, includes a target clipping template determination unit, which is used to determine a clipping template corresponding to a trigger operation as a target clipping template in response to a trigger operation of a template theme control, and a target clipping template acquisition unit, which is used to acquire the target clipping template.

[0105] In one embodiment of the present disclosure, the target clipping template acquisition module 62 further includes a video editing area display unit used to display a video editing area before responding to a triggering operation of a clipping template control, where the video editing area includes a template control, and a mask area display unit used to display a mask area and display at least one template theme control in the mask area in response to a triggering operation of the template control.

[0106] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a video compositing operation, and the target multimedia data generation module 63 is specifically used to synthesize the video fragments included in the target clipping template and the multimedia fragments included in the initial multimedia data based on the video compositing operation to obtain the target multimedia data.

[0107] In one embodiment of the present disclosure, the target multimedia data generation module 63 is specifically used to obtain target multimedia data by loading video fragments included in the target clipping template into set positions of multimedia fragments included in the initial multimedia data based on a video synthesis operation, where the set positions include before the first frame media data of the initial multimedia data and / or after the last frame media data of the initial multimedia data.

[0108] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a transition setting operation, and the target multimedia data generation module 63 is specifically used to add a transition effect to the multimedia fragment included in the initial multimedia data based on the transition setting operation to obtain the target multimedia data.

[0109] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a virtual object adding operation, and the target multimedia data generation module 63 is specifically used to add the virtual object included in the target clipping template to a preset position of the initial multimedia data based on the virtual object adding operation to obtain the target multimedia data.

[0110] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a background audio addition operation, and the target multimedia data generation module 63 is specifically used to obtain the target multimedia data by mixing the background audio included in the target clipping template and the reading audio included in the initial multimedia data based on the background audio addition operation.

[0111] In one embodiment of the present disclosure, the clipping operation indicated by the target clipping template includes a keyword extraction operation, and the target multimedia data generation module 63 is specifically used, for at least one target text fragment, to extract keywords in the target text fragment and add the keywords to the target multimedia fragment corresponding to the target text fragment.

[0112] In one embodiment of the present disclosure, the target multimedia data generation module 63 is specifically used to obtain key text information that matches the keyword, and add the keyword and the key text information to the target multimedia fragment corresponding to the target text fragment.

[0113] The video generation device provided by the embodiments of the present disclosure can perform the steps in the video generation method provided by the method embodiments of the present disclosure, and the performing steps and beneficial effects thereof will not be repeated here.

[0114] FIG. 7 is a schematic structural diagram of an electronic device in an embodiment of the present disclosure. Specifically referring to FIG. 7 below, a schematic structural diagram of an electronic device 700 suitable for implementing an embodiment of the present disclosure is shown. The electronic device 700 in the embodiment of the present disclosure includes, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and wearable terminals, as well as fixed terminals such as digital TVs, desktop computers, and smart home devices. The electronic device shown in FIG. 7 is merely an example and does not limit the functionality and scope of use of the embodiment of the present disclosure.

[0115] 7, electronic device 700 includes a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701 for performing various appropriate operations and processes according to a program stored in read-only memory (ROM) 702 or a program loaded from a storage device 708 into random access memory (RAM) 703 to realize the image rendering method of the embodiments of the present disclosure. RAM 703 further stores various programs and data necessary for the operation of terminal device 700. Processing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0116] Typically, input devices 706, such as a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, and gyroscope, output devices 707, such as a liquid crystal display (LCD), speaker, and vibrator, storage devices 708, such as a magnetic tape and hard disk, and communication devices 709 are connected to the I / O interface 705. The communication devices 709 enable the terminal device 700 to exchange data with other devices via wireless or wired communication. While FIG. 7 illustrates the terminal device 700 with various devices, it should be understood that it is not necessary for the terminal device 700 to implement or include all of the illustrated devices. Alternatively, more or fewer devices may be implemented or included.

[0117] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts may be implemented as a computer software program. For example, an embodiment of the present disclosure provides a computer program product including a computer program carried on a non-transitory computer-readable medium, the computer program including program code for executing the video generation method. In such an embodiment, the computer program may be downloaded and installed from a network via a communication device 709, or may be installed from the storage device 708, or may be installed from the ROM 702. When the computer program is executed by the processing device 701, the functions defined in the method of the embodiment of the present disclosure are realized.

[0118] It should be noted that the computer-readable medium described in this disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above. The computer-readable storage medium may be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier, carrying computer-readable program code. Such propagated data signals include, but are not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may be any computer-readable storage medium other than a computer-readable storage medium that transmits, propagates, or transfers a program used by or in conjunction with an instruction execution system, apparatus, or device. The program code contained in the computer-readable storage medium may be transferred by any suitable medium, such as, but not limited to, wire, fiber optic cable, RF (radio frequency), etc., or any suitable combination of the above.

[0119] In some embodiments, clients and servers may communicate using any now known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and may interconnect via any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), internetworks (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), and any now known or later developed networks.

[0120] The computer-readable storage medium may be included in the electronic device, or may be separate from the electronic device.

[0121] The computer-readable storage medium stores one or more programs, which, when executed by the terminal device, cause the terminal device to perform the following operations: generate initial multimedia data based on received text data, where the initial multimedia data includes a video image whose spoken audio of the text data matches the text data; the initial multimedia data includes at least one multimedia fragment, each of which corresponds to at least one text fragment divided by the text data; a target multimedia fragment in the at least one multimedia fragment corresponds to a target text fragment in the at least one text fragment; the target multimedia fragment includes a target video fragment and a target audio fragment, where the target video fragment includes a video image that matches the target text fragment and the target audio fragment includes spoken audio that matches the target text fragment; obtain a target clipping template in response to a clipping template acquisition request; apply a clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data; and generate a target video based on the target multimedia data.

[0122] Optionally, when the one or more programs are executed by the terminal device, they can cause the terminal device to perform other steps described in the above embodiments.

[0123] Computer program code for carrying out the operations of the present disclosure can be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​(such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages ​​(such as "C" or similar programming languages). The program code may run entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, such as a local area network (LAN) or wide area network (WAN), or may be connected to an external computer (e.g., connected via the Internet using an Internet Service Provider).

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowcharts or block diagrams may represent a module, program segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions depicted in the boxes may occur in a different order than that depicted in the accompanying drawings. For example, two boxes depicted in succession may actually be executed substantially in parallel, or may be executed in the reverse order, depending on the functionality involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs the specified function or operation, or in a combination of dedicated hardware and computer instructions.

[0125] The units described in the embodiments of the present disclosure may be implemented by software or hardware, and the names of the units do not constitute limitations on the units themselves in a given situation.

[0126] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.

[0127] In the context of this disclosure, a computer-readable storage medium may be a tangible medium that contains or stores a program used by or in combination with an instruction execution system, apparatus, or device. A computer-readable storage medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium includes, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of computer-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0128] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, the method including: generating initial multimedia data based on received text data, wherein the initial multimedia data includes a video image in which spoken audio of the text data matches the text data; the initial multimedia data includes at least one multimedia fragment, each of the at least one multimedia fragment corresponding to at least one text fragment separated by the text data; a target multimedia fragment in the at least one multimedia fragment corresponds to a target text fragment in the at least one text fragment, the target multimedia fragment including a target video fragment and a target audio fragment, the target video fragment including a video image matching the target text fragment, and the target audio fragment including spoken audio matching the target text fragment; obtaining a target clipping template in response to a clipping template acquisition request; applying a clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data; and generating a target video based on the target multimedia data.

[0129] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein a video image includes subtitle text that matches a target text fragment.

[0130] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein obtaining a target clipping template in response to a clipping template acquisition request includes determining a clipping template corresponding to the triggering operation as a target clipping template in response to a triggering operation of a template theme control, and obtaining the target clipping template.

[0131] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein, prior to responding to a triggering operation of a clipping template control, the method further includes displaying a video editing area, wherein the video editing area includes a template control, and, in response to the triggering operation of the template control, displaying a mask area and displaying at least one template theme control in the mask area.

[0132] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein the clipping operations indicated by the target clipping template include video compositing operations, and applying the clipping operations indicated by the target clipping template to initial multimedia data to obtain the target multimedia data includes compositing video fragments included in the target clipping template and multimedia fragments included in the initial multimedia data based on the video compositing operations to obtain the target multimedia data.

[0133] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein synthesizing a video fragment included in a target clipping template and a multimedia fragment included in initial multimedia data based on a video synthesis operation to obtain target multimedia data includes loading a video fragment included in the target clipping template into a set position of a multimedia fragment included in the initial multimedia data based on the video synthesis operation to obtain the target multimedia data, wherein the set position includes before the first frame media data of the initial multimedia data and / or after the last frame media data of the initial multimedia data.

[0134] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein the clipping operation indicated by the target clipping template includes a transition setting operation, and applying the clipping operation indicated by the target clipping template to initial multimedia data to obtain the target multimedia data includes adding a transition effect to a multimedia fragment included in the initial multimedia data based on the transition setting operation to obtain the target multimedia data.

[0135] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein the clipping operation indicated by the target clipping template includes a virtual object adding operation, and applying the clipping operation indicated by the target clipping template to initial multimedia data to obtain target multimedia data includes adding a virtual object included in the target clipping template to a preset position of the initial multimedia data based on the virtual object adding operation to obtain the target multimedia data.

[0136] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein the clipping operation indicated by the target clipping template includes a background audio addition operation, and applying the clipping operation indicated by the target clipping template to initial multimedia data to obtain target multimedia data includes mixing the background audio included in the target clipping template and the speech audio included in the initial multimedia data based on the background audio addition operation to obtain the target multimedia data.

[0137] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein the clipping operation indicated by the target clipping template includes a keyword extraction operation, and applying the clipping operation indicated by the target clipping template to the initial multimedia data includes, for at least one target text fragment, extracting keywords in the target text fragment and adding the keywords to the target multimedia fragment corresponding to the target text fragment.

[0138] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein adding a keyword to a target multimedia fragment corresponding to a target text fragment includes obtaining key text information that matches the keyword, and adding the keyword and the key text information to the target multimedia fragment corresponding to the target text fragment.

[0139] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, the device including: an initial multimedia data generation module for generating initial multimedia data based on received text data, where the initial multimedia data includes a video image in which a spoken voice of the text data matches the text data, the initial multimedia data includes at least one multimedia fragment, each of the at least one multimedia fragment corresponding to at least one text fragment divided by the text data, a target multimedia fragment in the at least one multimedia fragment corresponds to a target text fragment in the at least one text fragment, the target multimedia fragment includes a target video fragment and a target audio fragment, the target video fragment includes a video image that matches the target text fragment, and the target audio fragment includes a spoken voice that matches the target text fragment, and a target clipping template acquisition module for acquiring a target clipping template in response to a clipping template acquisition request; The system includes a target multimedia data generation module used to apply clipping operations indicated by the target clipping template to the initial multimedia data to obtain target multimedia data, and a target video generation module used to generate a target video based on the target multimedia data.

[0140] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, wherein a video image includes subtitle text that matches a target text fragment.

[0141] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, in which a target clipping template acquisition module is used to acquire a target clipping template in response to a clipping template acquisition request, a target clipping template determination unit is used to determine a clipping template corresponding to the trigger operation as a target clipping template in response to a trigger operation of a template theme control, and the target clipping template acquisition unit is used to acquire the target clipping template.

[0142] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, wherein the target clipping template acquisition module further includes: a video editing area display unit for displaying a video editing area before responding to a triggering operation of a clipping template control, wherein the video editing area includes a template control; and a mask area display unit for displaying a mask area and displaying at least one template theme control in the mask area in response to the triggering operation of the template control.

[0143] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, wherein the clipping operation indicated by the target clipping template includes a video compositing operation, and the target multimedia data generation module is specifically used to synthesize video fragments included in the target clipping template and multimedia fragments included in the initial multimedia data based on the video compositing operation to obtain target multimedia data.

[0144] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, wherein a target multimedia data generation module is specifically used to load a video fragment included in a target clipping template into a set position of a multimedia fragment included in initial multimedia data based on a video synthesis operation to obtain target multimedia data, wherein the set position includes before the first frame media data of the initial multimedia data and / or after the last frame media data of the initial multimedia data.

[0145] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, wherein the clipping operation indicated by the target clipping template includes a transition setting operation, and the target multimedia data generation module is specifically used to add a transition effect to the multimedia fragment included in the initial multimedia data based on the transition setting operation to obtain the target multimedia data.

[0146] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, wherein the clipping operation indicated by the target clipping template includes a virtual object adding operation, and the target multimedia data generation module is specifically used to add the virtual object included in the target clipping template to a preset position of the initial multimedia data based on the virtual object adding operation to obtain the target multimedia data.

[0147] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, wherein the clipping operation indicated by the target clipping template includes a background audio addition operation, and the target multimedia data generation module is specifically used to obtain target multimedia data by mixing the background audio included in the target clipping template and the reading audio included in the initial multimedia data based on the background audio addition operation.

[0148] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, wherein the clipping operation indicated by the target clipping template includes a keyword extraction operation, and the target multimedia data generation module is specifically used for at least one target text fragment to extract keywords in the target text fragment and add the keywords to the target multimedia fragment corresponding to the target text fragment.

[0149] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, in which a target multimedia data generation module is specifically used to obtain key text information matching a keyword, and add the keyword and the key text information to a target multimedia fragment corresponding to the target text fragment.

[0150] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, one or more processors; a memory for storing one or more programs; The one or more programs, when executed by the one or more processors, cause the one or more processors to perform any one of the video generation methods provided by this disclosure.

[0151] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, causing the computer to perform any one of the video generation methods provided by the present disclosure.

[0152] An embodiment of the present disclosure further provides a computer program product, the computer program product including computer programs or instructions that, when executed by a processor, cause the computer program or instructions to perform the above video generation method.

[0153] The above description is an illustrative example of the preferred embodiments of the present disclosure and the technical principles employed. Those skilled in the art should understand that the scope of the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the scope of the present disclosure. For example, it also covers (but is not limited to) technical solutions formed by replacing the above features with technical features having similar functions disclosed in the present disclosure.

[0154] Additionally, although operations are depicted using a particular order, this should not be construed as requiring the operations to be performed in the particular order shown or sequentially. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination.

[0155] Although the present subject matter has been described using language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. 1. A video generation method comprising: generating initial multimedia data based on received text data, the initial multimedia data including a video image whose spoken audio of the text data matches the text data, the initial multimedia data including at least one multimedia fragment, each of the at least one multimedia fragment corresponding to at least one text fragment divided by the text data, a target multimedia fragment in the at least one multimedia fragment corresponding to a target text fragment in the at least one text fragment, the target multimedia fragment including a target video fragment and a target audio fragment, the target video fragment including a video image that matches the target text fragment, and the target audio fragment including spoken audio that matches the target text fragment; Obtaining a target clipping template in response to a clipping template obtainment request; applying a clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data; and generating a target video based on the target multimedia data; The clipping operation includes at least one of a video synthesis operation, a transition setting operation, a virtual object addition operation, a background sound addition operation, and a keyword extraction operation. A method characterized by:

2. The method of claim 1 , wherein the video image includes subtitle text that matches the target text fragment.

3. Obtaining a target clipping template in response to the clipping template obtainment request includes: In response to a trigger operation of a template theme control, determining a clipping template corresponding to the trigger operation as the target clipping template; The method of claim 1 , further comprising obtaining the target clipping template.

4. Before responding to a Clipping Template Control trigger operation, displaying a video editing area, wherein said video editing area includes a template control; displaying a mask region in response to a trigger operation of the template control; The method of claim 3 , further comprising displaying at least one template theme control in the masked region.

5. the clipping operation indicated by the target clipping template includes a video compositing operation; applying a clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data, 2. The method of claim 1, further comprising: synthesizing video fragments included in the target clipping template and multimedia fragments included in the initial multimedia data based on the video synthesis operation to obtain target multimedia data.

6. synthesizing the video fragments included in the target clipping template and the multimedia fragments included in the initial multimedia data based on the video synthesis operation to obtain target multimedia data; The method of claim 5, further comprising: loading a video fragment included in the target clipping template into a set position of a multimedia fragment included in the initial multimedia data based on the video compositing operation to obtain target multimedia data, wherein the set position comprises before a first frame media data of the initial multimedia data and / or after a last frame media data of the initial multimedia data.

7. the clipping operations indicated by the target clipping template include a transition setting operation; applying a clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data, 2. The method of claim 1, further comprising adding a transition effect to a multimedia fragment included in the initial multimedia data based on the transition setting operation to obtain target multimedia data.

8. the clipping operation indicated by the target clipping template includes a virtual object adding operation; applying a clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data, 2. The method of claim 1, further comprising: directly adding a virtual object included in the target clipping template to a preset position of the initial multimedia data based on the virtual object adding operation to obtain target multimedia data.

9. the clipping operations indicated by the target clipping template include a background audio addition operation; applying a clipping operation indicated by the target clipping template to the initial multimedia data to obtain target multimedia data, 2. The method of claim 1, further comprising: mixing the background audio included in the target clipping template and the speech audio included in the initial multimedia data based on the background audio addition operation to obtain target multimedia data.

10. the clipping operation indicated by the target clipping template includes a keyword extraction operation; applying a clipping operation indicated by the target clipping template to the initial multimedia data, for at least one target text fragment, extracting keywords in the target text fragment; 2. The method of claim 1, further comprising adding the keywords to a target multimedia fragment corresponding to the target text fragment.

11. adding the keywords to a target multimedia fragment corresponding to the target text fragment, obtaining key text information that matches the keywords; 11. The method of claim 10, further comprising adding the keywords and the key text information to a target multimedia fragment corresponding to the target text fragment.

12. an initial multimedia data generating module for generating initial multimedia data based on received text data, wherein the initial multimedia data includes a video image whose spoken voice of the text data matches the text data, the initial multimedia data includes at least one multimedia fragment, each of the at least one multimedia fragment corresponding to at least one text fragment divided by the text data, a target multimedia fragment in the at least one multimedia fragment corresponding to a target text fragment in the at least one text fragment, the target multimedia fragment including a target video fragment and a target audio fragment, the target video fragment including a video image that matches the target text fragment, and the target audio fragment including spoken voice that matches the target text fragment; a target clipping template acquisition module for acquiring a target clipping template in response to the clipping template acquisition request; a target multimedia data generation module for applying clipping operations indicated by the target clipping template to the initial multimedia data to obtain target multimedia data; a target video generation module for generating a target video based on the target multimedia data; The video generating device, wherein the clipping operation includes at least one of a video synthesis operation, a transition setting operation, a virtual object addition operation, a background sound addition operation, and a keyword extraction operation.

13. one or more processors; a storage device for storing one or more programs; An electronic device, characterized in that, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having stored thereon a computer program, the computer program causing a processor to perform the method of any one of claims 1 to 11 when the computer program is executed by the processor.

15. A computer program comprising instructions which, when executed by a processor, causes the computer program to carry out the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Multimedia data processing method and device, electronic equipment and storage medium

    CN109756751A

  • Multimedia file material processing method and device, electronic equipment and storage medium

    CN112449231A

  • Video generation method and device, electronic equipment and storage medium

    CN113452941A

  • Video generation method and device, computer equipment and storage medium

    CN113473182A

  • Generation device, generation method, and generation program

    JP2021033367A