Video generation method and apparatus, and electronic device and storage medium

By receiving user instructions to generate model input instructions, and automatically generate videos using the target model and video editing tools, the problems of complex operation and low editing efficiency of traditional video editing solutions are solved, and efficient video generation is achieved.

WO2025108303A1PCT designated stage expired Publication Date: 2025-05-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/133207
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-11-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional video editing schemes are complex in operation and low in editing efficiency. Users need to manually edit videos, making it difficult to achieve efficient video generation.

Method used

By receiving user instructions, generating model input instructions, using the target model to obtain video clip parameters, and automatically generate target videos with video materials and target video editing tools.

Benefits of technology

It realizes that users can generate videos in one click, significantly improves video editing efficiency and simplifies the operation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133207_30052025_PF_FP_ABST
    Figure CN2024133207_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a video generation method and apparatus, and an electronic device and a storage medium. The video generation method comprises: receiving a user instruction input by a user; generating a model input instruction on the basis of the user instruction; obtaining target information on the basis of the model input instruction and a target model, wherein the target information includes a video editing parameter for a target video editing tool; and generating a target video on the basis of the target information, a video material and the target video editing tool, wherein the video material is associated with the video editing parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation method, device, electronic device and storage medium

[0001] This application claims priority to Chinese patent application No. 202311558669.8 filed on November 21, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby cited in their entirety as part of this application. Technical Field

[0002] Embodiments of the present disclosure relate to a video generation method, apparatus, electronic device, and storage medium. Background Art

[0003] Traditional video editing solutions require users to manually edit videos. For example, users manually edit video and audio tracks through the multi-track editing panel of video editing software, and manually add filters, stickers, special effects, etc. This solution has problems such as complex operation and low editing efficiency. Summary of the Invention

[0004] One or more embodiments of the present disclosure provide a video generation method, including:

[0005] Receive user instructions input by the user;

[0006] generating a model input instruction based on the user instruction;

[0007] Obtaining target information based on the model input instruction and the target model, wherein the target information includes video editing parameters for a target video editing tool;

[0008] A target video is generated based on the target information, the video material and the target video editing tool; wherein the video material is associated with the video editing parameters.

[0009] One or more embodiments of the present disclosure provide a video generation device, including:

[0010] An instruction receiving unit, configured to receive a user instruction input by a user;

[0011] an instruction generating unit, configured to generate a model input instruction based on the user instruction;

[0012] an information generating unit configured to obtain target information based on the model input instruction and the target model, wherein the target information includes video editing parameters for a target video editing tool;

[0013] The video generating unit is configured to generate a target video based on the target information, the video material and the target video editing tool; wherein the video material is associated with the video editing parameters.

[0014] One or more embodiments of the present disclosure provide an electronic device, comprising: at least one memory and at least one processor; wherein the memory is configured to store program code, and the processor is configured to call the program code stored in the memory to enable the electronic device to execute the method provided according to one or more embodiments of the present disclosure.

[0015] One or more embodiments of the present disclosure provide a non-transitory computer storage medium storing program code. When the program code is executed by a computer device, the computer device executes the method provided according to one or more embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0017] FIG1 is a flow chart of a video generation method provided by an embodiment of the present disclosure;

[0018] FIG2 is a flow chart of a video generation method provided by another embodiment of the present disclosure;

[0019] FIG3 is a schematic structural diagram of a video generating device provided by an embodiment of the present disclosure; and

[0020] FIG4 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0022] It should be understood that the steps described in the embodiments of the present disclosure can be performed in a different order and / or in parallel. In addition, the embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0023] As used herein, the term "including" and its variations are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The term "responsive to" and related terms refer to a signal or event being affected to a certain extent by another signal or event, but not necessarily completely or directly. If event x occurs "responsive to" event y, then x may be directly or indirectly responsive to y. For example, the occurrence of y may ultimately lead to the occurrence of x, but there may be other intermediate events and / or conditions. In other cases, y may not necessarily lead to the occurrence of x, and x may occur even if y has not yet occurred. In addition, the term "responsive to" may also mean "at least partially responsive to".

[0024] The term "determine" broadly encompasses a variety of actions, and may include obtaining, calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, database, or other data structure), ascertaining, and similar actions, and may also include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and similar actions, as well as parsing, selecting, choosing, establishing, and similar actions. Other terms are defined below. Other terms are defined below.

[0025] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the provisions of relevant laws and regulations.

[0026] It is understood that before using the technical solutions disclosed in each embodiment of the present disclosure, the user should be informed of the type, scope of use, and usage scenarios of the personal information involved in the present disclosure and obtain the user's authorization in an appropriate manner in accordance with relevant laws and regulations. For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information, so that the user can independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.

[0027] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0028] It is understood that images generated by the methods provided in various embodiments of the present disclosure should be processed in accordance with relevant laws and regulations. For example, technical measures may be taken in accordance with regulations to add a logo that does not affect user use, or prominent logos may be placed in reasonable locations and areas in accordance with regulations to inform the public of deep compositing.

[0029] It is understandable that the above notification and user authorization process, as well as image processing, are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0032] For the purposes of this disclosure, the phrase "A and / or B" means (A), (B), or (A and B).

[0033] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0034] 1 , which shows a flowchart of a video generation method 100 provided in an embodiment of the present disclosure. The method 100 includes steps S110 to S140 .

[0035] Step S110: receiving a user instruction input by a user.

[0036] Step S120: Generate a model input instruction based on the user instruction.

[0037] In some embodiments, a user instruction input by a user may be received through a chat session window. For example, the user may input the user instruction through text editing or voice input provided by the robot chat session window, but the present disclosure is not limited thereto.

[0038] In some embodiments, the user instruction input by the user may include a natural language instruction input by the user. The natural language instruction is used to indicate the theme and / or style of the video to be generated. For example, the user instruction may be "make a video about puppies" or "make a cute and warm video", etc., but the present disclosure is not limited thereto.

[0039] In some embodiments, the user instruction may further include an instruction indicating video display parameters, wherein the video display parameters include at least one of video duration, video resolution, video aspect ratio, and number of video segments.

[0040] In some embodiments, the model input instruction is an instruction for adapting the target model input, which conforms to the input format requirements of the target model or conforms to the training data distribution of the target model. In a specific embodiment, the model input instruction includes one or more preset fields, including but not limited to an input field, an audio (e.g., music) field, an aspect ratio field, a video duration field, and a video clip number field. The value of the one or more preset fields can be determined based on the user instruction input by the user, and the fields that cannot be determined based on the user instruction can adopt default values ​​or be assigned values ​​through a natural language model.

[0041] In some embodiments, the user instructions are adjusted and / or expanded to obtain model input instructions that adapt to the target model; wherein the adjustment processing includes adjusting the expression of the user instructions, and the expansion processing includes assigning values ​​to preset fields in the model input instructions.

[0042] In some embodiments, the user instruction can be expanded by a natural language model to obtain a model input instruction. The natural language model includes but is not limited to a Transformer-based model, an autoencoder-based model, a sequence-to-sequence model, a recursive neural network model, or a hierarchical model, but the present disclosure is not limited thereto. Exemplarily, the user instruction or the text generated based on the user instruction can be used as the input of the natural language model (for example, the prompt word "how to make a video about a puppy"), so that the natural language model recommends relevant video styles (such as "warm and cute"), video soundtracks, etc. based on the input prompt words, thereby expanding the user instruction to obtain the model input instruction text.

[0043] For example, the model input instruction may be as follows:

[0044] Input: Puppy

[0045] Music:**

[0046] Args: height:width = 3:4; duration, 13s

[0047] Segments:1

[0048] In some embodiments, keywords in user instructions can be extracted using a keyword extraction algorithm, and a model input instruction can be generated based on the keywords. Keyword extraction algorithms include, but are not limited to, word frequency-inverse document frequency algorithms, graph-based ranking algorithms for text, text analysis methods based on probabilistic graph models, keyword extraction algorithms based on statistics and natural language processing techniques, and keyword extraction algorithms based on machine learning techniques, and this disclosure does not limit these.

[0049] In some embodiments, the user instructions input by the user have a preset format. For example, the instructions input by the user may be required to correspond one-to-one with preset fields, so that model input instructions can be generated based on the values ​​input by the user under these preset fields.

[0050] Step S130: Obtain target information based on the model input instruction and the target model, where the target information includes video editing parameters for a target video editing tool.

[0051] In some embodiments, the target information includes structured text applicable to the target video editing tool, which includes video editing parameters readable by the target video editing tool. In some embodiments, the video editing parameters are used to control the style and / or time range of the video material presented in the video. In some embodiments, the video editing parameters include parameters involved in the video editing project, such as including one or more of the following: canvas size (such as width, height), video duration, transition mode, video resolution, number of video clips, identification (such as ID or key value) of the video material, the time range of the video material, the rotation mode of the video material, the scaling ratio of the video material. The video material includes but is not limited to images, music, stickers, animations, filters, special effects, masks, texts, or video templates.

[0052] The video template includes a collection of pre-configured video assets (e.g., music, stickers, animations, filters, special effects, masks, or text). After importing the original image or video, the video template can directly generate a video with added assets such as music, stickers, animations, filters, special effects, masks, or text. In some embodiments, the relevant video assets are arranged in a preset track or time range within the video template.

[0053] In some embodiments, the target information includes structured text describing the video editing project, which may describe relevant information about the video editing project (such as the aforementioned video editing parameters or related information) based on a protocol associated with the video editing tool. Exemplarily, the format of the target information is JSON (JavaScript Object Notation).

[0054] In some embodiments, the target model includes a text-to-text model, such as a large language model based on deep learning, including but not limited to an autoencoder-based model, a sequence-to-sequence model, a Transformer-based model, a recursive neural network model, and a hierarchical model, but the present disclosure is not limited thereto.

[0055] In some embodiments, when training a target model, a model input instruction sample can be obtained based on the label, title, description, soundtrack information (such as audio metadata) corresponding to the sample video or a keyword (query) used to search for the sample video as the model input, and the target information sample corresponding to the sample video (such as structured text describing the editing process of the sample video) can be used as the model's expected output to train the target model. Exemplarily, the target model parameters can be adjusted based on the loss between the model prediction result obtained by inputting the input instruction sample into the target model and the expected output until the loss function converges, but the present disclosure is not limited to this.

[0056] In some embodiments, the tags and descriptions of a video may include, but are not limited to, information about the video's theme, content, style, and filters. In some video publishing scenarios, when a user publishes a video, they may add relevant tags and descriptions to the video and publish them together with the video. These tags and descriptions can serve to introduce, categorize, or search for the video.

[0057] In some embodiments, at least one of the label, title, description, music information, and search keywords corresponding to the sample video can be directly used as an input instruction sample of the target model.

[0058] In some embodiments, at least one of the label, title, description, music information, or search keywords can be rewritten. For example, at least one of the label, title, description, music information, or search keywords can be input into a preset language model and rewritten to obtain a rewritten input instruction sample. In this embodiment, by rewriting the input instruction sample to perform data augmentation, the convergence of the target model can be accelerated, thereby improving the model training speed.

[0059] In some embodiments, the target video editing tool includes but is not limited to software, applications, applets, HTML (Hyper Text Markup Language) pages or services for video editing, which can be loaded locally on the terminal or in a cloud server, and the present disclosure does not impose any restrictions on this.

[0060] Step S140: generating a target video based on the target information, the video material and the target video editing tool; wherein the video material is associated with the video editing parameters.

[0061] The video material can be pre-stored in the target video editing tool, or downloaded from the server, or provided or selected by the user, or generated in real time based on user instructions, and this patent is not limited here. In some embodiments, the user can actively provide video material, for example, by uploading multimedia material through the aforementioned robot chat session window as video material. In some embodiments, candidate video materials can also be provided to the user for selection, and the video material can be determined based on the user's selection. In some embodiments, the video material can be determined from the local resource library of the target video editing tool or the server can obtain the video material based on the user's instructions. In some embodiments, the video material can be generated in real time based on the user's instructions, such as through a text graph model.

[0062] In this way, according to one or more embodiments of the present disclosure, by generating model input instructions based on user instructions, obtaining target information containing video editing parameters readable by the target video editing tool based on the model input instructions and the target model, and generating target video based on the target information, the target video editing tool and the video material associated with the video editing parameters, it is possible to ultimately achieve one-click generation of videos based on user instructions, thereby significantly improving video editing efficiency.

[0063] In some embodiments, multimedia materials input by the user can also be received, target tags of the multimedia materials can be obtained, and model input instructions can be generated based on the user instructions and the target tags, wherein the target tags are used to indicate the theme and / or style of the video to be generated. In this embodiment, the multimedia materials input by the user can be further combined to judge the user's theme or style of the video to be generated. For example, in a specific application scenario, the natural language instruction input by the user may only be a text instruction "cut a video" and input a corresponding image (such as a photo of a kitten or puppy). In this case, the user's video editing intention can be judged in combination with the image input by the user.

[0064] In some embodiments, a target detection algorithm can be used to identify multimedia material input by the user (e.g., an image or frame) to obtain a target label / description of the target object (e.g., the label "puppy"), and obtain a model input instruction based on the user instruction and the target label. The target detection algorithm includes but is not limited to a detection algorithm based on a multimodal Transformer model, a target detection algorithm based on a region-based convolutional neural network, a single-step multi-frame target detection algorithm, a target detection algorithm based on a retinal network, and the like.

[0065] In some embodiments, the user instruction and the target label can be input together into a Vincent model, such as a natural language model based on deep learning, to obtain a model input instruction. The Vincent model includes but is not limited to a Transformer-based model, an autoencoder-based model, a sequence-to-sequence model, a recursive neural network model, and a hierarchical model, but the present disclosure is not limited thereto. For example, in the aforementioned specific application scenario, in response to the user inputting the user instruction "cut a video" and a puppy image, a model input instruction can be finally obtained based on the aforementioned target detection algorithm and the Vincent model. The model input instruction includes information indicating that the video theme is "puppy" or "pet" and / or information indicating that the video style is "warm and cute", but the present disclosure is not limited thereto.

[0066] In some embodiments, the multimedia material input by the user includes image material (such as a static image, a dynamic image, or a video) or audio material, etc.

[0067] In some embodiments, step S120 includes:

[0068] Step A1: generating a first target instruction based on a user instruction, where the first target instruction includes information indicating a theme and / or style of a video to be generated;

[0069] Step A2: generating a second target instruction based on at least one of the user instruction, the first target instruction, and the multimedia material input by the user, wherein the second target instruction includes information indicating the audio of the video to be generated;

[0070] Step A3: Determine a model input instruction based on the first target instruction and the second target instruction.

[0071] In this embodiment, a second target instruction is further generated based on the user instruction, the first target instruction and at least one of the multimedia materials input by the user to determine the user's audio usage intention (such as the intention to create a soundtrack), so that the model input instruction can include information for indicating the audio, thereby making the final generated target video conform to the user's intention to create a soundtrack, thereby improving the efficiency of video soundtrack processing.

[0072] In some embodiments, the information indicating the audio for which the video is to be generated includes metadata of the audio, which is used to describe relevant information such as music, album or performer, including but not limited to song title, release date, track number, album name or performer name.

[0073] In some embodiments, a search keyword can be determined based on at least one of a user instruction, a first target instruction, and multimedia material input by the user; a target audio can be searched for based on the search keyword; and a second target instruction can be generated based on the target audio.

[0074] In some embodiments, a target detection algorithm can be used to identify multimedia material input by the user to obtain a target label / description of the target object (e.g., the label "puppy"), and a second target instruction can be obtained based on at least one of the user instruction, the first target instruction, and the multimedia material input by the user. The target detection algorithm includes but is not limited to a detection algorithm based on a multimodal Transformer model, a region-based convolutional neural network target detection algorithm, a single-step multi-frame target detection algorithm, a retinal network-based target detection algorithm, and the like.

[0075] In some embodiments, based on the user instruction (or first target instruction) and the target tag, an input text (such as a prompt word) suitable for a preset natural language processing model can be obtained, and the input text can be input into the natural language processing model to obtain the search keyword (such as the song name) recommended by the model, and the search keyword is used as a query word through a search engine to determine the target audio, and a second target instruction is generated based on the relevant information of the target audio (such as metadata). For example, in a specific application scenario, the user inputs the user instruction of "cut a video about a puppy", then the preset natural language processing model can be asked "please recommend a piece of music related to a puppy" to obtain the search keyword; in another specific application scenario, the user inputs the user instruction of "cut a video" and inputs an image containing a puppy, then the puppy in the image can be identified by the aforementioned target detection algorithm, and finally the input text of the natural language processing model of "please recommend a piece of music related to a puppy" is generated.

[0076] In some embodiments, the first target instruction is a model input instruction that does not include an audio instruction. The specific implementation of generating the first target instruction based on the user instruction in step A1 can refer to the above description of generating the model input instruction based on the user instruction, which will not be repeated here.

[0077] In some embodiments, step A2 includes: performing preset processing on the target audio, the preset processing including excerpt processing and / or segmentation processing; and generating a second target instruction based on the target audio and the preset processing result.

[0078] The excerpting process includes extracting a segment of a certain duration from the target audio. For example, the excerpting process can be performed based on the duration of the video or the duration of the video segment in the video.

[0079] The segmentation process includes segmenting the entire audio or the selected audio. For example, the audio can be segmented based on the number of image materials input by the user.

[0080] In some embodiments, the second target instruction includes one or more audio clip parameters of an audio identifier (eg, audio name, audio ID), an audio time range, and a number of audio segments.

[0081] For example, referring to FIG2 , the target detection algorithm can be used to identify the multimedia material input by the user to obtain a target label / description of the target object (e.g., the label “puppy”), and the user instruction (or first target instruction) and the target label are input into the natural language processing model to obtain the audio keywords (e.g., music titles) or template keywords (e.g., video template titles) recommended by the model. Finally, an audio search can be performed based on the audio keywords to obtain the target audio; alternatively, a template search can be performed based on the template keywords to obtain the target video template, and the audio used by the target video template is used as the target audio.

[0082] In some embodiments, step S130 includes:

[0083] Step B1: Inputting the model input instruction into the target model to obtain initial target information. In some embodiments, the initial target information is decoupled from the target video editing tool;

[0084] Step B2: Generate target information based on the initial target information and the protocol associated with the target video editing tool.

[0085] In this embodiment, the output of the target model (i.e., the initial target information) is decoupled from the target video editing tool, so the target model is not only adapted to the target video editing tool, thereby improving the versatility and applicability of the target model. When the target video editing tool is used, target information suitable for the target video editing tool can be further obtained based on the initial target information and the protocol associated with the target video editing tool.

[0086] In some embodiments, not all video editing parameters rely on target model prediction. To this end, the values ​​of some video editing parameters in the initial target information output by the target model may be empty. These video editing parameters with empty values ​​may be assigned preset values ​​(such as default values) or values ​​adapted to the target video editing tool in the target information, but the present disclosure is not limited to this.

[0087] Accordingly, referring to FIG3 , according to an embodiment of the present disclosure, a video generating apparatus 400 is provided, including:

[0088] The instruction receiving unit 301 is configured to receive a user instruction input by a user;

[0089] The instruction generation unit 302 is configured to generate a model input instruction based on the user instruction;

[0090] An information generating unit 303 is configured to obtain target information based on the model input instruction and the target model, wherein the target information includes video editing parameters for a target video editing tool;

[0091] The video generating unit 304 is configured to generate a target video based on target information, video material and target video editing tools; wherein the video material is associated with video editing parameters.

[0092] In some embodiments, the video generating apparatus further comprises:

[0093] a material receiving unit configured to receive multimedia material input by a user and determine a target object in the multimedia material;

[0094] The instruction generation unit is configured to generate a model input instruction based on a user instruction and a target object, wherein the target object is used to indicate a theme and / or style of a video to be generated.

[0095] In some embodiments, the instruction generation unit includes:

[0096] A first instruction generating subunit is configured to generate a first target instruction based on a user instruction, wherein the first target instruction includes information indicating a theme and / or style of a video to be generated;

[0097] A second instruction generating subunit is configured to generate a second target instruction based on at least one of the user instruction, the first target instruction, and the multimedia material input by the user, wherein the second target instruction includes information indicating the audio of the video to be generated;

[0098] The instruction generation subunit is configured to determine a model input instruction based on the first target instruction and the second target instruction.

[0099] In some embodiments, the second instruction generation subunit includes:

[0100] a keyword subunit configured to determine a search keyword based on at least one of a user instruction, a first target instruction, and multimedia material input by the user;

[0101] The audio subunit is configured to search for target audio based on a search keyword;

[0102] The instruction subunit is configured to generate a second target instruction based on the target audio.

[0103] In some embodiments, the search keyword includes a template keyword; the audio subunit is configured to perform a template search based on the template keyword to obtain a target video template, and use the audio used by the target video template as the target audio.

[0104] In some embodiments, the instruction subunit is configured to perform preset processing on the target audio, the preset processing including excerpt processing and / or segmentation processing; and generate a second target instruction based on the target audio and the preset processing result.

[0105] In some embodiments, the information generation unit includes:

[0106] an initial information generating subunit, configured to input the model input instruction into the target model to obtain initial target information;

[0107] The target information generating subunit is configured to generate target information based on the initial target information and a protocol associated with the target video editing tool.

[0108] In some embodiments, the video generation apparatus further includes a model training unit, including:

[0109] A sample acquisition subunit is configured to acquire an input instruction sample and a target information sample corresponding to a sample video;

[0110] The training subunit is configured to take the input instruction sample as the input of the target model and the target information sample as the expected output to train the target model.

[0111] In some embodiments, the sample acquisition subunit is configured to determine the input instruction sample based on at least one of a tag, a title, a description, audio information, and a search keyword associated with the sample video.

[0112] For the embodiments of the device, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, and the modules described as separation modules may or may not be separate. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art can understand and implement it without paying any creative work.

[0113] Accordingly, according to one or more embodiments of the present disclosure, there is provided an electronic device, including:

[0114] at least one memory and at least one processor;

[0115] The memory is configured to store program codes, and the processor is configured to call the program codes stored in the memory to enable the electronic device to execute the video generation method provided according to one or more embodiments of the present disclosure.

[0116] Accordingly, according to one or more embodiments of the present disclosure, a non-transitory computer storage medium is provided, which stores program code. The program code can be executed by a computer device to enable the computer device to perform the video generation method provided according to one or more embodiments of the present disclosure.

[0117] Reference is now made to FIG4 , which illustrates a schematic diagram of the structure of an electronic device (e.g., a terminal device or server) 800 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG4 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.

[0118] As shown in Figure 4, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0119] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although FIG4 shows the electronic device 800 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may be implemented or present instead.

[0120] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0121] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0122] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0123] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0124] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method of the present disclosure.

[0125] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0127] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0128] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0129] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0130] According to one or more embodiments of the present disclosure, a video generation method is provided, including: receiving a user instruction input by a user; generating a model input instruction based on the user instruction; obtaining target information based on the model input instruction and a target model, the target information including video editing parameters for a target video editing tool; generating a target video based on the target information, a video material and a target video editing tool; wherein the video material is associated with the video editing parameters.

[0131] The method provided according to one or more embodiments of the present disclosure further includes: receiving multimedia materials input by a user, and determining a target object in the multimedia materials; generating a model input instruction based on the user instruction, including: generating a model input instruction based on the user instruction and the target object, wherein the target object is used to indicate the theme and / or style of the video to be generated.

[0132] According to one or more embodiments of the present disclosure, generating a model input instruction based on a user instruction includes: generating a first target instruction based on the user instruction, the first target instruction including information indicating the theme and / or style of the video to be generated; generating a second target instruction based on the user instruction, the first target instruction and at least one of the multimedia materials input by the user, the second target instruction including information indicating the audio of the video to be generated; and determining the model input instruction based on the first target instruction and the second target instruction.

[0133] According to one or more embodiments of the present disclosure, a second target instruction is generated based on a user instruction, a first target instruction, and at least one of the multimedia materials input by the user, including: determining a search keyword based on the user instruction, the first target instruction, and at least one of the multimedia materials input by the user; searching for a target audio based on the search keyword; and generating a second target instruction based on the target audio.

[0134] According to one or more embodiments of the present disclosure, the search keyword includes a template keyword; based on the search keyword, the target audio is searched for, including: performing a template search based on the template keyword to obtain a target video template, and using the audio used by the target video template as the target audio.

[0135] According to one or more embodiments of the present disclosure, generating a second target instruction based on the target audio includes: performing preset processing on the target audio, the preset processing including excerpt processing and / or segmentation processing; and generating the second target instruction based on the target audio and the preset processing result.

[0136] According to one or more embodiments of the present disclosure, target information is obtained based on a model input instruction and a target model, including: inputting the model input instruction into the target model to obtain initial target information; and generating target information based on the initial target information and a protocol associated with a target video editing tool.

[0137] In some embodiments, the training method of the target model includes: obtaining input instruction samples and target information samples corresponding to the sample video; using the input instruction samples as the input of the target model and the target information samples as the expected output to train the target model.

[0138] In some embodiments, obtaining an input instruction sample corresponding to a sample video includes determining the input instruction sample based on at least one of a tag, title, description, audio information, and search keyword associated with the sample video.

[0139] According to one or more embodiments of the present disclosure, a video generating device is provided, comprising: an instruction receiving unit configured to receive a user instruction input by a user; an instruction generating unit configured to generate a model input instruction based on the user instruction; an information generating unit configured to obtain target information based on the model input instruction and a target model, wherein the target information includes video editing parameters for a target video editing tool; a video generating unit configured to generate a target video based on the target information, a video material and a target video editing tool; wherein the video material is associated with the video editing parameters.

[0140] According to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one memory and at least one processor; wherein the memory is configured to store program code, and the processor is configured to call the program code stored in the memory so that the electronic device executes the video generation method provided according to one or more embodiments of the present disclosure.

[0141] According to one or more embodiments of the present disclosure, a non-transitory computer storage medium is provided, which stores program code. When the program code is executed by a computer device, the computer device executes the video generation method provided according to one or more embodiments of the present disclosure.

[0142] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0143] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0144] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A video generation method, comprising: Receive user instructions input by the user; generating a model input instruction based on the user instruction; Obtaining target information based on the model input instruction and the target model, wherein the target information includes video editing parameters for a target video editing tool; A target video is generated based on the target information, the video material and the target video editing tool; wherein the video material is associated with the video editing parameters.

2. The method according to claim 1, further comprising: Receiving multimedia material input by a user, and determining a target object in the multimedia material; The generating the model input instruction based on the user instruction includes: generating the model input instruction based on the user instruction and the target object, wherein the target object is used to indicate the theme and / or style of the video to be generated.

3. The method according to claim 1, wherein: The step of generating a model input instruction based on the user instruction comprises: Generating a first target instruction based on the user instruction, wherein the first target instruction includes information indicating a theme and / or style of a video to be generated; Generate a second target instruction based on at least one of the user instruction, the first target instruction, and multimedia material input by the user, wherein the second target instruction includes information indicating the audio of the video to be generated; The model input command is determined based on the first target command and the second target command.

4. The method according to claim 3, wherein: The generating a second target instruction based on at least one of the user instruction, the first target instruction and the multimedia material input by the user comprises: Determine a search keyword based on at least one of the user instruction, the first target instruction, and the multimedia material input by the user; Based on the search keyword, search for the target audio; The second target instruction is generated based on the target audio.

5. The method according to claim 4, wherein: The search keywords include template keywords; The searching for the target audio based on the search keyword includes: performing a template search based on the template keyword to obtain a target video template, and using the audio used by the target video template as the target audio.

6. The method according to claim 4, wherein: The generating the second target instruction based on the target audio comprises: Performing a preset process on the target audio, wherein the preset process includes excerpt processing and / or segment processing; The second target instruction is generated based on the target audio and the preset processing result.

7. The method according to claim 1, wherein: The obtaining target information based on the model input instruction and the target model includes: Inputting the model input instruction into the target model to obtain initial target information; The target information is generated based on the initial target information and a protocol associated with the target video editing tool.

8. The method according to claim 1, wherein: The training method of the target model includes: Obtaining input instruction samples and target information samples corresponding to the sample video; The target model is trained by taking the input instruction sample as the input of the target model and the target information sample as the expected output.

9. The method according to claim 8, wherein: The step of obtaining an input instruction sample corresponding to the sample video includes: The input instruction sample is determined based on at least one of a tag, a title, a description, audio information, and a search keyword associated with the sample video.

10. A video generating device, comprising: An instruction receiving unit, configured to receive a user instruction input by a user; an instruction generating unit, configured to generate a model input instruction based on the user instruction; an information generating unit configured to obtain target information based on the model input instruction and the target model, wherein the target information includes video editing parameters for a target video editing tool; and The video generation unit is configured to generate a target video based on the target information, the video material and the target video editing tool; wherein the video material is associated with the video editing parameters.

11. An electronic device, comprising: at least one memory and at least one processor; The memory is configured to store program codes, and the processor is configured to call the program codes stored in the memory to enable the electronic device to execute the method according to any one of claims 1 to 9.

12. A non-transitory computer storage medium, wherein: The non-transitory computer storage medium stores a program code, and when the program code is executed by a computer device, the computer device executes the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Video generation method, device and equipment, and storage medium

    CN113794930A

  • Video generation method and device, terminal, server and storage medium

    CN114286169A

  • Video clip template searching method and device, electronic equipment and storage medium

    CN117076707A

  • Method and apparatus for automatic mash-up generation

    US20100209003A1

  • Short video generation method and platform, electronic device, and storage medium

    WO2021169459A1