Video generation method and device, electronic equipment and storage medium

By receiving user instructions and generating target information of target video clip parameters, and combining video materials and tools to automatically generate videos, the problems of complex and inefficient operation of traditional video clip solutions are solved, and efficient video generation is achieved.

CN120034693APending Publication Date: 2025-05-23BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311558669.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional video editing solutions are complex in operation and low in editing efficiency, making it difficult to meet users' needs for quickly generating videos.

Method used

By receiving user instructions input by users, generating model input instructions, and obtaining target information containing video clip parameters based on model input instructions and target model. Combining video materials and target video editing tools, the target video is automatically generated.

Benefits of technology

It realizes the generation of videos in one click based on user instructions, significantly improving the video editing efficiency and simplifying the operation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034693A_ABST
    Figure CN120034693A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, electronic equipment and a storage medium. The video generation method comprises the steps of receiving a user instruction input by a user; generating a model input instruction based on the user instruction; obtaining target information based on the model input instruction and the target model, wherein the target information comprises a video editing parameter for a target video editing tool; generating a target video based on the target information, the video material and the target video editing tool; wherein the video material is associated with the video editing parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to a video generation method, apparatus, electronic device, and storage medium. Background Art

[0002] Traditional video editing solutions require users to manually edit videos. For example, users manually edit video and audio tracks through the multi-track editing panel of video editing software, and manually add filters, stickers, special effects, etc. This solution has problems such as complex operations and low editing efficiency. Summary of the Invention

[0003] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the following Detailed Implementation section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, according to one or more embodiments of the present disclosure, there is provided a video generation method, including:

[0005] Receiving a user instruction input by a user;

[0006] Generating a model input instruction based on the user instruction;

[0007] Obtaining target information based on the model input instruction and a target model, where the target information includes video editing parameters for a target video editing tool;

[0008] Generating a target video based on the target information, video materials, and the target video editing tool; wherein, the video materials are associated with the video editing parameters.

[0009] In a second aspect, according to one or more embodiments of the present disclosure, there is provided a video generation apparatus, including:

[0010] An instruction receiving unit, configured to receive a user instruction input by a user;

[0011] An instruction generating unit, configured to generate a model input instruction based on the user instruction;

[0012] An information generating unit, configured to obtain target information based on the model input instruction and a target model, where the target information includes video editing parameters for a target video editing tool;

[0013] A video generating unit, configured to generate a target video based on the target information, video materials, and the target video editing tool; wherein, the video materials are associated with the video editing parameters.

[0014] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one memory and at least one processor; wherein the memory is used to store program code, and the processor is used to call the program code stored in the memory so that the electronic device executes the method provided according to one or more embodiments of the present disclosure.

[0015] In a fourth aspect, according to one or more embodiments of the present disclosure, a non-transitory computer storage medium is provided, wherein the non-transitory computer storage medium stores a program code, and when the program code is executed by a computer device, the computer device executes the method provided according to one or more embodiments of the present disclosure.

[0016] According to one or more embodiments of the present disclosure, by generating model input instructions based on user instructions, obtaining target information including video editing parameters readable by a target video editing tool based on the model input instructions and the target model, and generating a target video based on the target information, the target video editing tool and video materials associated with the video editing parameters, it is ultimately possible to achieve one-click video generation based on user instructions, thereby significantly improving video editing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0018] Figure 1 A flowchart of a video generation method provided by an embodiment of the present disclosure;

[0019] Figure 2 A flowchart of a video generation method provided by another embodiment of the present disclosure;

[0020] Figure 3 A schematic diagram of the structure of a video generating device provided by an embodiment of the present disclosure;

[0021] Figure 4 The figure is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0023] It should be understood that the steps described in the embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0024] As used herein, the term "including" and its variations are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The term "responsive to" and related terms refer to a signal or event being affected to a certain extent by another signal or event, but not necessarily completely or directly. If event x occurs "responsive to" event y, x may be directly or indirectly responsive to y. For example, the occurrence of y may ultimately lead to the occurrence of x, but there may be other intermediate events and / or conditions. In other cases, y may not necessarily lead to the occurrence of x, and x may occur even if y has not yet occurred. In addition, the term "responsive to" may also mean "at least partially responsive to".

[0025] The term "determine" broadly covers a variety of actions, which may include obtaining, calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, database, or other data structure), ascertaining, and similar actions, and may also include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory) and similar actions, as well as parsing, selecting, choosing, establishing, and similar actions, etc. The relevant definitions of other terms will be given in the following description. The relevant definitions of other terms will be given in the following description.

[0026] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the provisions of relevant laws and regulations.

[0027] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, and usage scenarios of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations. For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can autonomously choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0028] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information is sent to the user in a manner such as a pop-up window, in which the prompt information can be presented in text form. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0029] It is understandable that the images generated by the methods provided in the embodiments of the present disclosure should be processed in accordance with the provisions of relevant laws and regulations. For example, technical measures should be taken in accordance with regulations to add a logo that does not affect user use, or a prominent logo should be placed in a reasonable location or area in accordance with regulations to inform the public of the depth synthesis situation.

[0030] It is understandable that the above-mentioned notification and process of obtaining user authorization, as well as the processing of images, are merely illustrative and do not constitute a limitation on the implementation method of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation method of the present disclosure.

[0031] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0032] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0033] For the purposes of this disclosure, the phrase "A and / or B" means (A), (B), or (A and B).

[0034] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0035] refer to Figure 1, which shows a flowchart of a video generation method 100 provided in an embodiment of the present disclosure, and the method 100 includes steps S110 to S140.

[0036] Step S110: receiving a user instruction input by a user.

[0037] Step S120: Generate a model input instruction based on the user instruction.

[0038] In some embodiments, a user instruction input by a user may be received through a chat session window. For example, the user may input the user instruction through text editing or voice input provided by the robot chat session window, but the present disclosure is not limited thereto.

[0039] In some embodiments, the user instruction input by the user may include a natural language instruction arbitrarily input by the user. The natural language instruction is used to indicate the theme and / or style of the video to be generated. For example, the user instruction may be "make a video about a puppy" or "make a cute and warm video", etc., but the present disclosure is not limited thereto.

[0040] In some embodiments, the user instruction may further include an instruction indicating a video display parameter, wherein the video display parameter includes at least one of video duration, video resolution, video aspect ratio, and number of video segments.

[0041] In some embodiments, the model input instruction is an instruction to adapt the target model input, which meets the input format requirements of the target model or meets the training data distribution of the target model. In a specific embodiment, the model input instruction includes one or more preset fields, including but not limited to input fields, audio (e.g., music) fields, aspect ratio fields, video duration fields, and video clip number fields. The values ​​of the one or more preset fields can be determined based on the user instruction input by the user, and the fields that cannot be determined based on the user instruction can adopt default values ​​or be assigned through a natural language model.

[0042] In some embodiments, the user instructions are adjusted and / or expanded to obtain model input instructions that are adapted to the target model; wherein the adjustment includes adjusting the expression of the user instructions, and the expansion includes assigning values ​​to preset fields in the model input instructions.

[0043] In some embodiments, the user instruction can be expanded by a natural language model to obtain a model input instruction. The natural language model includes but is not limited to a Transformer-based model, an autoencoder-based model, a sequence-to-sequence model, a recursive neural network model, or a hierarchical model, but the present disclosure is not limited thereto. Exemplarily, the user instruction or the text generated based on the user instruction can be used as the input of the natural language model (for example, the prompt word "how to make a video about a puppy"), so that the natural language model recommends the relevant video style (such as "warm and cute"), video soundtrack, etc. according to the input prompt word, thereby expanding the user instruction to obtain the model input instruction text.

[0044] For example, the model input instruction may be as follows:

[0045] Input: Puppy

[0046] Music:**

[0047] Args: height: width = 3:4; duration, 13s

[0048] Segments:1

[0049] In some embodiments, keywords in user instructions can be extracted by a keyword extraction algorithm, and a model input instruction can be generated based on the keywords. The keyword extraction algorithm includes but is not limited to a word frequency-inverse document frequency algorithm, a graph-based sorting algorithm for text, a text analysis method based on a probabilistic graph model, a keyword extraction algorithm based on statistics and natural language processing technology, and a keyword extraction algorithm based on machine learning technology, which is not limited in the present disclosure.

[0050] In some embodiments, the user instructions input by the user have a preset format. For example, the instructions input by the user may be required to correspond one-to-one with preset fields, so that model input instructions can be generated based on the values ​​input by the user under these preset fields.

[0051] Step S130: Obtaining target information based on the model input instruction and the target model, wherein the target information includes video editing parameters for a target video editing tool.

[0052] In some embodiments, the target information includes structured text applicable to the target video editing tool, which includes video editing parameters readable by the target video editing tool. In some embodiments, the video editing parameters are used to control the style and / or time range of the video material presented in the video. In some embodiments, the video editing parameters include parameters involved in the video editing project, such as including one or more of the following: canvas size (such as width, height), video duration, transition mode, video resolution, number of video clips, identification of the video material (such as ID or key value), time range of the video material, rotation mode of the video material, and scaling ratio of the video material. Video material includes but is not limited to images, music, stickers, animations, filters, special effects, masks, texts, or video templates.

[0053] The video template includes a set of configured video materials (such as music, stickers, animations, filters, special effects, masks or texts). After the video template imports the original image or video, it can directly generate a video with music, stickers, animations, filters, special effects, masks or texts added. In some embodiments, the relevant video materials are arranged in a preset track or time range in the video template.

[0054] In some embodiments, the target information includes structured text describing the video editing project, which may describe relevant information of the video editing project (such as the above-mentioned video editing parameters or related information thereof) based on a protocol associated with the video editing tool. Exemplarily, the format of the target information is JSON (JavaScript Object Notation, JS object notation) format.

[0055] In some embodiments, the target model includes a text-to-text model, such as a large language model based on deep learning, including but not limited to an autoencoder-based model, a sequence-to-sequence model, a Transformer-based model, a recursive neural network model, and a hierarchical model, but the present disclosure is not limited thereto.

[0056] In some embodiments, when training a target model, the input instruction sample of the model can be obtained based on the label, title, description, soundtrack information (such as audio metadata) corresponding to the sample video or the keyword (query) used to search the sample video as the input of the model, and the target information sample corresponding to the sample video (such as structured text describing the editing project of the sample video) can be used as the expected output of the model to train the target model. Exemplarily, the parameters of the target model can be adjusted until the loss function converges based on the loss between the model prediction result obtained by inputting the input instruction sample into the target model and the expected output, but the present disclosure is not limited to this.

[0057] In some embodiments, the tags and descriptions of the video may include, but are not limited to, information about the theme, content, style, and filters of the video. In some video publishing scenarios, when a user publishes a video, he or she may add tags and descriptions related to the video to publish together with the video, and the tags and descriptions may serve to introduce, classify, or retrieve the video.

[0058] In some embodiments, at least one of the label, title, description, music information, and search keywords corresponding to the sample video can be directly used as an input instruction sample of the target model.

[0059] In some embodiments, at least one of the label, title, description, music information, or search keyword may be rewritten, for example, at least one of the label, title, description, music information, or search keyword is input into a preset language model and rewritten to obtain a rewritten input instruction sample. In this embodiment, by rewriting the input instruction sample for data enhancement, the convergence of the target model can be accelerated and the model training speed can be improved.

[0060] In some embodiments, the target video editing tool includes but is not limited to software, applications, applets, HTML (Hyper Text Markup Language) pages or services for video editing, which can be loaded locally on the terminal or in a cloud server, and the present disclosure does not make any limitations thereto.

[0061] Step S140: Generate a target video based on the target information, the video material and the target video editing tool; wherein the video material is associated with the video editing parameters.

[0062] The video material can be pre-stored in the target video editing tool, or downloaded from the server, or provided or selected by the user, or generated in real time based on the user's instructions, and this patent is not limited here. In some embodiments, the user can actively provide video material, for example, uploading multimedia material through the aforementioned robot chat session window as video material. In some embodiments, candidate video materials can also be provided to the user for selection, and the video material can be determined based on the user's selection. In some embodiments, the video material can be determined from the local resource library of the target video editing tool based on the user's instructions or the server obtains the video material. In some embodiments, the video material can be generated in real time based on the user's instructions, such as through a text graph model.

[0063] In this way, according to one or more embodiments of the present disclosure, by generating model input instructions based on user instructions, obtaining target information containing video editing parameters readable by the target video editing tool based on the model input instructions and the target model, and generating a target video based on the target information, the target video editing tool and the video material associated with the video editing parameters, it is ultimately possible to achieve one-click video generation based on user instructions, thereby significantly improving video editing efficiency.

[0064] In some embodiments, multimedia materials input by the user may also be received, a target tag of the multimedia materials may be obtained, and the model input instruction may be generated based on the user instruction and the target tag, wherein the target tag is used to indicate the theme and / or style of the video to be generated. In this embodiment, the multimedia materials input by the user may be further combined to judge the user's theme or style of the video to be generated. For example, in a specific application scenario, the natural language instruction input by the user may be only a text instruction "cut a video" and input a corresponding image (such as a photo of a kitten or puppy). In this case, the user's video editing intention may be judged in combination with the image input by the user.

[0065] In some embodiments, the target detection algorithm can be used to identify the multimedia material (e.g., one or one frame of an image) input by the user to obtain a target label / description (e.g., label "puppy") about the target object, and obtain the model input instruction based on the user instruction and the target label. Among them, the target detection algorithm includes but is not limited to a detection algorithm based on a multimodal Transformer model, a target detection algorithm based on a region-based convolutional neural network, a single-step multi-frame target detection algorithm, a target detection algorithm based on a retinal network, etc.

[0066] In some embodiments, the user instruction and the target label can be input together into a Vincent model, such as a natural language model based on deep learning, to obtain a model input instruction. The Vincent model includes but is not limited to a Transformer-based model, an autoencoder-based model, a sequence-to-sequence model, a recursive neural network model, and a hierarchical model, but the present disclosure is not limited thereto. Exemplarily, in the aforementioned specific application scenario, in response to the user inputting the user instruction "cut a video" and a puppy image, a model input instruction can be finally obtained based on the aforementioned target detection algorithm and the Vincent model, and the model input instruction includes information indicating that the theme of the video is "puppy" or "pet" and / or information indicating that the style of the video is "warm and cute", but the present disclosure is not limited thereto.

[0067] In some embodiments, the multimedia material input by the user includes image material (such as a static image, a dynamic image, or a video) or audio material, etc.

[0068] In some embodiments, step S120 includes:

[0069] Step A1: generating a first target instruction based on the user instruction, wherein the first target instruction includes information indicating a theme and / or style of a video to be generated;

[0070] Step A2: generating a second target instruction based on at least one of the user instruction, the first target instruction and the multimedia material input by the user, wherein the second target instruction includes information indicating the audio of the video to be generated;

[0071] Step A3: Determine the model input instruction based on the first target instruction and the second target instruction.

[0072] In this embodiment, a second target instruction is further generated based on the user instruction, the first target instruction and at least one of the multimedia materials input by the user to determine the user's audio usage intention (e.g., music matching intention), so that the model input instruction can include information for indicating the audio, thereby making the final generated target video conform to the user's music matching intention, thereby improving the efficiency of video music matching processing.

[0073] In some embodiments, the information indicating the audio for which the video is to be generated includes metadata of the audio, which is used to describe relevant information such as music, album or performer, including but not limited to song title, release date, track number, album name or performer name.

[0074] In some embodiments, a search keyword can be determined based on the user instruction, the first target instruction, and at least one of the multimedia materials input by the user; the target audio can be searched for based on the search keyword; and the second target instruction can be generated based on the target audio.

[0075] In some embodiments, the target detection algorithm can be used to identify the multimedia material input by the user to obtain a target label / description of the target object (e.g., the label "puppy"), and based on at least one of the user instruction, the first target instruction, and the multimedia material input by the user, a second target instruction is obtained. The target detection algorithm includes but is not limited to a detection algorithm based on a multimodal Transformer model, a target detection algorithm based on a region-based convolutional neural network, a single-step multi-frame target detection algorithm, a target detection algorithm based on a retinal network, and the like.

[0076] In some embodiments, based on the user instruction (or the first target instruction) and the target tag, an input text (such as a prompt word) suitable for a preset natural language processing model can be obtained, and the input text can be input into the natural language processing model to obtain a search keyword (such as a song name) recommended by the model, and the search keyword is used as a query word to search through a search engine to determine the target audio, and a second target instruction is generated based on the relevant information of the target audio (such as metadata). For example, in a specific application scenario, the user instruction input by the user is "cut a video about a puppy", then the preset natural language processing model can be asked "please recommend a piece of music related to a puppy" to obtain the search keyword; in another specific application scenario, the user instruction input by the user is "cut a video" and an image containing a puppy is input, then the puppy in the image can be identified by the aforementioned target detection algorithm, and finally the input text of the natural language processing model "please recommend a piece of music related to a puppy" is generated.

[0077] In some embodiments, the first target instruction is a model input instruction that does not include an audio instruction. The specific implementation of generating the first target instruction based on the user instruction in step A1 can refer to the above description of generating the model input instruction based on the user instruction, which will not be repeated here.

[0078] In some embodiments, step A2 includes: performing preset processing on the target audio, the preset processing including excerpt processing and / or segment processing; and generating the second target instruction based on the target audio and the preset processing result.

[0079] The excerpting process includes extracting a segment of a certain duration from the target audio. Exemplarily, the excerpting process can be performed based on the duration of the video or the duration of the video segment in the video.

[0080] The segmentation process includes segmenting the entire audio or the selected audio. Exemplarily, the audio can be segmented based on the number of image materials input by the user.

[0081] In some embodiments, the second target instruction includes one or more audio clipping parameters of an audio identifier (eg, an audio name, an audio ID), an audio time range, and a number of audio segments.

[0082] For example, reference Figure 2, the target detection algorithm can be used to identify the multimedia material input by the user to obtain the target label / description of the target object (for example, the label "puppy"), and the user instruction (or the first target instruction) and the target label are input into the natural language processing model to obtain the audio keywords (such as the name of the music) or template keywords (such as the name of the video template) recommended by the model. Finally, an audio search can be performed based on the audio keywords to obtain the target audio; or a template search can be performed based on the template keywords to obtain the target video template, and the audio used by the target video template is used as the target audio.

[0083] In some embodiments, step S130 includes:

[0084] Step B1: Input the model input instruction into the target model to obtain initial target information. In some embodiments, the initial target information is decoupled from the target video editing tool;

[0085] Step B2: Generate the target information based on the initial target information and a protocol associated with the target video editing tool.

[0086] In this embodiment, the output of the target model (i.e., the initial target information) is decoupled from the target video editing tool, so the target model is not only adapted to the target video editing tool, thereby improving the versatility and applicability of the target model. When the target video editing tool is used, the target information applicable to the target video editing tool can be further obtained based on the initial target information and the protocol associated with the target video editing tool.

[0087] In some embodiments, not all video editing parameters rely on target model prediction. In this regard, the values ​​of some video editing parameters in the initial target information output by the target model may be empty. These video editing parameters with empty values ​​may be assigned preset values ​​(such as default values) or values ​​adapted to the target video editing tool in the target information, but the present disclosure is not limited to this.

[0088] Accordingly, refer to Figure 3 According to an embodiment of the present disclosure, a video generating device 400 is provided, including:

[0089] The instruction receiving unit 301 is used to receive a user instruction input by a user;

[0090] An instruction generating unit 302, configured to generate a model input instruction based on the user instruction;

[0091] An information generating unit 303, configured to obtain target information based on the model input instruction and the target model, wherein the target information includes video editing parameters for a target video editing tool;

[0092] The video generation unit 304 is used to generate a target video based on the target information, the video material and the target video editing tool; wherein the video material is associated with the video editing parameters.

[0093] In some embodiments, the video generating device further comprises:

[0094] A material receiving unit, configured to receive multimedia material input by a user and determine a target object in the multimedia material;

[0095] The instruction generation unit is used to generate the model input instruction based on the user instruction and the target object, wherein the target object is used to indicate the theme and / or style of the video to be generated.

[0096] In some embodiments, the instruction generation unit includes:

[0097] A first instruction generating subunit is used to generate a first target instruction based on the user instruction, wherein the first target instruction includes information indicating a theme and / or style of a video to be generated;

[0098] A second instruction generating subunit is used to generate a second target instruction based on at least one of the user instruction, the first target instruction and the multimedia material input by the user, wherein the second target instruction includes information indicating the audio of the video to be generated;

[0099] The instruction generation subunit is used to determine the model input instruction based on the first target instruction and the second target instruction.

[0100] In some embodiments, the second instruction generation subunit includes:

[0101] A keyword subunit, configured to determine a search keyword based on at least one of the user instruction, the first target instruction, and the multimedia material input by the user;

[0102] An audio subunit, used to search for target audio based on the search keyword;

[0103] An instruction subunit is used to generate the second target instruction based on the target audio.

[0104] In some embodiments, the search keyword includes a template keyword; the audio subunit is used to perform a template search based on the template keyword to obtain a target video template, and use the audio used by the target video template as the target audio.

[0105] In some embodiments, the instruction subunit is used to perform preset processing on the target audio, and the preset processing includes excerpt processing and / or segment processing; and the second target instruction is generated based on the target audio and the preset processing result.

[0106] In some embodiments, the information generating unit includes:

[0107] An initial information generating subunit, used for inputting the model input instruction into the target model to obtain initial target information;

[0108] The target information generating subunit is used to generate the target information based on the initial target information and a protocol associated with the target video editing tool.

[0109] In some embodiments, the video generation device further includes a model training unit, including:

[0110] A sample acquisition subunit, used to acquire input instruction samples and target information samples corresponding to the sample video;

[0111] The training subunit is used to train the target model by taking the input instruction sample as the input of the target model and the target information sample as the expected output.

[0112] In some embodiments, the sample acquisition subunit is used to determine the input instruction sample based on at least one of a tag, a title, a description, audio information, and a search keyword associated with the sample video.

[0113] For the embodiments of the device, since they basically correspond to the method embodiments, the relevant parts refer to the partial description of the method embodiments. The device embodiments described above are merely illustrative, and the modules described as separation modules may or may not be separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present embodiment. Those of ordinary skill in the art can understand and implement without paying creative work.

[0114] Accordingly, according to one or more embodiments of the present disclosure, there is provided an electronic device, including:

[0115] at least one memory and at least one processor;

[0116] The memory is used to store program codes, and the processor is used to call the program codes stored in the memory to enable the electronic device to execute the video generation method provided according to one or more embodiments of the present disclosure.

[0117] Accordingly, according to one or more embodiments of the present disclosure, a non-transitory computer storage medium is provided, which stores a program code. The program code can be executed by a computer device to enable the computer device to perform a video generation method provided according to one or more embodiments of the present disclosure.

[0118] Reference below Figure 4 , which shows a schematic diagram of the structure of an electronic device (such as a terminal device or a server) 800 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0119] like Figure 4 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0120] Typically, the following devices may be connected to the I / O interface 805: input devices 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 808 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 800 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0121] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0122] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0123] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0124] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0125] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method of the present disclosure.

[0126] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0127] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0128] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, constitute a limitation on the unit itself.

[0129] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0130] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0131] According to one or more embodiments of the present disclosure, a video generation method is provided, including: receiving a user instruction input by a user; generating a model input instruction based on the user instruction; obtaining target information based on the model input instruction and a target model, where the target information includes video editing parameters for a target video editing tool; generating a target video based on the target information, video materials, and the target video editing tool; where the video materials are associated with the video editing parameters.

[0132] The method provided according to one or more embodiments of the present disclosure further includes: receiving multimedia materials input by a user and determining a target object in the multimedia materials; the generating a model input instruction based on the user instruction includes: generating the model input instruction based on the user instruction and the target object, where the target object is used to indicate the theme and / or style of the video to be generated.

[0133] According to one or more embodiments of the present disclosure, the generating a model input instruction based on the user instruction includes: generating a first target instruction based on the user instruction, where the first target instruction includes information indicating the theme and / or style of the video to be generated; generating a second target instruction based on at least one of the user instruction, the first target instruction, and the multimedia materials input by the user, where the second target instruction includes information indicating the audio of the video to be generated; determining the model input instruction based on the first target instruction and the second target instruction.

[0134] According to one or more embodiments of the present disclosure, the generating a second target instruction based on at least one of the user instruction, the first target instruction, and the multimedia materials input by the user includes: determining a search keyword based on at least one of the user instruction, the first target instruction, and the multimedia materials input by the user; searching for a target audio based on the search keyword; generating the second target instruction based on the target audio.

[0135] According to one or more embodiments of the present disclosure, the search keyword includes a template keyword; the searching for a target audio based on the search keyword includes: performing a template search based on the template keyword to obtain a target video template, and using the audio used in the target video template as the target audio.

[0136] According to one or more embodiments of the present disclosure, the generating a second target instruction based on the target audio includes: performing a preset process on the target audio, where the preset process includes an excerpt process and / or a segmentation process; generating the second target instruction based on the target audio and the preset process result.

[0137] According to one or more embodiments of the present disclosure, obtaining the target information based on the model input instruction and the target model includes: inputting the model input instruction into the target model to obtain initial target information; and generating the target information based on the initial target information and a protocol associated with the target video editing tool.

[0138] In some embodiments, the training method of the target model includes: obtaining input instruction samples and target information samples corresponding to the sample video; using the input instruction samples as the input of the target model and the target information samples as the expected output to train the target model.

[0139] In some embodiments, obtaining the input instruction sample corresponding to the sample video includes: determining the input instruction sample based on at least one of a tag, title, description, audio information, and search keyword associated with the sample video.

[0140] According to one or more embodiments of the present disclosure, a video generating device is provided, comprising: an instruction receiving unit for receiving a user instruction input by a user; an instruction generating unit for generating a model input instruction based on the user instruction; an information generating unit for obtaining target information based on the model input instruction and a target model, wherein the target information includes video editing parameters for a target video editing tool; a video generating unit for generating a target video based on the target information, a video material and the target video editing tool; wherein the video material is associated with the video editing parameters.

[0141] According to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one memory and at least one processor; wherein the memory is used to store program code, and the processor is used to call the program code stored in the memory so that the electronic device executes the video generation method provided according to one or more embodiments of the present disclosure.

[0142] According to one or more embodiments of the present disclosure, a non-transitory computer storage medium is provided, wherein the non-transitory computer storage medium stores program code, and when the program code is executed by a computer device, the computer device executes the video generation method provided according to one or more embodiments of the present disclosure.

[0143] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0144] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0145] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.

Claims

1. A video generation method, It is characterized in that include: Receive user instructions input by the user; generating a model input instruction based on the user instruction; Obtaining target information based on the model input instruction and the target model, wherein the target information includes video editing parameters for a target video editing tool; A target video is generated based on the target information, the video material and the target video editing tool; wherein the video material is associated with the video editing parameters.

2. The method according to claim 1, It is characterized in that Also includes: Receiving multimedia material input by a user, and determining a target object in the multimedia material; The generating the model input instruction based on the user instruction includes: generating the model input instruction based on the user instruction and the target object, wherein the target object is used to indicate the theme and / or style of the video to be generated.

3. The method according to claim 1, It is characterized in that The step of generating a model input instruction based on the user instruction comprises: generating a first target instruction based on the user instruction, wherein the first target instruction includes information indicating a theme and / or style of a video to be generated; Based on at least one of the user instruction, the first target instruction and the multimedia material input by the user, generating a second target instruction, wherein the second target instruction includes information indicating the audio of the video to be generated; The model input command is determined based on the first target command and the second target command.

4. The method according to claim 3, It is characterized in that The generating a second target instruction based on at least one of the user instruction, the first target instruction and the multimedia material input by the user comprises: Determine a search keyword based on at least one of the user instruction, the first target instruction, and the multimedia material input by the user; Based on the search keyword, search for the target audio; The second target instruction is generated based on the target audio.

5. The method according to claim 4, It is characterized in that The search keywords include template keywords; The searching for the target audio based on the search keyword includes: performing a template search based on the template keyword to obtain a target video template, and using the audio used by the target video template as the target audio.

6. The method according to claim 4, It is characterized in that The generating the second target instruction based on the target audio comprises: Performing a preset process on the target audio, wherein the preset process includes excerpt processing and / or segment processing; The second target instruction is generated based on the target audio and the preset processing result.

7. The method according to claim 1, It is characterized in that The obtaining target information based on the model input instruction and the target model includes: Inputting the model input instruction into the target model to obtain initial target information; The target information is generated based on the initial target information and a protocol associated with the target video editing tool.

8. The method according to claim 1, It is characterized in that The training method of the target model includes: Obtaining input instruction samples and target information samples corresponding to the sample video; The target model is trained by taking the input instruction sample as the input of the target model and the target information sample as the expected output.

9. The method according to claim 8, It is characterized in that The step of obtaining an input instruction sample corresponding to the sample video includes: The input instruction sample is determined based on at least one of a tag, a title, a description, audio information, and a search keyword associated with the sample video.

10. A video generating device, It is characterized in that include: An instruction receiving unit, used for receiving a user instruction input by a user; An instruction generating unit, used for generating a model input instruction based on the user instruction; an information generating unit, configured to obtain target information based on the model input instruction and the target model, wherein the target information includes video editing parameters for a target video editing tool; A video generation unit is used to generate a target video based on the target information, the video material and the target video editing tool; wherein the video material is associated with the video editing parameters.

11. An electronic device, It is characterized in that include: at least one memory and at least one processor; The memory is used to store program codes, and the processor is used to call the program codes stored in the memory to enable the electronic device to execute the method according to any one of claims 1 to 9.

12. A non-transitory computer storage medium, It is characterized in that The non-transitory computer storage medium stores a program code, and when the program code is executed by a computer device, the computer device executes the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Voice-driven intelligent picture book generation method and device, electronic equipment and storage medium

    CN121393444A